Monitoring coverage was 3 of 89 active entities. Three bugs, each hidden by discarded errors in checkdefaults: - writeCheck generated a fresh uuid, inserted the check entity ON CONFLICT (slug) DO NOTHING, then wrote a check_defs row referencing it. On any re-seed the slug already existed, the entity insert no-oped, and the FK violated — aborting the ingest transaction and surfacing as an unrelated failure several entities later. Re-seeding has been broken since; prod's coverage was frozen at its first successful seed. This is what TestSeedIngestIdempotentAndNoDuplicateEdges had been reporting. - shortSlug truncated to the last 8 chars, so all 21 ingress routes collapsed to ".network" and overwrote each other; service:jellyfin collided with lxc:jellyfin. - The ssh-script checker never read the `args` config checkdefaults wrote, so process_check.sh always ran without its unit name and returned "unknown". Coverage is now 75/89. Monitoring is declared per entity type in seeds/ontology.yaml and resolved through the is-a hierarchy, so a type can say it warrants nothing (site, lan, mesh, cluster) and never be reported as a gap. coverageSweep raises an `unmonitored` signal only where a type declares monitoring it lacks — 8 real gaps, no false positives. Also: - entity_types.attribute_schema was never ingested: the seed loader read "attribute_schema" but the YAML says "attributes", so all 60 types stored JSON null. - ListExecutions ignored its declared target/action/correlation_id filters and paginated on a non-unique target slug, dropping and repeating rows. - started_at was captured but only written at terminal state, so a running execution reported NULL for its whole life. The three MCP auto-run copies wrote no timing at all; they are now one autoRun helper. - SSH output was buffered to completion and discarded entirely on timeout. Both sshExec copies now stream through a shared execlog sink into execution_logs, and keep partial output when a command is cancelled. - executions.correlation_id was a random per-execution uuid that correlated nothing; it is now the chat session id, which is what lets the chat tail live output. - reversible_low had no auto-run branch despite policy declaring it unattended. Since computeCommandRisk never returns it, the class only arises when an agent declares it over a read_only command — so gating it penalised candor without adding safety. - backup-target gains a backup-freshness checker (portable find -mmin, since the first target is on macOS), resolving its host by walking backs-up-to backwards. The pre-deploy pg_dump is now a tracked backup target. UI: an Executions section on entity detail with live output tailing, and streamed output under a running `run` call in the chat timeline. Migrations 022-024. Ops.svelte and context.ts exclude execution.output from their refetch triggers, which would otherwise fire once a second per command. Co-Authored-By: Claude <noreply@anthropic.com>
65 lines
2.1 KiB
Go
65 lines
2.1 KiB
Go
package httpapi
|
|
|
|
import (
|
|
"testing"
|
|
"time"
|
|
|
|
"github.com/google/uuid"
|
|
)
|
|
|
|
// The cursor carries both created_at and entity_id because executions are
|
|
// ordered by the pair. created_at alone is not unique — several executions can
|
|
// share a millisecond — and paginating on a non-unique key silently drops or
|
|
// repeats rows at page boundaries. The previous cursor was the target slug,
|
|
// which is far less unique still: every execution against the same host shares
|
|
// it.
|
|
func TestExecutionCursorRoundTrips(t *testing.T) {
|
|
created := time.Date(2026, 7, 28, 9, 15, 30, 123456789, time.UTC)
|
|
id := uuid.MustParse("018f3a2b-0000-7000-8000-000000000042")
|
|
|
|
cursor := formatExecutionCursor(created, id)
|
|
|
|
gotTime, gotID, err := parseExecutionCursor(&cursor)
|
|
if err != nil {
|
|
t.Fatalf("parse: %v", err)
|
|
}
|
|
if !gotTime.Equal(created) {
|
|
t.Errorf("time round-trip: got %v, want %v", gotTime, created)
|
|
}
|
|
if *gotID != id {
|
|
t.Errorf("id round-trip: got %v, want %v", *gotID, id)
|
|
}
|
|
}
|
|
|
|
func TestExecutionCursorNanosecondsSurvive(t *testing.T) {
|
|
// Truncating to seconds would make the cursor ambiguous for executions
|
|
// started in the same second, which is the normal case for a plan whose
|
|
// steps run back to back.
|
|
a := time.Date(2026, 7, 28, 9, 15, 30, 1, time.UTC)
|
|
b := time.Date(2026, 7, 28, 9, 15, 30, 2, time.UTC)
|
|
id := uuid.New()
|
|
|
|
if formatExecutionCursor(a, id) == formatExecutionCursor(b, id) {
|
|
t.Error("cursors one nanosecond apart must not collide")
|
|
}
|
|
}
|
|
|
|
func TestExecutionCursorRejectsGarbage(t *testing.T) {
|
|
empty := ""
|
|
tm, id, err := parseExecutionCursor(&empty)
|
|
if err != nil || tm != nil || id != nil {
|
|
t.Errorf("empty cursor should mean 'no cursor', got %v/%v/%v", tm, id, err)
|
|
}
|
|
|
|
if tm, id, err := parseExecutionCursor(nil); err != nil || tm != nil || id != nil {
|
|
t.Errorf("nil cursor should mean 'no cursor', got %v/%v/%v", tm, id, err)
|
|
}
|
|
|
|
for _, bad := range []string{"nonsense", "2026-07-28T09:15:30Z", "notatime,018f3a2b-0000-7000-8000-000000000042", "2026-07-28T09:15:30Z,notauuid"} {
|
|
b := bad
|
|
if _, _, err := parseExecutionCursor(&b); err == nil {
|
|
t.Errorf("cursor %q should have been rejected", bad)
|
|
}
|
|
}
|
|
}
|