feat(observability): restore monitoring coverage, make gaps visible, stream executions

Monitoring coverage was 3 of 89 active entities. Three bugs, each hidden by
discarded errors in checkdefaults:

- writeCheck generated a fresh uuid, inserted the check entity ON CONFLICT
  (slug) DO NOTHING, then wrote a check_defs row referencing it. On any
  re-seed the slug already existed, the entity insert no-oped, and the FK
  violated — aborting the ingest transaction and surfacing as an unrelated
  failure several entities later. Re-seeding has been broken since; prod's
  coverage was frozen at its first successful seed. This is what
  TestSeedIngestIdempotentAndNoDuplicateEdges had been reporting.
- shortSlug truncated to the last 8 chars, so all 21 ingress routes collapsed
  to ".network" and overwrote each other; service:jellyfin collided with
  lxc:jellyfin.
- The ssh-script checker never read the `args` config checkdefaults wrote, so
  process_check.sh always ran without its unit name and returned "unknown".

Coverage is now 75/89. Monitoring is declared per entity type in
seeds/ontology.yaml and resolved through the is-a hierarchy, so a type can say
it warrants nothing (site, lan, mesh, cluster) and never be reported as a gap.
coverageSweep raises an `unmonitored` signal only where a type declares
monitoring it lacks — 8 real gaps, no false positives.

Also:
- entity_types.attribute_schema was never ingested: the seed loader read
  "attribute_schema" but the YAML says "attributes", so all 60 types stored
  JSON null.
- ListExecutions ignored its declared target/action/correlation_id filters and
  paginated on a non-unique target slug, dropping and repeating rows.
- started_at was captured but only written at terminal state, so a running
  execution reported NULL for its whole life. The three MCP auto-run copies
  wrote no timing at all; they are now one autoRun helper.
- SSH output was buffered to completion and discarded entirely on timeout.
  Both sshExec copies now stream through a shared execlog sink into
  execution_logs, and keep partial output when a command is cancelled.
- executions.correlation_id was a random per-execution uuid that correlated
  nothing; it is now the chat session id, which is what lets the chat tail
  live output.
- reversible_low had no auto-run branch despite policy declaring it
  unattended. Since computeCommandRisk never returns it, the class only arises
  when an agent declares it over a read_only command — so gating it penalised
  candor without adding safety.
- backup-target gains a backup-freshness checker (portable find -mmin, since
  the first target is on macOS), resolving its host by walking backs-up-to
  backwards. The pre-deploy pg_dump is now a tracked backup target.

UI: an Executions section on entity detail with live output tailing, and
streamed output under a running `run` call in the chat timeline.

Migrations 022-024. Ops.svelte and context.ts exclude execution.output from
their refetch triggers, which would otherwise fire once a second per command.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2026-07-28 13:51:14 +02:00
parent 873b00ac42
commit 1dca2cfd7a
39 changed files with 3105 additions and 273 deletions

View File

@@ -53,6 +53,18 @@
stepToggles.set(step.id, !stepOpen(step))
stepToggles = new Map(stepToggles)
}
// Pin each streaming output pane to its tail as chunks arrive. Keyed by
// tool id because several run entries can be on screen, though only the
// newest one is ever actually streaming.
let liveOutputEls = $state<Record<string, HTMLPreElement | null>>({})
$effect(() => {
for (const e of entries) {
if (!e.liveOutput) continue
const el = liveOutputEls[e.id]
if (el) el.scrollTop = el.scrollHeight
}
})
function toggleTool(id: string) {
if (expandedTools.has(id)) expandedTools.delete(id)
else expandedTools.add(id)
@@ -366,7 +378,7 @@
{#if expandedWithTools}
<div transition:slide={{ duration: 150 }} class="flex flex-col">
{#each item.tools as tool (tool.id)}
{@const tOpen = expandedTools.has(tool.id)}
{@const tOpen = expandedTools.has(tool.id) || !!tool.liveOutput}
<div class="relative" data-tl-id={tool.id}>
<!-- Branch stub: backbone → tool -->
<span
@@ -379,10 +391,12 @@
<button
type="button"
class="flex w-full items-center gap-1.5 py-1 pl-9 pr-3 text-left text-[11px] {tool.args ||
tool.detail
tool.detail ||
tool.liveOutput
? 'cursor-pointer hover:bg-muted/20'
: 'cursor-default'}"
onclick={() => (tool.args || tool.detail) && toggleTool(tool.id)}
onclick={() =>
(tool.args || tool.detail || tool.liveOutput) && toggleTool(tool.id)}
>
<span class="flex size-3 shrink-0 items-center justify-center">
{#if tool.status === 'running'}
@@ -428,6 +442,13 @@
tool.args
)}</pre>
{/if}
{#if tool.liveOutput}
<!-- Streaming while the command runs. Bound so it
can be pinned to the tail as chunks arrive. -->
<pre
bind:this={liveOutputEls[tool.id]}
class="max-h-36 overflow-auto whitespace-pre-wrap break-words rounded bg-muted/50 p-1.5 font-mono text-[9px] leading-relaxed text-muted-foreground">{tool.liveOutput}</pre>
{/if}
{#if tool.detail}
<pre
class="max-h-36 overflow-auto whitespace-pre-wrap break-words rounded bg-muted/50 p-1.5 font-mono text-[9px] leading-relaxed {tool.status ===