Monitoring coverage was 3 of 89 active entities. Three bugs, each hidden by discarded errors in checkdefaults: - writeCheck generated a fresh uuid, inserted the check entity ON CONFLICT (slug) DO NOTHING, then wrote a check_defs row referencing it. On any re-seed the slug already existed, the entity insert no-oped, and the FK violated — aborting the ingest transaction and surfacing as an unrelated failure several entities later. Re-seeding has been broken since; prod's coverage was frozen at its first successful seed. This is what TestSeedIngestIdempotentAndNoDuplicateEdges had been reporting. - shortSlug truncated to the last 8 chars, so all 21 ingress routes collapsed to ".network" and overwrote each other; service:jellyfin collided with lxc:jellyfin. - The ssh-script checker never read the `args` config checkdefaults wrote, so process_check.sh always ran without its unit name and returned "unknown". Coverage is now 75/89. Monitoring is declared per entity type in seeds/ontology.yaml and resolved through the is-a hierarchy, so a type can say it warrants nothing (site, lan, mesh, cluster) and never be reported as a gap. coverageSweep raises an `unmonitored` signal only where a type declares monitoring it lacks — 8 real gaps, no false positives. Also: - entity_types.attribute_schema was never ingested: the seed loader read "attribute_schema" but the YAML says "attributes", so all 60 types stored JSON null. - ListExecutions ignored its declared target/action/correlation_id filters and paginated on a non-unique target slug, dropping and repeating rows. - started_at was captured but only written at terminal state, so a running execution reported NULL for its whole life. The three MCP auto-run copies wrote no timing at all; they are now one autoRun helper. - SSH output was buffered to completion and discarded entirely on timeout. Both sshExec copies now stream through a shared execlog sink into execution_logs, and keep partial output when a command is cancelled. - executions.correlation_id was a random per-execution uuid that correlated nothing; it is now the chat session id, which is what lets the chat tail live output. - reversible_low had no auto-run branch despite policy declaring it unattended. Since computeCommandRisk never returns it, the class only arises when an agent declares it over a read_only command — so gating it penalised candor without adding safety. - backup-target gains a backup-freshness checker (portable find -mmin, since the first target is on macOS), resolving its host by walking backs-up-to backwards. The pre-deploy pg_dump is now a tracked backup target. UI: an Executions section on entity detail with live output tailing, and streamed output under a running `run` call in the chat timeline. Migrations 022-024. Ops.svelte and context.ts exclude execution.output from their refetch triggers, which would otherwise fire once a second per command. Co-Authored-By: Claude <noreply@anthropic.com>
91 lines
3.4 KiB
Go
91 lines
3.4 KiB
Go
package ontology
|
|
|
|
// Monitoring resolution: which check kinds an entity type warrants.
|
|
//
|
|
// Coverage is not uniform. A `service` warrants an HTTP probe; a `site` is a
|
|
// physical location with nothing to probe; a `dns-zone` warrants a check whose
|
|
// checker does not exist yet. Collapsing those three into "has no check_def"
|
|
// is what made the fleet's monitoring gap invisible — 86 of 89 active entities
|
|
// had no check, and staleSweep's INNER JOIN against check_defs meant none of
|
|
// them could ever be marked stale.
|
|
//
|
|
// So the declaration lives on the entity TYPE, in seeds/ontology.yaml, and
|
|
// resolves through the same is-a hierarchy the validator already walks:
|
|
// declaring `monitoring: [ping, resource]` on abstract `machine` covers
|
|
// proxmox-host, standalone-server, workstation and appliance.
|
|
|
|
// MonitoringResolution is the outcome of resolving a type's monitoring
|
|
// declaration. The three states are deliberately distinguishable:
|
|
//
|
|
// Declared=false — nobody in the chain said anything. An ontology
|
|
// gap: report it, but as a modelling problem
|
|
// rather than as a fleet monitoring problem.
|
|
// Declared=true, len(0) — explicitly unmonitorable. Working as intended;
|
|
// never raise an `unmonitored` signal for it.
|
|
// Declared=true, len(n) — these kinds are expected to exist.
|
|
type MonitoringResolution struct {
|
|
Kinds []string
|
|
|
|
// Declared reports whether anything in the chain (or the layer default)
|
|
// settled the question.
|
|
Declared bool
|
|
|
|
// Source names the type that supplied the answer — the type itself, an
|
|
// ancestor, or "" when the layer default applied. Useful in log lines
|
|
// that explain why an entity has the checks it has.
|
|
Source string
|
|
}
|
|
|
|
// None reports an explicit "this type is not monitored".
|
|
func (m MonitoringResolution) None() bool {
|
|
return m.Declared && len(m.Kinds) == 0
|
|
}
|
|
|
|
// Wants reports whether the type expects a check of this kind.
|
|
func (m MonitoringResolution) Wants(kind string) bool {
|
|
for _, k := range m.Kinds {
|
|
if k == kind {
|
|
return true
|
|
}
|
|
}
|
|
return false
|
|
}
|
|
|
|
// Monitoring resolves the check kinds a type warrants, walking parent types
|
|
// until one carries a declaration.
|
|
//
|
|
// Types outside the infrastructure layer (governance, cognition, meta) fall
|
|
// back to an implicit "none": a signal, an approval, a document and a person
|
|
// are records, not running things. That default keeps ~20 record types out of
|
|
// the ontology without needing an explicit `monitoring: none` on each, while
|
|
// still treating an undeclared *infrastructure* type as a genuine gap — those
|
|
// are the ones somebody should have made a decision about. A non-infrastructure
|
|
// type that really is probeable (agent, which serves a gateway on :8092) just
|
|
// declares its kinds explicitly and wins on the first rule.
|
|
func (t *TypeTree) Monitoring(typ string) MonitoringResolution {
|
|
seen := map[string]bool{}
|
|
for cur := typ; cur != ""; cur = t.Types[cur].Parent {
|
|
info, ok := t.Types[cur]
|
|
if !ok {
|
|
break
|
|
}
|
|
if seen[cur] {
|
|
break // cycle guard — ingest rejects cycles, belt and braces
|
|
}
|
|
seen[cur] = true
|
|
|
|
if info.Monitoring != nil {
|
|
return MonitoringResolution{
|
|
Kinds: *info.Monitoring,
|
|
Declared: true,
|
|
Source: cur,
|
|
}
|
|
}
|
|
}
|
|
|
|
if info, ok := t.Types[typ]; ok && info.Layer != "infrastructure" {
|
|
return MonitoringResolution{Declared: true}
|
|
}
|
|
return MonitoringResolution{}
|
|
}
|