Files
oikos/plans/done/2026-07-29-health-check-reality-and-knowledge-graph.md
dtoro 428f4fe945
Some checks failed
ci / build-test (push) Has been cancelled
ci / docker-build (push) Has been cancelled
ci / web (push) Has been cancelled
Desktop App / Build Linux (amd64) (push) Has been cancelled
Desktop App / Attach to Release (push) Has been cancelled
docs(plans): add status notes missed by the rename (chat-full-polish, health-check-reality)
git mv staged the pre-edit index content; the status edits to these two
files landed in the working tree but not the archive commit. Amending the
status now so the archived copies reflect Implemented.
2026-08-03 22:53:29 +02:00

26 KiB
Raw Blame History

Plan: Make health reflect reality + complete the knowledge graph

Status: Implemented (v0.14.x0.16.x). Shipped across c9a00a9 (per-entity monitoring override), a3914eb/8eb1ca2 (process check opt-in + probe_unit), 0929c17 (discover_infra_drift), and the vm-status/layered-probe/route-via- proxmox-host decisions now in project memory. Created 2026-07-29.

Context

ws:mac-mini reports health down despite being the healthy control-plane host. Investigation showed the problem is systemic, not local: 49 enabled checks report down, almost all ssh-script, because the resource/updates probes assume scripts are deployed at /opt/oikos/checks/ AND root SSH works on every target — both false for macOS, non-enrolled LXCs, and mesh-only entities. The knowledge graph also has real gaps (unmodeled TLS certs, empty skills table, seed drift, a capped topology view).

The DB is the source of truth; live state was verified via the REST API (Authorization: Bearer $OIKOS_MCP_BEARER_TOKEN, token in oikos-api-1 container env) and docker exec oikos-postgres-1 psql. Direct psql access is available for cleanup.

Decisions (confirmed with operator)

  1. Monitoring philosophy: make checks work everywhere — via the proven pct exec/ qm guest exec host-routing the MCP run tool already uses (no per-guest SSH keys), plus deploy the check scripts INTO each guest and make them macOS-aware. Hosts/workstations use direct SSH with the correct per-target user.
  2. Canonical host-hop accesspct exec/qm guest exec through the proxmox host is the ONLY execution path for any LXC/VM command (scheduler + MCP run + agent). Direct guest SSH is retired for execution; lan_ip stays for network probes only. (A1.)
  3. Auto-provision monitoring for new entities — wire script-deploy + the health-check-answering lifecycle gate into entity creation so any entity Nomos creates becomes monitorable with zero manual steps (Track E).
  4. Lifecycle gate: skip monitoring for deprecated/destroyed targets — no permanent false alarms from retired things.
  5. Knowledge graph: address ALL gaps — model TLS certificates, fix dns-zone gap, re-export seeds, seed skills, raise graph cap.
  6. Read-only audit skill — a read_only operator skill discovers live infra and diffs it against the DB graph, producing a ranked drift report; the operator acts on findings via existing lifecycle runbooks. No auto-fix. (Track F.)

Findings (evidence)

A. Health-check reality gaps (49 checks down)

Root cause is a routing mismatch, verified live (tests use the scheduler's own key -i /etc/oikos/ssh_key, not a default-key test):

The MCP run tool already reaches every guest correctly via resolveExecTarget (internal/mcp/server.go:582): resolve the proxmox host (attributes.hosthosts edge → hubris default), SSH there, run pct exec <pve_id> -- bash -c 'echo <b64> | base64 -d | bash' (VMs: qm guest exec). That path needs no per-guest lan_ip, no per-guest authorized_keys, no per-guest sshd.

The scheduler's checkSSHScript does not use it — it SSHes directly to each entity's own resolved address (internal/scheduler/scheduler.go:758, internal/checkdefaults/defaults.go:376 resolveHost) and runs /opt/oikos/checks/<script>. That is the bug. Decomposed by class:

Class Targets (verified) Root cause
Guests reached wrong lxc:rclone (mesh-only, no lan_ip), lxc:nfs-export (192.168.8.200: ssh port 22 timeout — no sshd), lxc:teddycloud (key not authorized — "not a homelab client"), lxc:grimmory/romm/seanime (strong: pct-exec reachable, scripts not inside) scheduler SSHes the guest directly; should route via proxmox host pct exec like resolveExecTarget. rclone is correctly parented on hubris (hosts edge verified) and IS reachable via pct exec 132 — the mesh fqdn is a red herring.
macOS host ws:mac-mini (5 resource/updates checks down) root SSH disabled (macOS); user: dtoro never read by resolver (defaults.go:406 reads attrs["ssh"]["user"] only); scripts not deployed; scripts Linux-only
External / mesh-only host:netbird-vps (no lan_ip; mesh unreachable from container) resolveHost picks mesh IP over public_ipv4 (defaults.go:376); sshd also "locked to hubris pubkey"
Dead route ingress:secrets.hubris.network http down service:secrets-issuance is deprecated but its ingress check still enabled — no lifecycle gate
ICMP-blocked vm:haos ping down while up HAOS blocks ICMP

Working (prove the host-SSH model is sound): host:hubris, host:strong SSH with the scheduler key → SCRIPTS_PRESENT; lxc:gitea direct-SSH → SCRIPTS_PRESENT (it's a homelab client with root key + scripts). So the host hop is the reliable path.

Parentage verified correct (all hosts edges checked in DB): strong guests on strong, hubris guests on hubris. No misplaced parents — the gap is routing + in-guest script deployment, not topology.

Health aggregation itself is correct: WorstHealthForTarget (internal/db/sqlcgen/operations.sql.go:1472) = worst enabled check. One failing ssh-script drags an otherwise-healthy entity to down.

B. Dead/stale data

  • 24 orphan check_defs + check entities, slugs ^check:(ping|ssh-script|disk):[0-9a-f]{8}$ (e.g. check:ssh-script:0d31fdd1), enabled=false, last_health=NULL, state=NULL. Leftover from the old shortSlug() collision bug (fixed in defaults.go:263).
  • service:secrets-issuance = deprecated; ingress:secrets.hubris.network still routes to it and alarms permanently.

C. Knowledge-graph gaps

  • TLS certificates unmodeled: certificate type + uses-certificate edge + cert-expiry checker all exist, but 0 certificate entities. Cert expiry is invisible.
  • dns-zone declares monitoring: [dns] (seeds/ontology.yaml:382) but no dns checker exists → every zone is an unmonitored signal.
  • Seed drift: 23 dns-record entities in DB, 0 in seeds/inventory.yaml.
  • skills table = 0 despite .agents/skills/*/SKILL.md on disk (runbooks = 15).
  • Graph capped at 500 nodes (internal/httpapi/impl.go:27 graphNodeCap = 500); 299 execution + 87 task rows dominate, so /graph is not a faithful topology view.

Work breakdown

Track A — Make ssh-script checks work everywhere (route through the proxmox host)

Core idea: stop having the scheduler SSH each guest directly. Reuse the MCP run tool's proven resolveExecTarget pattern — reach every LXC/VM through its proxmox host via pct exec/qm guest exec. This fixes rclone (no lan_ip), nfs-export (no sshd), teddycloud (no key), and every strong guest in one stroke, because the host hop already has working root SSH. Hosts/workstations keep direct SSH.

A1. Canonicalize host-hop as the ONLY execution path for LXC/VM (the real fix + simplification).

Principle: never SSH directly into a guest to run a command. Every LXC/VM command execution — scheduler checks, the MCP run tool, and the agent — routes through the owning proxmox host via pct exec <pve_id> -- ... (VMs: qm guest exec). One SSH credential per host (root key, already authorized on hubris/strong), no per-guest keys, sshd, or lan_ip needed for execution. Verified this works: pct exec 132 reaches rclone; the MCP run tool already does it for every guest (internal/mcp/server.go:582).

  • Network probes (http/ping) keep hitting the guest's lan_ip/URL directly — they don't execute inside the guest, so they're unaffected. For LXCs all checks are ssh-script, so they all route via the host; lan_ip becomes optional metadata, not a monitoring prereq.
  • Extract resolveExecTarget/resolveProxmoxHostSlug out of internal/mcp into a shared package (e.g. internal/remote) so the scheduler's checkSSHScript (internal/scheduler/scheduler.go:710) and checkBackupFreshness (backup.go:79, the other direct-SSH path) and the MCP run tool share ONE resolver. Today they diverge — the scheduler SSHes guests directly (broken), MCP host-hops (works).
  • checkSSHScript/checkBackupFreshness: when the target is lxc:/vm:, resolve the proxmox host and wrap the invocation as pct exec <pve_id> -- bash -c 'echo <b64> | base64 -d | bash' (VMs: the qm guest exec form at server.go:625). For host:/ws: keep direct SSH (they ARE the host).
  • Risk class: config_mutation (changes how probes reach every guest) → operator approval. Verify one LXC end-to-end (rclone) before fanning out.

A2. Deploy check scripts INTO guests (via pct push), not just to the host.

  • Verified: scripts exist on hubris/strong (the hosts) but NO_SCRIPTS inside grimmory, romm, seanime, rclone. A pct exec-routed check still runs inside the guest, so the scripts must live in the guest.
  • Add a fleet-deploy tool (tools/deploy-checks.sh): for each LXC, from its proxmox host, pct push <id> checks/<script> /opt/oikos/checks/<script> + chmod 755 (loop the checks/*.sh set). For VMs, scp/agent; for hosts/workstations, run checks/install.sh.
  • Backfill once now (all guests + mac-mini). See Track E for the automated version.

A3. Fix per-target SSH user + resolver (hosts/workstations only).

  • internal/checkdefaults/defaults.go:406 resolveSSHUser: also read top-level attrs["user"] (workstations carry user: dtoro, not ssh.user). Returns dtoro for mac-mini. Re-derive mac-mini's check_defs so config carries the user.
  • Do NOT enable root SSH on mac-mini — use dtoro (keeps macOS hardening).

A4. macOS-aware check scripts.

  • checks/cpu_check.sh:5 top -bn1 (Linux) → branch on uname -s == Darwin (top -l 1/sysctl). Same for memory_check.sh, load_check.sh, disk_usage_check.sh (df differs), updates_check.sh (already apt-guarded; on Darwin report healthy with security_updates=0 or read softwareupdate --list).
  • Each must still emit {"health":..,"metrics":{..}} JSON (internal/scheduler/scheduler.go:767).

A5. Reachability for external/mesh-only hosts.

  • internal/checkdefaults/defaults.go:376 resolveHost: prefer public_ipv4 over mesh IP for standalone-server/external so host:netbird-vps (82.165.190.79) is probeable. Note sshd is "locked to hubris pubkey" (inventory.yaml:89) — either add the scheduler key or proxy via hubris. Confirm before assuming direct SSH works.
  • ws:republic-laptop: roving laptop on mesh only. ping-down when asleep is real; keep ping-only and accept transient down, or set monitoring: none. (Decision in Open Questions.)
  • lxc:rclone no longer a special case — handled by A1's pct routing.

A6. ICMP-blocked VMs.

  • vm:haos ping down while up: optional tcp-ping fallback in checkPing (internal/scheduler/scheduler.go:604) for VMs that block ICMP, gated by an attribute. Lower priority — confirm haos blocks ICMP before building.

Track B — Lifecycle monitoring gate

B1. Skip monitoring for deprecated/destroyed targets.

  • Disable (set enabled=false) and skip-scheduling check_defs whose target entity state ∈ {deprecated,destroyed}.
  • Implement by joining target state in ListEnabledCheckDefs (internal/db/sqlcgen/operations.sql.go, the ListEnabledCheckDefs query) — exclude rows whose target is retired — or in a housekeeping sweep (internal/scheduler/scheduler.go:302) that disables them. Prefer the query filter (no write needed at runtime).
  • Matches policy.yaml lifecycle philosophy (destroyed.refuse: all); extend the comment.
  • Effect: dead ingress:secrets.hubris.network alarm goes silent automatically.

Track C — Dead-data cleanup

C1. Delete 24 orphan check_defs + check entities.

  • Direct SQL (have psql access): delete check_defs then entities matching slug ~ '^check:(ping|ssh-script|disk):[0-9a-f]{8}$'. Confirm state IS NULL / enabled=false first (already verified).
  • Wrap as a one-shot migration or scripts/cleanup-orphan-checks.sh. Risk class: read the rows first; this is config_mutation → operator approval.

C2. Retire the secrets route.

  • With B1 in place the alarm silences. Optionally set ingress:secrets.hubris.networkdeprecated/destroyed and remove its routes-to edge to service:secrets-issuance (or keep as archaeology). Decide with operator.

C3. Destroy 7 stray test LXCs (active cruft in the graph).

  • DB shows these with live hosts edges on strong, never cleaned up: lxc:preflight-test, lxc:preflight-test2, lxc:test-autocontinue, lxc:test-decompose3, lxc:test-livewatch, lxc:test-livewatch2, lxc:typetype.
  • First confirm they're really gone in Proxmox (pct list on strong); if so, set their entity state → destroyed (move to archaeology) and drop the hosts edges. If any container still exists, destroy via pct destroy first (destructive → approval).
  • They currently generate checks and pollute the graph/health view.

Track D — Knowledge graph

D1. Model TLS certificates.

  • Seed certificate entities (one per *.hubris.network route, or per Caddy-managed cert) + uses-certificate edges from each ingress-route.
  • Source real data: read Caddy's cert store (LXC 121) expiry via the existing cert-expiry checker's discovery, or seed from Caddyfile and backfill expires live.
  • Wires the cert-expiry checker (internal/scheduler/scheduler.go, cert-expiry kind) against real entities instead of nothing.

D2. dns-zone monitoring gap.

  • seeds/ontology.yaml:382: change dns-zone monitoring: [dns]monitoring: none with a comment "no dns checker yet; revisit when implemented". Stops the per-zone unmonitored noise. Re-seed.

D3. Re-export seeds to fix drift.

  • Run oikos export (or the export endpoint) so the 23 runtime dns-record entities + other runtime-created topology land in seeds/inventory.yaml. Diff, review, commit.

D4. Seed skills from disk.

  • Ingest .agents/skills/*/SKILL.md as skill entities (mirror how runbooks seed → 15 exist). Add to the knowledge seed ingest path (internal/db/seed.go) or a one-shot ingest. get_skills() then returns data.

D5. Raise graph node cap.

  • internal/httpapi/impl.go:27 graphNodeCap = 500 → raise (e.g. 5000) and/or paginate /api/v1/graph. Ensure the query stays performant (it already limits by default; confirm no full-table risk). Optionally exclude cognition rows (execution/task) from the default topology view via a ?layer=infrastructure filter so infra isn't crowded out.

Track E — Auto-provision monitoring when a new entity is created

Goal: the operator's request — "make sure this is handled automatically in the future when the agent creates new entities." Today ensureDefaultChecks (internal/httpapi/default_checks.go:9) writes check_defs on entity creation but does NOT make the target probe-ready (no script deploy, no host-routing). Its own comment admits the gap. A new entity should become monitorable with zero manual steps.

E1. Hook script-deploy into entity creation / provisioning.

  • Extend ensureDefaultChecks (called on entity create, default_checks.go) so that, after writing check_defs, it also ensures the target can answer:
    • LXC/VM: pct push the checks/*.sh set into the guest from its proxmox host (reuse the host resolution from A1). Idempotent (skip if present + unchanged).
    • host/workstation: ensure scripts at /opt/oikos/checks/ (run checks/install.sh over SSH; locally on mac-mini).
  • Because the check itself is routed via pct exec (Track A), no per-guest SSH key or sshd is needed — host hop + in-guest scripts are the only prerequisites, both now automated. mac-mini still needs its dtoro key (A3) once.

E2. Tie into the lifecycle provisioning → active gate.

  • The ontology already requires health-check-answering for provisioning → active (seeds/ontology.yaml:39, checked by internal/ontology/validate.go:167).
  • Make that gate actually run one check against the new entity and require a non-down verdict before the transition is allowed. This closes the loop: an entity isn't "active" (and isn't trusted for blast-radius/auto decisions) until monitoring proves it answers.

E3. Re-run on re-seed / attribute change.

  • checkdefaults.Ensure already re-derives check config from the seed on re-ingest (internal/checkdefaults/defaults.go:301, seed wins, enabled preserved). Mirror that for script deploy: when pve_id/host/address attributes change, re-target the check and re-deploy scripts to the new guest.

Net effect: a new LXC provisioned by Nomos (via pct_create, which registers the entity + hosts edge, internal/httpapi/actuator.go:615) automatically gets script-pushed + check_defs + a passing health-check-answering gate before going active.

Track F — Read-only knowledge-graph audit skill

Goal: the operator's request — a skill that auto-discovers live infra and validates the knowledge graph (entities, parentage, checks, scripts, seeds, certs) against reality, producing a ranked drift report. Read-only; no auto-fix — the operator routes each finding to the relevant lifecycle runbook.

Precedent (reuse, don't duplicate): existing drift/quality machinery is fragmented and knowledge-content focused. The audit orchestrates these + fills the topology/script gaps:

  • internal/httpapi/knowledge_drift.go — duplicate notes, orphan notes, tag splits (already endpoints).
  • internal/scheduler/coverage.go coverageSweep — unmonitored declared types (re-use its logic/SQL).
  • MCP discovery: list_lxcs (internal/mcp/tools.go:478), get_lxc_state, list_entities, get_relations, http_get. These already enumerate live LXC/VM state from the proxmox host.

F1. Add an on-demand audit primitive (MCP tool + endpoint).

  • New MCP tool audit_knowledge_graph (+ GET /api/v1/audit/drift) — read-only, runs the discovery+diff in one pass and returns a ranked report. Each finding = {category, severity, entities, evidence, suggested_runbook}.
  • Discovery sources (all via the canonical host-hop / existing tools): pct list + pct config on hubris & strong (guests, net0 IP, onboot state); qm list (VMs); Caddy admin API / Caddyfile (routes → certs); docker ps on compose hosts; the checks/*.sh set vs what's deployed at /opt/oikos/checks/ per target.
  • Report categories (the gaps this investigation found):
    1. Ghost entities — in DB but not in Proxmox (e.g. stray lxc:test-*).
    2. Missing entities — in Proxmox/Caddy/docker but no DB entity.
    3. Misplaced parenthosts edge disagrees with where the guest actually runs (the rclone class — though rclone's parent is correct; this catches real migrations).
    4. Orphan/dead checkscheck_defs whose target is deprecated/destroyed, or random-slug orphans (^check:(ping|ssh-script|disk):[0-9a-f]{8}$).
    5. Undeployed scripts — checks expect /opt/oikos/checks/<script> but it's absent in the guest (the strong-guest/rclone class).
    6. Unmonitored declared types — reuse coverageSweep SQL (dns-zone today, agents).
    7. Seed drift — entities/edges in DB but not in seeds/inventory.yaml (23 dns-records), via oikos export diff.
    8. Unmodeled certs — Caddy serves a cert with no certificate entity + uses-certificate edge.
    9. Knowledge rot — delegate to the existing knowledge_drift endpoints (duplicates/orphans/tags).

F2. Author the skill.

  • .agents/skills/knowledge-graph-audit/SKILL.md — front-matter risk_class: read_only, inputs: [scope?], verification: "drift report returns ok". Body: run audit_knowledge_graph, read the ranked report, and for each category point at the remediation runbook (lifecycle-deprecate-node, lifecycle-destroy-node, config-change-deploy for scripts, lifecycle-migrate-node for parents, this plan's tracks for cert/seed/graph-cap work). No mutating steps.
  • Seed a matching runbook:knowledge-graph-audit entity in seeds/knowledge.yaml (bound by applies_to_type) so search_knowledge/get_skills surface it (also fixes the empty-skills-table gap, Track D4).

F3. Optional: periodic sweep (later). Wrap categories 4/6 as a scheduler housekeeping sweep that raises drift signals, mirroring coverageSweep. Out of scope for this plan unless the operator wants continuous drift signals; the on-demand skill is the deliverable.

Risk class: read_only. The audit only reads (pct list/config, docker ps, Caddy API, DB selects, an oikos export to a temp file). No writes. Safe to run unattended.


Validation

After each track, verify via API (read-only, no approval):

  • GET /api/v1/entities/ws:mac-minihealth ∈ {healthy,degraded} (not down).
  • GET /api/v1/entities/lxc:rclonehealth healthy (proves pct-routing through hubris; rclone currently unreachable because it resolves to a mesh fqdn). Verify its checks now route via pct exec 132 on hubris.
  • Strong guests (lxc:grimmory, lxc:romm, lxc:seanime) → ssh-script checks healthy after scripts pushed inside + routed via strong's pct exec.
  • GET /api/v1/checks?include_disabled=falsedown count drops from 49 to the genuinely-down set (republic-laptop asleep, real outages only). Re-run the per-class table.
  • GET /api/v1/entities/service:secrets-issuance + its ingress → no enabled check.
  • Orphan cleanup: SELECT count(*) FROM check_defs cd JOIN entities e ON e.id=cd.entity_id WHERE e.slug ~ '^check:(ping|ssh-script|disk):[0-9a-f]{8}$'; → 0.
  • Test LXCs (C3): SELECT count(*) FROM entities WHERE slug IN ('lxc:preflight-test','lxc:test-livewatch',...) AND state<>'destroyed'; → 0.
  • Provision a throwaway LXC via Nomos → it auto-gets scripts + check_defs + passes health-check-answering before reaching active (E1/E2).
  • GET /api/v1/entities?type=certificate&limit=1 → >0; cert-expiry checks created.
  • GET /api/v1/entities?type=skill&limit=50 → >0.
  • GET /api/v1/graph node count > 500 (or infra fully represented with a layer filter).
  • oikos export diff shows dns-record entities present; git diff seeds/inventory.yaml.
  • Scheduler logs: checkdefaults: declared check not created warnings gone for dns-zone.
  • Canonical access (A1): no scheduler code path SSHes a guest directly — grep -rn "sshExec" internal/scheduler shows it only for host:/ws: targets; LXC/VM go through the shared pct exec/qm guest exec resolver.
  • Audit skill (F1/F2): audit_knowledge_graph MCP tool returns a ranked report with the 9 categories; running it against current state reproduces this plan's findings (orphan checks, stray test LXCs, undeployed scripts, seed drift, 0 certs). The skill is read-only — confirm it performs no DB writes (audit-log shows only reads).

Unit/integration tests to add/update:

  • internal/checkdefaults / shared internal/remote resolver: LXC/VM check routes via pct exec/qm guest exec to the resolved proxmox host; resolver reads top-level user; public_ipv4 preferred for standalone-server (defaults_test.go).
  • internal/scheduler: ListEnabledCheckDefs excludes deprecated/destroyed targets (new test); coverage_test.go still green; sshExec no longer called for guest slugs.
  • macOS script branches: assert JSON shape unchanged on Darwin (shunit2 or a smoke run).
  • E1: new-entity creation triggers script push (mock pct/SSH in test).
  • F1: audit_knowledge_graph against a fixture DB+mock discovery returns the expected category counts (ghost, missing, orphan, undeployed, drift).

Risks

  • Canonical host-hop (A1) makes each proxmox host the single SSH dependency for all its guests. This is already true (pct exec requires the host up) and is a net improvement (one credential vs many), but a host outage now fails all its guest checks together — which is the correct blast radius (guests are unreachable when their host is down).
  • Routing change (A1) alters how probes reach every guest — config_mutation. Verify one LXC end-to-end (rclone via pct exec 132) before fanning out. Extracting resolveExecTarget into a shared package keeps scheduler + MCP in lockstep.
  • Script push into guests (A2/E1) writes to guest filesystems — config_mutation. Idempotent + content-checked; never clobber a same-named operator script without diffing.
  • mac-mini root SSH: do NOT enable root login; use dtoro (A3) — keeps macOS hardening.
  • netbird-vps sshd locked to hubris pubkey: may need the scheduler key added or proxying via hubris; confirm before assuming direct SSH works (A5).
  • health-check-answering gate (E2) could block a legitimately-active entity whose only working check is ICMP-blocked (haos). Allow the gate to pass on any non-down reachable probe, or grant an operator override.
  • Audit skill (F1) discovers infra via pct/Caddy/docker reads — keep it strictly read_only; ensure discovery commands are in the read-only allowlist (no state change).
  • Seed re-export can surface large diffs (cognition entities) — scope export to topology entities, or review carefully before commit. Bump VERSION per repo rule.
  • Graph cap raise: large node sets may slow the graph render; pair with a layer filter.

Open questions (none blocking; confirm during implementation)

  • republic-laptop: mesh-only roving laptop — keep ping-only (accept transient down) or monitoring: none? (A5)
  • secrets ingress: keep as archaeology or destroy the route? (C2)
  • certificates: seed statically from Caddyfile, or auto-discover live from Caddy store? (D1)
  • netbird-vps: add scheduler key to its sshd, or always proxy through hubris? (A5)
  • Audit discovery for docker hosts/stacks: enumerate via docker ps, or model compose stacks only? (F1)

Suggested order

A1 (canonical host-hop routing — unblocks rclone + all guests) → A2 (push scripts into guests) → A3 → A4 (mac-mini) → A5 → E1/E2 (automate for new entities) → B1 → C1 → C3 → C2 → D2 (quick, silences dns noise) → F1/F2 (audit skill — also validates the above worked) → D1 → D4 → D3 → D5. Validate after each track.