# 2026-08-04 — Unified plan: external MCP agents + Nomos reliability **Status:** Complete. All 5 stages implemented (2026-08-04). --- ## 1. Context Two audiences, two gaps: | Audience | Current state | Goal | |---|---|---| | **Nomos** (internal agent) | 85% session success, but 38% plan adherence, 9 timeout failures, premature success, empty learning table | Reliable, self-correcting, leaves a trace | | **External agents** (Claude, Goose, etc.) | 40 MCP tools — heavy on observe, light on act. Can't manage signals, checks, executions, or knowledge beyond `upsert`. | Full Oikos surface: observe + act + curate | The two plans share infrastructure (`internal/mcp/server.go`) and have cross-task dependencies. This document merges them into one sequenced plan. --- ## 2. Cross-plan dependencies Three audit tasks are **prerequisites** for external-agent mutation tools: | Audit task | Enables | Why | |---|---|---| | **Task 3** — async `run` + timeout | `cancel_execution`, `list_executions` with live status | Without async `run`, every command >30s hits a client timeout. The agent loses track of the execution and can't cancel it. | | **Task 9** — execution linkage | `list_executions` filtered by session/entity | `nomos_plan_executions` table is empty; `audit_log.session_id` is NULL. Without linkage, execution queries are blind. | | **Task 2** — target validation in `run` | All Phase 2 mutation tools | Can't let external agents mutate state on targets that can't execute the command (audit found `qm` run on `lxc:dns`). | Two items were **dropped** from the original MCP expansion: - **`approve_execution` / `deny_execution`** — separation-of-duties violation. An MCP agent approving its own queued commands breaks the approval model. The correct fix is raising `auto_act` for `reversible_low` (policy change, zero code — listed in the audit plan's out-of-scope). The existing assent-window path in `run()` already auto-approves reversible-low commands. - **Execution log streaming** — MCP has no push model. `get_execution_status` returns the latest output; full streaming stays WebUI-only. --- ## 3. Unified sequence (5 stages, 31 tasks) ### Stage 1: Foundation — shared infrastructure (3 tasks) These unlock everything downstream. Do first. **T1 — Target validation in `run`** *(audit Task 2)* `internal/mcp/server.go`: before dispatching a `run` command, validate prefix against target type. `qm`/`pct`/`pvesh` → `host:*` only. `systemctl`/`docker` → `host:*` or `lxc:*`. Mismatch returns error without execution. **T2 — Async `run` + 120s timeout** *(audit Task 3)* `cmd/nomos`: raise MCP client timeout 30s → 120s. `internal/mcp/server.go`: when classifier detects `sleep`/`wait`/poll loops in the command, start execution and return `execution_id` immediately. Agent polls with `get_execution_status`. Execution continues server-side even if client timeout. **T3 — Execution linkage** *(audit Task 9)* `internal/mcp/server.go`: when `run` creates an execution, write into `nomos_plan_executions` (session_id + plan_step_seq + execution_id). Pass `session_id` from MCP request headers into `audit_log` writes. --- ### Stage 2: External agent observe (8 tasks) All read-only MCP tools. Zero risk, ship fast. Unblocks external agents from understanding system state. **T4 — `get_dashboard_summary`** Fleet health counts, signals by severity, pending approvals, event rate. One call instead of 4. Wraps existing `GetDashboardSummary` DB query. **T5 — `get_ontology`** Entity types, relationship types, lifecycle states, monitoring specs. Agents need this to reason about the schema. **T6 — `list_checks`** Per-entity health checks with verdict, last run, probe output. Filter by entity slug or check state. **T7 — `list_executions`** Cursor-paginated execution history. Filter by entity slug, status, risk class. Depends on T3 (execution linkage) for session/entity filtering. **T8 — `get_knowledge_revisions`** Version history for a knowledge entity. Agent can see what changed and when. **T9 — `get_knowledge_duplicates`** Near-duplicate knowledge entries via trigram clustering. Wraps existing `knowledgeDuplicates` query. **T10 — `get_knowledge_orphans`** Knowledge entries not linked to any entity. Agent can suggest cleanup. **T11 — `list_knowledge_tags`** All tags with counts. Agent can see the taxonomy. **T12 — `list_entity_sessions`** *(Phase 3 polish)* Active Nomos sessions linked to an entity. Depends on T3 (execution linkage). **T13 — `find_entities_by`** *(Phase 3 polish)* Search entities by attribute (IP, port, version, tag). More flexible than `list_entities` (type/state only). **T14 — MCP resources** Expose entities, knowledge entries, and executions as MCP resource templates: `oikos://entity/{slug}`, `oikos://knowledge/{id}`, `oikos://execution/{id}`. MCP clients that support resources (Claude Desktop, Goose) can browse and attach them to conversations. --- ### Stage 3: Nomos reliability (6 tasks) Fixes the worst failure modes found in the 5-session audit. **T15 — Prevent premature `complete_task(success)`** *(audit Task 1)* `internal/mcp/tools.go`: when `complete_task` with `outcome=success`, verify the summary doesn't contradict known state. If goal mentions a URL but the last probe shows non-200, log a warning. `nomos/SOUL.md`: explicit rule — restate goal, verify every condition before calling success. If any condition is "probably works," use `outcome=partial`. **T16 — Plan step integrity: require `replaced_reason`** *(audit Task 4)* Schema: add `replaced_reason TEXT` to `session_plan_steps`. `cmd/nomos`: when agent emits `update_plan_step(status=replaced)`, require non-empty reason (enum: `wrong_diagnosis`, `scope_change`, `blocked`, `superseded`, `operator_override`). `nomos/SOUL.md`: explicit rule — complete or skip steps. Replacing all steps with no reason is a session-quality violation. **T17 — Force `propose_plan` on session resume** *(audit Task 5)* `cmd/nomos`: when a session with status `done`/`failed` receives a new user message, reset plan state. Require fresh `propose_plan` before any `run`. The "must have plan" guard treats resumed sessions as plan-less. **T18 — Scope gate: surface `session_questions` on context switch** *(audit Task 6)* `cmd/nomos` system prompt: "When investigation leads to a subsystem unrelated to the expressed goal, call `session_questions` before taking action." `nomos/SOUL.md`: explicit rule — ask before pivoting. **T19 — Command syntax validation in `run`** *(audit Task 13)* `internal/mcp/server.go`: before executing, reject literal `\n` in commands, backslash-continuation on last line, `head - n`/`grep - i` space-before-flag typos, and `&& \n` patterns from LLM formatting errors. **T20 — Stuck-session reaping** *(audit Task 14, from 2026-08-03 plan)* Reap sessions with `closed_at IS NULL` and no message in 30 minutes. Set status=failed, outcome=failure. --- ### Stage 4: External agent act (9 tasks) Mutation MCP tools. Each writes audit log + emits event. Requires Stage 1 infrastructure (T1 target validation, T2 async run, T3 execution linkage). Follows existing direct-DB patterns — no HTTP API calls. **T21 — `ack_signal(signal_id)`** Acknowledge an open signal. Agent investigating an alert marks it acknowledged. **T22 — `resolve_signal(signal_id, resolution?)`** Resolve a signal with optional resolution note. **T23 — `mute_signal(signal_id, duration?)`** Temporarily mute a signal. Optional duration (default 1h). **T24 — `cancel_execution(execution_id, reason)`** Cancel a queued/running execution. Depends on T2 (async `run` returns `execution_id`) and T3 (execution linkage for audit context). **T25 — `update_check(check_id, enabled)`** Enable/disable a health check. Agent suppresses a noisy probe. **T26 — `delete_knowledge(knowledge_id)`** Soft-delete a knowledge entry (move to trash, restorable). **T27 — `restore_knowledge(knowledge_id)`** Restore a trashed knowledge entry. **T28 — `merge_knowledge(source_id, target_id)`** Fold one knowledge entry into another. Source gets soft-deleted, content appended to target. **T29 — `rename_knowledge_tag(old_name, new_name)`** Bulk-rename a tag across all knowledge entries. --- ### Stage 5: Close the learning loop (5 tasks) Turn execution data into persistent knowledge. Currently all learning tables are empty (0 classifications, 0 feedback, 0 patterns, 0 skills). **T30 — Auto-classify every `run` → `classifications` table** *(audit Task 11)* `internal/mcp/server.go`: `run` already calls `classifyCommand`. Write the result to the `classifications` table (risk_class + route + patterns matched). Currently 0 rows despite 1,884 executions. **T31 — Auto-upsert knowledge on session close** *(audit Task 7)* `cmd/nomos`: on `complete_task` (any outcome), auto-generate a knowledge entry: title=`: `, content with Outcome/Root cause/What was done/Unresolved sections, tags=`[session:]`, linked to involved entities. **T32 — Auto-feedback on session close** *(audit Task 12)* `cmd/nomos`: on `complete_task`, generate a `feedback` entry: session_id, outcome, observation, lesson, side_effects. Daily cron job reads recent feedback and extracts patterns (recurring root causes, same-fix-applied-multiple-times). **T33 — Token tracking** *(audit Task 8)* `cmd/nomos`: after each LLM call, extract `usage.total_tokens` from the response and write to `agent_activity.token_count`. Currently NULL for all rows. **T34 — Plan quality metric** *(audit Task 10)* `cmd/nomos`: at session close, compute `completed_steps / total_steps` (currently ~38%). Write as session attribute. Track over time to measure impact of T16+T17. --- ## 4. Validation | Task | Test | |---|---| | T1 | `run("lxc:dns", "qm stop 100")` → error: "qm is a Proxmox host command" | | T2 | `run` with `sleep 45; echo done` → returns `execution_id` immediately; `get_execution_status` eventually shows completed | | T3 | After `run`, `nomos_plan_executions` has row linking session + step + execution | | T4 | `get_dashboard_summary()` returns health counts, signal counts, approval count in one call | | T5 | `get_ontology()` returns entity_types, relationship_types, lifecycle_states | | T6 | `list_checks(entity_slug="lxc:jellyfin")` returns all checks with verdict + last run | | T7 | `list_executions(entity_slug="host:hubris", limit=10)` returns cursor-paginated list | | T8 | `get_knowledge_revisions(id)` returns ordered revision list with timestamps | | T9 | `get_knowledge_duplicates()` returns clusters with similarity scores | | T10 | `get_knowledge_orphans()` returns knowledge entries with zero entity links | | T11 | `list_knowledge_tags()` returns {name, count} for all tags | | T12 | `list_entity_sessions("lxc:jellyfin")` returns active sessions with goal + status | | T13 | `find_entities_by(ip="10.0.0.5")` returns matching entities | | T14 | MCP client can browse `oikos://entity/*` resources | | T15 | Session with goal "make X reachable" where last ping shows 502 → `complete_task(success)` warns or rejects | | T16 | `update_plan_step(status=replaced)` with no reason → rejected | | T17 | Resumed session calls `run` before `propose_plan` → blocked | | T18 | Agent pivots to unrelated subsystem → `session_questions` is called | | T19 | `run` with `head - n /etc/hosts` → rejected with syntax error | | T20 | Session idle for 30+ min with no `closed_at` → reaped (status=failed) | | T21 | `ack_signal(id)` → signal status transitions to acknowledged, audit logged | | T22 | `resolve_signal(id, "fixed DNS")` → resolved with note | | T23 | `mute_signal(id, 3600)` → muted for 1 hour, auto-unmutes | | T24 | `cancel_execution(id, "wrong target")` → execution cancelled, audit logged | | T25 | `update_check(id, false)` → check disabled, scheduler stops probing | | T26 | `delete_knowledge(id)` → soft-deleted (trashed), restorable | | T27 | `restore_knowledge(id)` → restored from trash, reappears in list | | T28 | `merge_knowledge(src, dst)` → src deleted, content appended to dst | | T29 | `rename_knowledge_tag("old", "new")` → all entries updated | | T30 | After any `run`, `classifications` has row with risk_class + route | | T31 | `complete_task` → knowledge entry created automatically with session link | | T32 | `complete_task` → feedback entry created; daily job extracts pattern if same root cause appears ≥3 times | | T33 | `agent_activity.token_count` is non-NULL after LLM call | | T34 | Session close writes `plan_adherence` attribute (% steps completed) | --- ## 5. Files touched | File | Tasks | |---|---| | `internal/mcp/server.go` | T1, T2, T3, T19, T24, T30 | | `internal/mcp/tools.go` | T4–T14, T21–T29 | | `internal/mcp/discover.go` | (no changes — references for T10/T11 patterns) | | `cmd/nomos/main.go` (or config) | T2 (timeout), T16, T17, T18, T31, T32, T33, T34 | | `nomos/SOUL.md` | T15, T16, T18 | | `internal/httpapi/` | T3 (audit_log.session_id plumbing) | | DB migrations | T16 (replaced_reason column) | --- ## 6. What stays WebUI-only | Feature | Reason | |---|---| | FleetMap visual graph | Canvas rendering — not an MCP concern | | uPlot metric charts | Raw data available via `query_metrics`/`get_trend` | | Desktop shell, Cluck, App Store | Pure UI layer | | SSE event streaming | MCP has no push model; polling covers it | | Execution log streaming | `get_execution_status` returns latest output | | Nomos session chat | MCP is a tool interface, not a chat agent | | Knowledge wiki editor (revision browse, cleanup UI) | MCP tools expose the data + mutations; UI provides the editing experience | | Approval queue with Approve/Deny buttons | Approvals stay operator-gated via WebUI/Matrix | | Client enrollment flow | Enrollment is IP-gated, not an MCP tool | --- ## 7. Out of scope - **`approve_execution` / `deny_execution` MCP tools** — dropped. The fix is raising `auto_act` for `reversible_low` (policy change, no code). - **`delete_entity` MCP tool** — separate lifecycle management concern. - **Per-client bearer tokens** — open item tracked in CLIENTS.md. Until they exist, external agents share the same `OIKOS_MCP_BEARER_TOKEN`. - **Exactly-once pattern extraction from feedback** — T32 seeds the pipeline; the full pattern-mining algorithm (TF-IDF clustering, causal inference from event timelines) is future work.