Stage 1 — Foundation:
- Target validation: iptables + systemctl/docker target-type gates
- Async run for long-running commands (sleep/wait/poll loops)
- audit_log.session_id plumbing (SQL, sqlcgen, 18 call sites)
Stage 2 — External agent observe (11 tools):
- get_dashboard_summary, get_ontology, list_checks, list_executions
- get_knowledge_revisions, get_knowledge_duplicates, get_knowledge_orphans
- list_knowledge_tags, list_entity_sessions, find_entities_by
- 3 resource templates: oikos://entity/{slug}, knowledge/{id}, execution/{id}
Stage 3 — Nomos reliability:
- complete_task(success) refused without verification (upgraded from warn)
- sessionHasPlan excludes replaced steps (forces propose_plan after reopen)
- Bash syntax validation in run() (rejects literal \n, flag-space typos)
- Scope gate in SOUL.md (ask before pivoting to unrelated subsystem)
Stage 4 — External agent act (9 mutation tools):
- ack_signal, resolve_signal, mute_signal, cancel_execution
- update_check, delete_knowledge, restore_knowledge
- merge_knowledge, rename_knowledge_tag
14 KiB
2026-08-04 — Unified plan: external MCP agents + Nomos reliability
Status: Complete. All 5 stages implemented (2026-08-04).
1. Context
Two audiences, two gaps:
| Audience | Current state | Goal |
|---|---|---|
| Nomos (internal agent) | 85% session success, but 38% plan adherence, 9 timeout failures, premature success, empty learning table | Reliable, self-correcting, leaves a trace |
| External agents (Claude, Goose, etc.) | 40 MCP tools — heavy on observe, light on act. Can't manage signals, checks, executions, or knowledge beyond upsert. |
Full Oikos surface: observe + act + curate |
The two plans share infrastructure (internal/mcp/server.go) and have cross-task
dependencies. This document merges them into one sequenced plan.
2. Cross-plan dependencies
Three audit tasks are prerequisites for external-agent mutation tools:
| Audit task | Enables | Why |
|---|---|---|
Task 3 — async run + timeout |
cancel_execution, list_executions with live status |
Without async run, every command >30s hits a client timeout. The agent loses track of the execution and can't cancel it. |
| Task 9 — execution linkage | list_executions filtered by session/entity |
nomos_plan_executions table is empty; audit_log.session_id is NULL. Without linkage, execution queries are blind. |
Task 2 — target validation in run |
All Phase 2 mutation tools | Can't let external agents mutate state on targets that can't execute the command (audit found qm run on lxc:dns). |
Two items were dropped from the original MCP expansion:
approve_execution/deny_execution— separation-of-duties violation. An MCP agent approving its own queued commands breaks the approval model. The correct fix is raisingauto_actforreversible_low(policy change, zero code — listed in the audit plan's out-of-scope). The existing assent-window path inrun()already auto-approves reversible-low commands.- Execution log streaming — MCP has no push model.
get_execution_statusreturns the latest output; full streaming stays WebUI-only.
3. Unified sequence (5 stages, 31 tasks)
Stage 1: Foundation — shared infrastructure (3 tasks)
These unlock everything downstream. Do first.
T1 — Target validation in run (audit Task 2)
internal/mcp/server.go: before dispatching a run command, validate prefix
against target type. qm/pct/pvesh → host:* only. systemctl/docker →
host:* or lxc:*. Mismatch returns error without execution.
T2 — Async run + 120s timeout (audit Task 3)
cmd/nomos: raise MCP client timeout 30s → 120s.
internal/mcp/server.go: when classifier detects sleep/wait/poll loops in
the command, start execution and return execution_id immediately. Agent polls
with get_execution_status. Execution continues server-side even if client
timeout.
T3 — Execution linkage (audit Task 9)
internal/mcp/server.go: when run creates an execution, write into
nomos_plan_executions (session_id + plan_step_seq + execution_id). Pass
session_id from MCP request headers into audit_log writes.
Stage 2: External agent observe (8 tasks)
All read-only MCP tools. Zero risk, ship fast. Unblocks external agents from understanding system state.
T4 — get_dashboard_summary
Fleet health counts, signals by severity, pending approvals, event rate. One
call instead of 4. Wraps existing GetDashboardSummary DB query.
T5 — get_ontology
Entity types, relationship types, lifecycle states, monitoring specs. Agents
need this to reason about the schema.
T6 — list_checks
Per-entity health checks with verdict, last run, probe output. Filter by entity
slug or check state.
T7 — list_executions
Cursor-paginated execution history. Filter by entity slug, status, risk class.
Depends on T3 (execution linkage) for session/entity filtering.
T8 — get_knowledge_revisions
Version history for a knowledge entity. Agent can see what changed and when.
T9 — get_knowledge_duplicates
Near-duplicate knowledge entries via trigram clustering. Wraps existing
knowledgeDuplicates query.
T10 — get_knowledge_orphans
Knowledge entries not linked to any entity. Agent can suggest cleanup.
T11 — list_knowledge_tags
All tags with counts. Agent can see the taxonomy.
T12 — list_entity_sessions (Phase 3 polish)
Active Nomos sessions linked to an entity. Depends on T3 (execution linkage).
T13 — find_entities_by (Phase 3 polish)
Search entities by attribute (IP, port, version, tag). More flexible than
list_entities (type/state only).
T14 — MCP resources
Expose entities, knowledge entries, and executions as MCP resource templates:
oikos://entity/{slug}, oikos://knowledge/{id}, oikos://execution/{id}.
MCP clients that support resources (Claude Desktop, Goose) can browse and attach
them to conversations.
Stage 3: Nomos reliability (6 tasks)
Fixes the worst failure modes found in the 5-session audit.
T15 — Prevent premature complete_task(success) (audit Task 1)
internal/mcp/tools.go: when complete_task with outcome=success, verify the
summary doesn't contradict known state. If goal mentions a URL but the last
probe shows non-200, log a warning.
nomos/SOUL.md: explicit rule — restate goal, verify every condition before
calling success. If any condition is "probably works," use outcome=partial.
T16 — Plan step integrity: require replaced_reason (audit Task 4)
Schema: add replaced_reason TEXT to session_plan_steps.
cmd/nomos: when agent emits update_plan_step(status=replaced), require
non-empty reason (enum: wrong_diagnosis, scope_change, blocked,
superseded, operator_override).
nomos/SOUL.md: explicit rule — complete or skip steps. Replacing all steps
with no reason is a session-quality violation.
T17 — Force propose_plan on session resume (audit Task 5)
cmd/nomos: when a session with status done/failed receives a new user
message, reset plan state. Require fresh propose_plan before any run.
The "must have plan" guard treats resumed sessions as plan-less.
T18 — Scope gate: surface session_questions on context switch (audit Task 6)
cmd/nomos system prompt: "When investigation leads to a subsystem unrelated to
the expressed goal, call session_questions before taking action."
nomos/SOUL.md: explicit rule — ask before pivoting.
T19 — Command syntax validation in run (audit Task 13)
internal/mcp/server.go: before executing, reject literal \n in commands,
backslash-continuation on last line, head - n/grep - i space-before-flag
typos, and && \n patterns from LLM formatting errors.
T20 — Stuck-session reaping (audit Task 14, from 2026-08-03 plan)
Reap sessions with closed_at IS NULL and no message in 30 minutes. Set
status=failed, outcome=failure.
Stage 4: External agent act (9 tasks)
Mutation MCP tools. Each writes audit log + emits event. Requires Stage 1 infrastructure (T1 target validation, T2 async run, T3 execution linkage). Follows existing direct-DB patterns — no HTTP API calls.
T21 — ack_signal(signal_id)
Acknowledge an open signal. Agent investigating an alert marks it acknowledged.
T22 — resolve_signal(signal_id, resolution?)
Resolve a signal with optional resolution note.
T23 — mute_signal(signal_id, duration?)
Temporarily mute a signal. Optional duration (default 1h).
T24 — cancel_execution(execution_id, reason)
Cancel a queued/running execution. Depends on T2 (async run returns
execution_id) and T3 (execution linkage for audit context).
T25 — update_check(check_id, enabled)
Enable/disable a health check. Agent suppresses a noisy probe.
T26 — delete_knowledge(knowledge_id)
Soft-delete a knowledge entry (move to trash, restorable).
T27 — restore_knowledge(knowledge_id)
Restore a trashed knowledge entry.
T28 — merge_knowledge(source_id, target_id)
Fold one knowledge entry into another. Source gets soft-deleted, content
appended to target.
T29 — rename_knowledge_tag(old_name, new_name)
Bulk-rename a tag across all knowledge entries.
Stage 5: Close the learning loop (5 tasks)
Turn execution data into persistent knowledge. Currently all learning tables are empty (0 classifications, 0 feedback, 0 patterns, 0 skills).
T30 — Auto-classify every run → classifications table (audit Task 11)
internal/mcp/server.go: run already calls classifyCommand. Write the
result to the classifications table (risk_class + route + patterns matched).
Currently 0 rows despite 1,884 executions.
T31 — Auto-upsert knowledge on session close (audit Task 7)
cmd/nomos: on complete_task (any outcome), auto-generate a knowledge entry:
title=<date>: <goal>, content with Outcome/Root cause/What was done/Unresolved
sections, tags=[session:<id>], linked to involved entities.
T32 — Auto-feedback on session close (audit Task 12)
cmd/nomos: on complete_task, generate a feedback entry: session_id,
outcome, observation, lesson, side_effects. Daily cron job reads recent feedback
and extracts patterns (recurring root causes, same-fix-applied-multiple-times).
T33 — Token tracking (audit Task 8)
cmd/nomos: after each LLM call, extract usage.total_tokens from the response
and write to agent_activity.token_count. Currently NULL for all rows.
T34 — Plan quality metric (audit Task 10)
cmd/nomos: at session close, compute completed_steps / total_steps (currently
~38%). Write as session attribute. Track over time to measure impact of T16+T17.
4. Validation
| Task | Test |
|---|---|
| T1 | run("lxc:dns", "qm stop 100") → error: "qm is a Proxmox host command" |
| T2 | run with sleep 45; echo done → returns execution_id immediately; get_execution_status eventually shows completed |
| T3 | After run, nomos_plan_executions has row linking session + step + execution |
| T4 | get_dashboard_summary() returns health counts, signal counts, approval count in one call |
| T5 | get_ontology() returns entity_types, relationship_types, lifecycle_states |
| T6 | list_checks(entity_slug="lxc:jellyfin") returns all checks with verdict + last run |
| T7 | list_executions(entity_slug="host:hubris", limit=10) returns cursor-paginated list |
| T8 | get_knowledge_revisions(id) returns ordered revision list with timestamps |
| T9 | get_knowledge_duplicates() returns clusters with similarity scores |
| T10 | get_knowledge_orphans() returns knowledge entries with zero entity links |
| T11 | list_knowledge_tags() returns {name, count} for all tags |
| T12 | list_entity_sessions("lxc:jellyfin") returns active sessions with goal + status |
| T13 | find_entities_by(ip="10.0.0.5") returns matching entities |
| T14 | MCP client can browse oikos://entity/* resources |
| T15 | Session with goal "make X reachable" where last ping shows 502 → complete_task(success) warns or rejects |
| T16 | update_plan_step(status=replaced) with no reason → rejected |
| T17 | Resumed session calls run before propose_plan → blocked |
| T18 | Agent pivots to unrelated subsystem → session_questions is called |
| T19 | run with head - n /etc/hosts → rejected with syntax error |
| T20 | Session idle for 30+ min with no closed_at → reaped (status=failed) |
| T21 | ack_signal(id) → signal status transitions to acknowledged, audit logged |
| T22 | resolve_signal(id, "fixed DNS") → resolved with note |
| T23 | mute_signal(id, 3600) → muted for 1 hour, auto-unmutes |
| T24 | cancel_execution(id, "wrong target") → execution cancelled, audit logged |
| T25 | update_check(id, false) → check disabled, scheduler stops probing |
| T26 | delete_knowledge(id) → soft-deleted (trashed), restorable |
| T27 | restore_knowledge(id) → restored from trash, reappears in list |
| T28 | merge_knowledge(src, dst) → src deleted, content appended to dst |
| T29 | rename_knowledge_tag("old", "new") → all entries updated |
| T30 | After any run, classifications has row with risk_class + route |
| T31 | complete_task → knowledge entry created automatically with session link |
| T32 | complete_task → feedback entry created; daily job extracts pattern if same root cause appears ≥3 times |
| T33 | agent_activity.token_count is non-NULL after LLM call |
| T34 | Session close writes plan_adherence attribute (% steps completed) |
5. Files touched
| File | Tasks |
|---|---|
internal/mcp/server.go |
T1, T2, T3, T19, T24, T30 |
internal/mcp/tools.go |
T4–T14, T21–T29 |
internal/mcp/discover.go |
(no changes — references for T10/T11 patterns) |
cmd/nomos/main.go (or config) |
T2 (timeout), T16, T17, T18, T31, T32, T33, T34 |
nomos/SOUL.md |
T15, T16, T18 |
internal/httpapi/ |
T3 (audit_log.session_id plumbing) |
| DB migrations | T16 (replaced_reason column) |
6. What stays WebUI-only
| Feature | Reason |
|---|---|
| FleetMap visual graph | Canvas rendering — not an MCP concern |
| uPlot metric charts | Raw data available via query_metrics/get_trend |
| Desktop shell, Cluck, App Store | Pure UI layer |
| SSE event streaming | MCP has no push model; polling covers it |
| Execution log streaming | get_execution_status returns latest output |
| Nomos session chat | MCP is a tool interface, not a chat agent |
| Knowledge wiki editor (revision browse, cleanup UI) | MCP tools expose the data + mutations; UI provides the editing experience |
| Approval queue with Approve/Deny buttons | Approvals stay operator-gated via WebUI/Matrix |
| Client enrollment flow | Enrollment is IP-gated, not an MCP tool |
7. Out of scope
approve_execution/deny_executionMCP tools — dropped. The fix is raisingauto_actforreversible_low(policy change, no code).delete_entityMCP tool — separate lifecycle management concern.- Per-client bearer tokens — open item tracked in CLIENTS.md. Until they
exist, external agents share the same
OIKOS_MCP_BEARER_TOKEN. - Exactly-once pattern extraction from feedback — T32 seeds the pipeline; the full pattern-mining algorithm (TF-IDF clustering, causal inference from event timelines) is future work.