# 2026-08-04 — Session audit: agent reliability, plan system, and learning loop gaps **Status:** Done — all 12 tasks implemented, tested, deployed in v0.21.0. Verified in production with live session tests (classifications, feedback, token tracking, execution linkage all confirmed). **Reviewed sessions:** Past 5 completed plus ZimaOS continuation (268895a5) **Method:** Direct Postgres read of `agent_sessions`/`agent_messages`/ `agent_activity`/`session_plan_steps`/`executions`/`classifications`/`feedback`/ `patterns`/`skills`/`approvals`/`nomos_plan_executions`/`audit_log` on the prod mac-mini. Cross-referenced with `internal/audit/`, `internal/mcp/`, `internal/httpapi/`, `internal/policy/`. --- ## 1. Sessions audited | # | Session | Turns | Run calls | Failures | Plan steps (done/total) | Duration | Outcome | |---|---------|-------|-----------|----------|-------------------------|----------|---------| | S1 | Pocket-pascal deploy (a433b386) | 8 | 111 | 1 | 10/11 | 1.5h | Success (truncated by turn limit) | | S2 | SSH re-investigation (6f0ade08) | 4 | 42 | 0 | 0/4 | 11m | Success (all steps replaced) | | S3 | Webhook HMAC (0f509508) | 2 | 40 | 0 | 5/5 | 2m | Success (clean; best session) | | S4 | Check scripts (d458a5f8) | 23 | 96 | 4 | 0/14 | 2h | Success (3 plan gens; 0 steps done) | | S5 | ZimaOS outage (268895a5) | 18+10 | 209 | 8 | 0/5 | 2h30m | **Marked success; dashboard still broken** | **Headline:** 85% session success rate, but plan adherence is ~38% (15/39 steps ever reached `done`). The ZimaOS session was the worst: marked `success` while the dashboard was still down, then burned 81 more calls and 30 minutes chasing irrelevant DHCP reservations. --- ## 2. Cross-cutting findings ### F1 — MCP 30s client timeout kills every long-running command (P0) All **9** `run` failures across the audit are `Post "http://api:8090/mcp": context deadline exceeded` at exactly 30s. Commands with `sleep 30`, `qm shutdown`+wait, or async poll loops always hit this. The agent retries with longer sleeps and hits the same wall. **Root cause:** `cmd/nomos` MCP client uses a 30s timeout; the `run` tool blocks synchronously waiting for command completion. No async path exists for long-running commands. ### F2 — Premature success: `complete_task` fires before the goal is actually met (P0) Session 268895a5 called `complete_task(success)` at 18:29 with summary "zimaos returns 200." The endpoint was serving ttyd (terminal), not the ZimaOS dashboard. The agent conflated "HTTPS 200" with "dashboard works." The user had to resume the session. **Root cause:** No pre-completion validation. The agent can mark success with a summary that doesn't match reality. `complete_task` is a write-and-forget operation with no state check. ### F3 — Wrong target: `run` doesn't validate the target can execute the command (P0) In the ZimaOS continuation, `qm stop 100 --skiplock` was executed on `lxc:dns`. The command failed (`qm: command not found`) because the agent copied the command from a previous run but forgot to change the target. Similarly, `cmd/nomos` tried `docker exec` on `host:hubris` (no docker). **Root cause:** `run(target, command)` in `internal/mcp/server.go` doesn't validate that the target type can execute the given command. A simple allowlist would catch `qm`/`pct`/`pvesh` on non-host targets. ### F4 — Plan system is decorative: 62% of steps never reach `done` (P1) Across 5 sessions: 39 plan steps. 24 (62%) were `replaced`, 15 (38%) reached `done`. Session d458a5f8 had 3 complete plan regenerations with **zero** completed steps. The agent replaces plans instead of completing or explicitly skipping steps. **Root cause:** Plan steps carry status (`pending`/`running`/`done`/`failed`/ `skipped`/`blocked`/`replaced`) but `replaced` has no `replaced_reason` field. The model can silently replace every step and the system doesn't flag it. No constraint ties `propose_plan` to existing plan state. ### F5 — No plan on session resume: continuation sessions run ad-hoc (P1) The ZimaOS continuation (18:29→19:01) had 81 calls with **zero** `propose_plan` calls. The agent ran ad-hoc tool calls with no structure. **Root cause:** When a session resumes, the `must_have_plan` guard is already satisfied by the old (completed) plan. The agent doesn't re-plan on resume. ### F6 — Scope expansion / rabbit holes: agent chases irrelevant sub-goals (P2) In the ZimaOS continuation, the agent spent ~40 calls trying to fix Technitium DHCP reservations — a completely different subsystem from the goal ("make the dashboard reachable"). The ZimaOS dashboard hadn't started since July 19 (3-week-old issue), making the DHCP reservation effort moot. The agent never surfaced a question like "This is pre-existing — should I still fix DHCP?" **Root cause:** No scope gate. When the agent pivots to a subsystem unrelated to the stated goal, nothing stops it. The `session_questions` mechanism exists (2 calls in 193 sessions) but the model never uses it. ### F7 — Command generation errors: malformed bash from the LLM (P2) In the ZimaOS continuation, the agent generated: - `head - n` instead of `head -n` → bash syntax error - `echo "---" && \n curl ...` → literal `\n` in command → ambiguous redirect **Root cause:** The LLM generates bash commands inline in content blocks. No syntax validation, no escape-character handling. The `run` tool should reject malformed commands before execution. ### F8 — Learning pipeline completely empty (P4) | Table | Rows | |-------|------| | `classifications` | **0** | | `feedback` | **0** | | `patterns` | **0** | | `skills` | **0** | Despite 1,884 executions and 263 approvals, the system learns nothing from outcomes. The ZimaOS session discovered that the dashboard hadn't started since July 19 — this was never persisted. The HMAC trailing-newline discovery was persisted manually; if the agent forgot `upsert_knowledge`, it would be lost. ### F9 — Observability gaps (P3) | Gap | Detail | |-----|--------| | Token tracking | `agent_activity.token_count` is NULL for every row | | Execution linkage | `nomos_plan_executions` is empty; `audit_log.session_id` is null | | Plan quality | No metric for step completion rate (currently 38%) | | MCP timeout rate | No counter for `run` calls that hit the client timeout | ### F10 — Execution success rate is 74% (P2) 1,884 executions: 1,400 completed (74%), 295 failed (15.6%), 152 cancelled (8%), 37 denied (2%). One in four execution attempts doesn't complete. --- ## 3. Improvement plan (ordered by impact/effort) ### Task 1 — Prevent premature `complete_task(success)` *(fixes F2)* **`internal/mcp/tools.go`** — In the `complete_task` handler, when `outcome=success`: require the summary field to contain a verifiable state assertion. Minimum: if the goal mentions a URL, check that the summary doesn't contradict known state. Lightweight: log a warning if the summary says "returns 200" but the last `ping_service` or `run` result says otherwise. **Coach:** `nomos/SOUL.md` — explicit rule: *"Before calling complete_task(success), restate the user's original goal in your own words and verify each condition. If any condition is 'probably works' rather than 'verified,' ask the operator or set outcome=partial."* ### Task 2 — Target validation in `run` *(fixes F3)* **`internal/mcp/server.go`** — In the `run` handler, before dispatching: validate that the command prefix matches the target type. ``` pct/qm/pvesh/iptables → only host:* targets systemctl/docker → host:* or lxc:* targets curl/nmap/ss → any target ``` If mismatched, return a clear error: *"Cannot run `qm` on lxc:dns — `qm` is a Proxmox host command. Use target host:hubris or host:strong."* Do not classify or execute. **Test:** `TestClassifyCommand_WrongTarget` → commands with host-only prefixes on LXC targets return error without execution. ### Task 3 — Raise MCP client timeout; add async path for long-running commands *(fixes F1)* **`cmd/nomos`** — Raise the MCP client timeout from 30s to 120s. **`internal/mcp/server.go`** — For `run` commands that the classifier determines will exceed the client timeout (presence of `sleep`, `wait`, `timeout` in the command), return immediately with an `execution_id` and status `running`. The agent already has `get_execution_status` — use it: 1. Classify the command; if it contains `sleep`, `wait`, or shell constructs that imply polling, flag it as `async_potential`. 2. Start the command, return the `execution_id` immediately. 3. Agent polls with `get_execution_status(execution_id)`. 4. If the client timeout is hit mid-poll, the execution continues on the server — it's not lost. **Test:** `run` with `sleep 60; echo done` on host:hubris → returns immediately (not 30s timeout), `get_execution_status` eventually returns `completed`. ### Task 4 — Plan step integrity: require `replaced_reason` on replacement *(fixes F4)* **`session_plan_steps` migration** — Add `replaced_reason TEXT` column. **`cmd/nomos`** — When the agent emits `update_plan_step` with `status=replaced`, require a non-empty `replaced_reason`. Valid reasons: `wrong_diagnosis`, `scope_change`, `blocked`, `superseded`, `operator_override`. **Coach:** `nomos/SOUL.md` — explicit rule: *"Complete (status=done) or explicitly skip (status=skipped) steps. Use status=replaced only when the entire plan generation is wrong; include the reason. Replacing all steps with no reason is a session-quality violation."* ### Task 5 — Force `propose_plan` on session resume *(fixes F5)* **`cmd/nomos`** — When a session with status `done` or `failed` receives a new user message, reset the plan state: clear step status, require a new `propose_plan` call before any `run` calls. The "must have plan" guard should consider the resumed session as plan-less until a fresh `propose_plan` is called. **Guard:** `set_goal` + `propose_plan` must be called before any `run` in a resumed session. Reuse the existing "No plan — call set_goal then propose_plan" error from `internal/mcp/server.go`. ### Task 6 — Scope gate: surface `session_questions` on context switch *(fixes F6)* **Coach:** `nomos/SOUL.md` — explicit rule: *"Before pivoting to a subsystem not mentioned in the user's goal, ask via session_questions. Example: 'The dashboard logs show it hasn't started since July 19. Do you want me to debug the dashboard service itself [A], skip it and just stabilize the IP [B], or stop here [C]?'"* **`cmd/nomos` prompt** — Add to the system prompt: *"When the investigation leads to a subsystem or root cause unrelated to the expressed goal, surface a session_question before taking action."* ### Task 7 — Auto-upsert knowledge on session close *(fixes F8)* **`cmd/nomos`** — On `complete_task` (any outcome: success, partial, failure), auto-generate a knowledge entry: ``` title: ": " content: "## Outcome\n\n## Root cause\n\n## What was done\n\n## What was left\n" tags: [session:] about: [entities involved] ``` This ensures every session leaves a trace regardless of whether the agent remembered to call `upsert_knowledge`. ### Task 8 — Token tracking *(fixes F9)* **`cmd/nomos`** — After each LLM call, extract `usage.prompt_tokens`, `usage.completion_tokens`, `usage.total_tokens` from the response and write to `agent_activity.token_count`. Currently the field exists but is never populated (NULL for all rows). ### Task 9 — Execution linkage *(fixes F9)* **`internal/mcp/server.go`** — When `run` creates an execution, write a row into `nomos_plan_executions` linking `session_id`, `plan_step_seq`, and `execution_id`. **`internal/mcp/server.go`** — Pass `session_id` (from MCP request headers) into `audit_log` writes. Currently `audit_log.session_id` is NULL — the `createAuditLog` function in `internal/httpapi/impl.go` receives the correlation_id but not the session_id from the MCP path. ### Task 10 — Plan quality metric *(fixes F9)* **`cmd/nomos`** — At session close, compute: `completed_steps / total_steps_per_plan` (currently ~38%). Log as a metric or write as a session attribute. Track over time to measure plan-adherence improvements from Tasks 4+5. ### Task 11 — Auto-classify every `run` call *(fixes F8)* **`internal/mcp/server.go`** — The `run` handler already calls the classifier (`classifyCommand` in `internal/policy/command.go`) to determine risk_class and approval route. Write the result to the `classifications` table. Currently the table is empty (0 rows) despite 1,884 executions being classified. ### Task 12 — Auto-feedback on session close *(fixes F8)* **`cmd/nomos`** — On `complete_task`, generate a `feedback` entry: ``` session_id: outcome: observation: lesson: side_effects: ``` **`cmd/oikos`** — Add a daily cron or scheduler job that reads recent `feedback` entries and extracts `patterns` (recurring root causes, same-fix applied multiple times, known-broken services). Seed the pattern table. ### Task 13 — Command syntax validation in `run` *(fixes F7)* **`internal/mcp/server.go`** — Before executing a `run` command, do lightweight bash syntax validation: ``` - Reject literal \n in commands (should be ; or &&) - Reject commands where the last line ends with \ (backslash-continuation) but no next line - Warn on common typos: "head - n", "grep - i", spaces before flags - Reject `&& \n` patterns (the LLM sometimes inserts literal \n between && chains) ``` ### Task 14 — Stuck-session reaping (from prior plan; re-confirmed) This session exhibited the same idle zombie pattern (23da10db — no closed_at, status `failed` but outcome `failure`). Task 5 from the 2026-08-03 plan is still open. Copying here for completeness. --- ## 4. Recommended sequence ``` P0 (blocks operational waste): 1 → 2 → 3 P1 (fixes plan architecture): 4 → 5 P2 (cognitive guardrails): 6 → 7 → 13 P3 (observability): 8 → 9 → 10 P4 (learning loop): 11 → 12 ``` Sequence rationale: Tasks 1-3 stop the worst outcomes (premature success, wrong-target execution, MCP timeouts). Tasks 4-5 make the plan system actually useful instead of decorative. Tasks 6-7 add guardrails that prevent the ZimaOS rabbit-hole class of failure. Tasks 8-10 give us visibility into whether any of the previous tasks are working. Tasks 11-12 close the learning loop. --- ## 5. Validation | Task | Test | |------|------| | 1 | Session with goal "make X reachable" where last ping shows 502 → `complete_task(success)` is rejected or warns | | 2 | `run("lxc:dns", "qm stop 100")` → error: "qm is a Proxmox host command" | | 3 | `run` with `sleep 45; echo done` → returns execution_id immediately, `get_execution_status` shows final result | | 4 | `update_plan_step(status=replaced)` with no reason → rejected; with reason → accepted | | 5 | Resumed session calls `run` before `propose_plan` → blocked: "No plan — call propose_plan" | | 6 | Agent pivots to unrelated subsystem → `session_questions` is called before action | | 7 | `complete_task` → knowledge entry created automatically with session link | | 8 | `agent_activity.token_count` is non-NULL after any LLM call | | 9 | `nomos_plan_executions` has rows linking session + step + execution | | 10 | Session close writes `plan_adherence` attribute (step-completion %) | | 11 | `classifications` table has 1 row per `run` call with risk_class + route | | 12 | `complete_task` → auto `feedback` entry; daily pattern job finds recurring issues | | 13 | `run` with `head - n /etc/hosts` → rejected with clear error about malformed command | --- ## 6. Out of scope / open questions - Whether to raise `auto_act` from `off` for `reversible_low` actions (separate policy decision; would reduce approval pileup without code changes). - Whether to add a `delete_entity` MCP tool for lifecycle management (separate from this reliability plan). - The exact TTL for stuck-session reaping (30 min recommended, confirmed in 2026-08-03 plan). - Whether `run` async mode should be opt-in (command contains sleep/wait) or universal (every run returns immediately, agent always polls). Recommend opt-in for now — most commands complete in <5s.