Files
oikos/plans/done/2026-08-04-session-audit-agent-reliability.md
dtoro 85f0bb67fa
Some checks failed
ci / build-test (push) Has been cancelled
ci / docker-build (push) Has been cancelled
ci / web (push) Has been cancelled
Desktop App / Build Linux (amd64) (push) Has been cancelled
Desktop App / Attach to Release (push) Has been cancelled
docs(plans): move 2026-08-04 session audit plan to done (v0.21.0 shipped)
2026-08-04 23:43:28 +02:00

370 lines
16 KiB
Markdown

# 2026-08-04 — Session audit: agent reliability, plan system, and learning loop gaps
**Status:** Done — all 12 tasks implemented, tested, deployed in v0.21.0. Verified in production
with live session tests (classifications, feedback, token tracking, execution linkage all confirmed).
**Reviewed sessions:** Past 5 completed plus ZimaOS continuation (268895a5)
**Method:** Direct Postgres read of `agent_sessions`/`agent_messages`/
`agent_activity`/`session_plan_steps`/`executions`/`classifications`/`feedback`/
`patterns`/`skills`/`approvals`/`nomos_plan_executions`/`audit_log` on the prod
mac-mini. Cross-referenced with `internal/audit/`, `internal/mcp/`,
`internal/httpapi/`, `internal/policy/`.
---
## 1. Sessions audited
| # | Session | Turns | Run calls | Failures | Plan steps (done/total) | Duration | Outcome |
|---|---------|-------|-----------|----------|-------------------------|----------|---------|
| S1 | Pocket-pascal deploy (a433b386) | 8 | 111 | 1 | 10/11 | 1.5h | Success (truncated by turn limit) |
| S2 | SSH re-investigation (6f0ade08) | 4 | 42 | 0 | 0/4 | 11m | Success (all steps replaced) |
| S3 | Webhook HMAC (0f509508) | 2 | 40 | 0 | 5/5 | 2m | Success (clean; best session) |
| S4 | Check scripts (d458a5f8) | 23 | 96 | 4 | 0/14 | 2h | Success (3 plan gens; 0 steps done) |
| S5 | ZimaOS outage (268895a5) | 18+10 | 209 | 8 | 0/5 | 2h30m | **Marked success; dashboard still broken** |
**Headline:** 85% session success rate, but plan adherence is ~38% (15/39 steps
ever reached `done`). The ZimaOS session was the worst: marked `success` while
the dashboard was still down, then burned 81 more calls and 30 minutes chasing
irrelevant DHCP reservations.
---
## 2. Cross-cutting findings
### F1 — MCP 30s client timeout kills every long-running command (P0)
All **9** `run` failures across the audit are `Post "http://api:8090/mcp":
context deadline exceeded` at exactly 30s. Commands with `sleep 30`, `qm
shutdown`+wait, or async poll loops always hit this. The agent retries with
longer sleeps and hits the same wall.
**Root cause:** `cmd/nomos` MCP client uses a 30s timeout; the `run` tool
blocks synchronously waiting for command completion. No async path exists for
long-running commands.
### F2 — Premature success: `complete_task` fires before the goal is actually met (P0)
Session 268895a5 called `complete_task(success)` at 18:29 with summary "zimaos
returns 200." The endpoint was serving ttyd (terminal), not the ZimaOS
dashboard. The agent conflated "HTTPS 200" with "dashboard works." The user had
to resume the session.
**Root cause:** No pre-completion validation. The agent can mark success with a
summary that doesn't match reality. `complete_task` is a write-and-forget
operation with no state check.
### F3 — Wrong target: `run` doesn't validate the target can execute the command (P0)
In the ZimaOS continuation, `qm stop 100 --skiplock` was executed on
`lxc:dns`. The command failed (`qm: command not found`) because the agent
copied the command from a previous run but forgot to change the target.
Similarly, `cmd/nomos` tried `docker exec` on `host:hubris` (no docker).
**Root cause:** `run(target, command)` in `internal/mcp/server.go` doesn't
validate that the target type can execute the given command. A simple
allowlist would catch `qm`/`pct`/`pvesh` on non-host targets.
### F4 — Plan system is decorative: 62% of steps never reach `done` (P1)
Across 5 sessions: 39 plan steps. 24 (62%) were `replaced`, 15 (38%) reached
`done`. Session d458a5f8 had 3 complete plan regenerations with **zero**
completed steps. The agent replaces plans instead of completing or explicitly
skipping steps.
**Root cause:** Plan steps carry status (`pending`/`running`/`done`/`failed`/
`skipped`/`blocked`/`replaced`) but `replaced` has no `replaced_reason`
field. The model can silently replace every step and the system doesn't flag
it. No constraint ties `propose_plan` to existing plan state.
### F5 — No plan on session resume: continuation sessions run ad-hoc (P1)
The ZimaOS continuation (18:29→19:01) had 81 calls with **zero**
`propose_plan` calls. The agent ran ad-hoc tool calls with no structure.
**Root cause:** When a session resumes, the `must_have_plan` guard is already
satisfied by the old (completed) plan. The agent doesn't re-plan on resume.
### F6 — Scope expansion / rabbit holes: agent chases irrelevant sub-goals (P2)
In the ZimaOS continuation, the agent spent ~40 calls trying to fix Technitium
DHCP reservations — a completely different subsystem from the goal ("make the
dashboard reachable"). The ZimaOS dashboard hadn't started since July 19
(3-week-old issue), making the DHCP reservation effort moot. The agent never
surfaced a question like "This is pre-existing — should I still fix DHCP?"
**Root cause:** No scope gate. When the agent pivots to a subsystem unrelated
to the stated goal, nothing stops it. The `session_questions` mechanism exists
(2 calls in 193 sessions) but the model never uses it.
### F7 — Command generation errors: malformed bash from the LLM (P2)
In the ZimaOS continuation, the agent generated:
- `head - n` instead of `head -n` → bash syntax error
- `echo "---" && \n curl ...` → literal `\n` in command → ambiguous redirect
**Root cause:** The LLM generates bash commands inline in content blocks. No
syntax validation, no escape-character handling. The `run` tool should reject
malformed commands before execution.
### F8 — Learning pipeline completely empty (P4)
| Table | Rows |
|-------|------|
| `classifications` | **0** |
| `feedback` | **0** |
| `patterns` | **0** |
| `skills` | **0** |
Despite 1,884 executions and 263 approvals, the system learns nothing from
outcomes. The ZimaOS session discovered that the dashboard hadn't started since
July 19 — this was never persisted. The HMAC trailing-newline discovery was
persisted manually; if the agent forgot `upsert_knowledge`, it would be lost.
### F9 — Observability gaps (P3)
| Gap | Detail |
|-----|--------|
| Token tracking | `agent_activity.token_count` is NULL for every row |
| Execution linkage | `nomos_plan_executions` is empty; `audit_log.session_id` is null |
| Plan quality | No metric for step completion rate (currently 38%) |
| MCP timeout rate | No counter for `run` calls that hit the client timeout |
### F10 — Execution success rate is 74% (P2)
1,884 executions: 1,400 completed (74%), 295 failed (15.6%), 152 cancelled
(8%), 37 denied (2%). One in four execution attempts doesn't complete.
---
## 3. Improvement plan (ordered by impact/effort)
### Task 1 — Prevent premature `complete_task(success)` *(fixes F2)*
**`internal/mcp/tools.go`** — In the `complete_task` handler, when
`outcome=success`: require the summary field to contain a verifiable state
assertion. Minimum: if the goal mentions a URL, check that the summary doesn't
contradict known state. Lightweight: log a warning if the summary says "returns
200" but the last `ping_service` or `run` result says otherwise.
**Coach:** `nomos/SOUL.md` — explicit rule: *"Before calling
complete_task(success), restate the user's original goal in your own words and
verify each condition. If any condition is 'probably works' rather than
'verified,' ask the operator or set outcome=partial."*
### Task 2 — Target validation in `run` *(fixes F3)*
**`internal/mcp/server.go`** — In the `run` handler, before dispatching:
validate that the command prefix matches the target type.
```
pct/qm/pvesh/iptables → only host:* targets
systemctl/docker → host:* or lxc:* targets
curl/nmap/ss → any target
```
If mismatched, return a clear error: *"Cannot run `qm` on lxc:dns — `qm` is a
Proxmox host command. Use target host:hubris or host:strong."* Do not classify
or execute.
**Test:** `TestClassifyCommand_WrongTarget` → commands with host-only prefixes
on LXC targets return error without execution.
### Task 3 — Raise MCP client timeout; add async path for long-running commands *(fixes F1)*
**`cmd/nomos`** — Raise the MCP client timeout from 30s to 120s.
**`internal/mcp/server.go`** — For `run` commands that the classifier
determines will exceed the client timeout (presence of `sleep`, `wait`,
`timeout` in the command), return immediately with an `execution_id` and status
`running`. The agent already has `get_execution_status` — use it:
1. Classify the command; if it contains `sleep`, `wait`, or shell constructs
that imply polling, flag it as `async_potential`.
2. Start the command, return the `execution_id` immediately.
3. Agent polls with `get_execution_status(execution_id)`.
4. If the client timeout is hit mid-poll, the execution continues on the server
— it's not lost.
**Test:** `run` with `sleep 60; echo done` on host:hubris → returns
immediately (not 30s timeout), `get_execution_status` eventually returns
`completed`.
### Task 4 — Plan step integrity: require `replaced_reason` on replacement *(fixes F4)*
**`session_plan_steps` migration** — Add `replaced_reason TEXT` column.
**`cmd/nomos`** — When the agent emits `update_plan_step` with
`status=replaced`, require a non-empty `replaced_reason`. Valid reasons:
`wrong_diagnosis`, `scope_change`, `blocked`, `superseded`, `operator_override`.
**Coach:** `nomos/SOUL.md` — explicit rule: *"Complete (status=done) or
explicitly skip (status=skipped) steps. Use status=replaced only when the
entire plan generation is wrong; include the reason. Replacing all steps with
no reason is a session-quality violation."*
### Task 5 — Force `propose_plan` on session resume *(fixes F5)*
**`cmd/nomos`** — When a session with status `done` or `failed` receives a new
user message, reset the plan state: clear step status, require a new
`propose_plan` call before any `run` calls. The "must have plan" guard should
consider the resumed session as plan-less until a fresh `propose_plan` is
called.
**Guard:** `set_goal` + `propose_plan` must be called before any `run` in a
resumed session. Reuse the existing "No plan — call set_goal then propose_plan"
error from `internal/mcp/server.go`.
### Task 6 — Scope gate: surface `session_questions` on context switch *(fixes F6)*
**Coach:** `nomos/SOUL.md` — explicit rule: *"Before pivoting to a subsystem
not mentioned in the user's goal, ask via session_questions. Example: 'The
dashboard logs show it hasn't started since July 19. Do you want me to debug
the dashboard service itself [A], skip it and just stabilize the IP [B], or
stop here [C]?'"*
**`cmd/nomos` prompt** — Add to the system prompt: *"When the investigation
leads to a subsystem or root cause unrelated to the expressed goal, surface a
session_question before taking action."*
### Task 7 — Auto-upsert knowledge on session close *(fixes F8)*
**`cmd/nomos`** — On `complete_task` (any outcome: success, partial, failure),
auto-generate a knowledge entry:
```
title: "<date>: <goal summary>"
content: "## Outcome\n<outcome>\n## Root cause\n<extracted>\n## What was done\n<summary>\n## What was left\n<unresolved>"
tags: [session:<id>]
about: [entities involved]
```
This ensures every session leaves a trace regardless of whether the agent
remembered to call `upsert_knowledge`.
### Task 8 — Token tracking *(fixes F9)*
**`cmd/nomos`** — After each LLM call, extract `usage.prompt_tokens`,
`usage.completion_tokens`, `usage.total_tokens` from the response and write to
`agent_activity.token_count`. Currently the field exists but is never populated
(NULL for all rows).
### Task 9 — Execution linkage *(fixes F9)*
**`internal/mcp/server.go`** — When `run` creates an execution, write a row
into `nomos_plan_executions` linking `session_id`, `plan_step_seq`, and
`execution_id`.
**`internal/mcp/server.go`** — Pass `session_id` (from MCP request headers)
into `audit_log` writes. Currently `audit_log.session_id` is NULL — the
`createAuditLog` function in `internal/httpapi/impl.go` receives the
correlation_id but not the session_id from the MCP path.
### Task 10 — Plan quality metric *(fixes F9)*
**`cmd/nomos`** — At session close, compute: `completed_steps /
total_steps_per_plan` (currently ~38%). Log as a metric or write as a session
attribute. Track over time to measure plan-adherence improvements from Tasks
4+5.
### Task 11 — Auto-classify every `run` call *(fixes F8)*
**`internal/mcp/server.go`** — The `run` handler already calls the classifier
(`classifyCommand` in `internal/policy/command.go`) to determine risk_class and
approval route. Write the result to the `classifications` table. Currently the
table is empty (0 rows) despite 1,884 executions being classified.
### Task 12 — Auto-feedback on session close *(fixes F8)*
**`cmd/nomos`** — On `complete_task`, generate a `feedback` entry:
```
session_id: <id>
outcome: <outcome>
observation: <summary>
lesson: <extracted from complete_task.summary>
side_effects: <entities created/modified during session>
```
**`cmd/oikos`** — Add a daily cron or scheduler job that reads recent
`feedback` entries and extracts `patterns` (recurring root causes, same-fix
applied multiple times, known-broken services). Seed the pattern table.
### Task 13 — Command syntax validation in `run` *(fixes F7)*
**`internal/mcp/server.go`** — Before executing a `run` command, do
lightweight bash syntax validation:
```
- Reject literal \n in commands (should be ; or &&)
- Reject commands where the last line ends with \ (backslash-continuation)
but no next line
- Warn on common typos: "head - n", "grep - i", spaces before flags
- Reject `&& \n` patterns (the LLM sometimes inserts literal \n between && chains)
```
### Task 14 — Stuck-session reaping (from prior plan; re-confirmed)
This session exhibited the same idle zombie pattern (23da10db — no closed_at,
status `failed` but outcome `failure`). Task 5 from the 2026-08-03 plan is
still open. Copying here for completeness.
---
## 4. Recommended sequence
```
P0 (blocks operational waste):
1 → 2 → 3
P1 (fixes plan architecture):
4 → 5
P2 (cognitive guardrails):
6 → 7 → 13
P3 (observability):
8 → 9 → 10
P4 (learning loop):
11 → 12
```
Sequence rationale: Tasks 1-3 stop the worst outcomes (premature success,
wrong-target execution, MCP timeouts). Tasks 4-5 make the plan system actually
useful instead of decorative. Tasks 6-7 add guardrails that prevent the ZimaOS
rabbit-hole class of failure. Tasks 8-10 give us visibility into whether any of
the previous tasks are working. Tasks 11-12 close the learning loop.
---
## 5. Validation
| Task | Test |
|------|------|
| 1 | Session with goal "make X reachable" where last ping shows 502 → `complete_task(success)` is rejected or warns |
| 2 | `run("lxc:dns", "qm stop 100")` → error: "qm is a Proxmox host command" |
| 3 | `run` with `sleep 45; echo done` → returns execution_id immediately, `get_execution_status` shows final result |
| 4 | `update_plan_step(status=replaced)` with no reason → rejected; with reason → accepted |
| 5 | Resumed session calls `run` before `propose_plan` → blocked: "No plan — call propose_plan" |
| 6 | Agent pivots to unrelated subsystem → `session_questions` is called before action |
| 7 | `complete_task` → knowledge entry created automatically with session link |
| 8 | `agent_activity.token_count` is non-NULL after any LLM call |
| 9 | `nomos_plan_executions` has rows linking session + step + execution |
| 10 | Session close writes `plan_adherence` attribute (step-completion %) |
| 11 | `classifications` table has 1 row per `run` call with risk_class + route |
| 12 | `complete_task` → auto `feedback` entry; daily pattern job finds recurring issues |
| 13 | `run` with `head - n /etc/hosts` → rejected with clear error about malformed command |
---
## 6. Out of scope / open questions
- Whether to raise `auto_act` from `off` for `reversible_low` actions (separate
policy decision; would reduce approval pileup without code changes).
- Whether to add a `delete_entity` MCP tool for lifecycle management (separate
from this reliability plan).
- The exact TTL for stuck-session reaping (30 min recommended, confirmed in
2026-08-03 plan).
- Whether `run` async mode should be opt-in (command contains sleep/wait) or
universal (every run returns immediately, agent always polls). Recommend
opt-in for now — most commands complete in <5s.