diff --git a/plans/2026-07-11-task-completion-safety-net.md b/plans/2026-07-11-task-completion-safety-net.md new file mode 100644 index 0000000..2649e54 --- /dev/null +++ b/plans/2026-07-11-task-completion-safety-net.md @@ -0,0 +1,204 @@ +# Task completion safety net: every live task is stuck "Running" + +Status: Planned — 2026-07-11. + +## Scope + +Fix the root cause of a production-wide defect found while UI-testing +[`2026-07-11-ui-review-ia-usability.md`](done/2026-07-11-ui-review-ia-usability.md): +every session on the live task board shows as "Running" forever. Traced +through `cmd/nomos/` and confirmed against the running database — this is +not a frontend bug (the board correctly reflects real `agent_sessions.status` +values). It's an agent-behavior gap: the model almost never calls the +lifecycle tools (`set_goal` / `propose_plan` / `complete_task`) that the +task-board feature (shipped today, +[`done/2026-07-11-goal-oriented-chat-control-panel.md`](done/2026-07-11-goal-oriented-chat-control-panel.md)) +depends on to know a task is finished. + +## Evidence + +Queried the live nomos API directly (`curl localhost:8092/sessions` and +per-session transcripts) against the running mac-mini stack: + +- **50/50 live sessions**: 49 `active`, 1 `planning`. Zero have ever reached + `executing`, `awaiting_input`, `done`, or `failed`. +- Across all 50 sessions: **`set_goal` called once. `propose_plan` called + zero times. `complete_task` called zero times.** +- The dominant pattern (43/50 sessions, 2-message transcripts) is a single + quick exchange: operator asks something narrow ("what's the hostname of + lxc:caddy?"), the model runs one read tool (`run hostname`), answers in + plain text, and the turn ends — no lifecycle tool call at all. This is + exactly the case + [`nomos/SOUL.md:105-109`](../nomos/SOUL.md#L105) calls out by name + ("a trivial read-only task... is a degenerate case... answer it and + `complete_task` with a one-line summary") — the instruction exists and is + explicit, and the model skips it anyway, consistently. +- The one session that *did* call `set_goal` (a fleet health check) did + substantial real research (`get_health_summary`, `get_state_snapshot`, + `get_signal_history`, `list_lxcs`), gave the operator a full structured + answer, and then also just stopped — no `propose_plan`, no + `complete_task`. Status: stuck at `planning` since 2026-07-11T11:35, still + showing "Running" on the board. + +This means the board's "N Running / 0 Done / 0 Failed" isn't a fluke or an +edge case — it's the default outcome for essentially every task the system +has ever run. The feature as designed (terminal state is 100% dependent on +the model remembering to call one specific tool) doesn't hold up against +real model behavior, even with an explicit prompt instruction already in +place. + +## Where this lives in the code + +`cmd/nomos/agent.go`'s `chatWith` has exactly one place a turn ends with a +plain-text answer and no tool calls: + +```go +// agent.go:359-368 +if len(msg.ToolCalls) == 0 { + emit(agentEvent{Type: "text", Data: msg.Content, SessionID: sessionID}) + emit(agentEvent{Type: "done", ...}) + return +} +``` + +This is reached for the trivial-Q&A case (a turn that made zero or a few +read-only tool calls this iteration, then answered in text) and is where +43/50 of the stuck sessions are produced. There's a second, rarer exit at +the step-limit fallback (`agent.go:487-494`, `finalSummary`) with the same +gap. + +Neither exit currently checks whether the session ever reached a terminal +state — the turn just ends, and `agent_sessions.status` is left wherever it +was (usually `active`, its creation-time default, +[`store.go:71,84`](../cmd/nomos/store.go#L71)). + +## Design + +Two different failure shapes need two different fixes — collapsing them +into one heuristic would either auto-close genuinely in-progress structured +tasks or fail to catch the trivial-Q&A majority. + +**1. Trivial/no-lifecycle-tool sessions (the 43/50 case) — auto-complete + inline, same turn.** +If a turn ends with a plain-text response (`len(msg.ToolCalls) == 0`, the +existing exit at `agent.go:359`) AND this session has never called +`set_goal` in its history, that's strong evidence this was never meant to +be a structured multi-step task — it's a one-shot question that got +answered. Call `store.completeTask` server-side right there, before the +`return`, with `outcome="success"` and a summary derived from the response +text (first ~120 chars, same truncation pattern +`buildContinuationNote` already uses at +[`continue.go:203-205`](../cmd/nomos/continue.go#L203)). No LLM call needed +— this is a mechanical default, not a judgment call, matching the "trivial +task" case SOUL.md already describes. + +If the session *has* called `set_goal` (meaning the model explicitly framed +this as a task, e.g. the fleet-health-check session), auto-completing on +the very next plain-text turn is riskier — the model may reasonably expect +to be asked something next. Skip the inline auto-complete for these; case 2 +covers them. + +**2. Structured (goal/plan set) sessions that stall — idle sweep, not + inline.** +Extend the existing `runContinuationWorker` ticker +([`continue.go:41-57`](../cmd/nomos/continue.go#L41), already polling every +4s for a different purpose) with a second, coarser sweep — e.g. every 5 +minutes — that finds sessions where: +- `status` is `active`, `planning`, or `executing` (not already terminal or + `awaiting_input`, which has its own resolution path), AND +- `set_goal` was called (this is a real task, not case 1), AND +- `last_active_at` is older than some idle threshold (start with 15 + minutes — long enough that it's not still mid-turn, short enough that the + board doesn't lie for hours). + +First idle hit: inject a system note next time nothing else touches the +session ("[System: this task has been idle for N minutes with no +`complete_task` call. If the goal is done, call it now with a summary. If +you're genuinely still working, ignore this.]") the same way +`buildContinuationNote` already injects notes into resumed sessions — reuse +`resumeSession`'s live-persist pattern +([`continue.go:106-196`](../cmd/nomos/continue.go#L106)) so the nudge and +the model's response show up in the transcript, not silently. + +If a second idle sweep finds the same session still not completed (i.e. +the nudge didn't take), auto-complete it directly with +`outcome="partial"` and a summary noting it was auto-closed after an +unanswered nudge — same reasoning as `resumeSession`'s existing +"give the task a real, operator-visible terminal state instead of leaving +it silently stuck forever" logic at +[`continue.go:179-193`](../cmd/nomos/continue.go#L179), which already does +exactly this for a different failure mode (a resume that produces no +response). This is the same architectural pattern, applied to a session +that produces responses but never a terminal tool call. + +**3. Leave `ask_operator` and gated-execution flows alone.** Those already +have real terminal signals (`awaiting_input` status, the continuation +worker's assent-window logic) — this plan only targets sessions that fall +through with no lifecycle signal at all. + +## Fix plan + +1. **Inline safety net (case 1)** — in `chatWith`'s plain-text exit + (`agent.go:359`), check `set_goal` was never called for this session + (cheap: track a bool while replaying `history` in the same function, no + extra query — the loop at `agent.go:209-226` already walks every + persisted message and could flag `sawSetGoal` while extracting tool + calls). If not sawSetGoal, call `completeTask` before returning. +2. **Idle sweep (case 2)** — new ticker in `continue.go` (or extend the + existing one with a slower secondary tick), a new store query + (`store.staleGoalSessions(ctx, idleThreshold)` mirroring + `pendingContinuations`'s shape), and reuse of `resumeSession`'s + live-persist injection for the nudge. +3. **Second-strike auto-close (case 2, continued)** — track nudge count (a + new `agent_sessions` column, e.g. `completion_nudges int default 0`, or + reuse the existing `summary`/attributes json instead of a schema change + if that's preferable) so the sweep can tell "never nudged" from "nudged + once already, still stuck." +4. **Backfill** — the 50 already-stuck live sessions won't get fixed by new + code alone (they're historical). One-time cleanup: run the same + case-1/case-2 classification against existing rows once the code ships, + so the board doesn't show 50 permanently-orphaned "Running" cards on top + of new correctly-terminating ones. This should be a script, not a manual + UPDATE — the classification logic will already exist in Go. + +## Implementation order + +1. Fix 1 (inline safety net) first — it's the highest-leverage, lowest-risk + change (self-contained, no schema change, covers 43/50 of the evidence). +2. Fix 2+3 (idle sweep + second-strike) — needs the schema decision + (new column vs. attribute) settled first; smaller blast radius than 1 + but touches the ticker/worker machinery, deserves its own review pass. +3. Fix 4 (backfill) last, once 1-3 are deployed and verified live — running + it before the code ships would just recreate the same gap for new + sessions created in between. + +## Verification + +- After fix 1: start a few trivial one-shot chats against the live agent + (`hostname`-style questions), confirm each session reaches `status=done` + immediately after the answer, via `curl localhost:8092/sessions/:id` or + the task board. +- After fix 2+3: manually let a goal-bearing session go idle past the + threshold (or lower the threshold for a local test run), confirm the + nudge appears in the transcript, then confirm second-strike auto-close + fires if the nudge is ignored. +- Re-run the same audit query used to find this bug + (`curl localhost:8092/sessions` → status histogram) a day after deploy; + the "stuck active/planning forever" count should track only genuinely + in-flight tasks, not accumulate. + +## Open questions + +- **Outcome for case-1 auto-complete**: always `"success"`, or worth a + cheap heuristic (e.g. scan the final text for obvious failure language)? + Recommend starting with always-`"success"` — SOUL.md's own trivial-task + guidance doesn't distinguish, and a wrong "success" on a genuinely-failed + one-shot lookup is low-stakes (the transcript still shows the real + answer; nothing acts on the outcome besides the board's color). +- **Idle threshold (15 min) and nudge-to-close gap**: arbitrary starting + points, not measured against real task durations — worth revisiting after + a week of the new sessions' real timing data exists. +- **Schema change for nudge tracking**: a new column is simpler to query + than packing state into existing JSON, but adds a migration — worth + confirming that's acceptable before starting fix 2+3 (this plan defers + that call to whoever implements it, per Implementation order above). diff --git a/plans/index.md b/plans/index.md index 22fbe99..a9a36ea 100644 --- a/plans/index.md +++ b/plans/index.md @@ -14,6 +14,7 @@ went sideways, open an investigation. | 2026-07-08 | [Liveness, drift, and UX cohesion](2026-07-08-liveness-drift-and-ux-cohesion.md) | In Progress — Phase 5 deferred | | 2026-07-10 | [General gated execution: unlimited actions, gated by risk](2026-07-10-general-gated-execution.md) | In Progress — enum retirement + auto-act revival still open | | 2026-07-11 | [Nomos agent code review: gaps and improvement plan](2026-07-11-nomos-agent-code-review.md) | In Progress — only C1 (unauthenticated nomos gateway) still open, deferred | +| 2026-07-11 | [Task completion safety net: every live task is stuck "Running"](2026-07-11-task-completion-safety-net.md) | Planned — not started | ## Done