# Task completion safety net: every live task is stuck "Running" Status: In Progress — 2026-07-11. Fixes 1-3 implemented, built, tested (`go build ./...`, `go test ./cmd/nomos/...`), and committed (`3b9c75f`). Not yet deployed or verified live. Fix 4 (backfill of the 50 already-stuck live sessions) intentionally not started — per the implementation order below, it needs 1-3 deployed and verified first. ## Scope Fix the root cause of a production-wide defect found while UI-testing [`2026-07-11-ui-review-ia-usability.md`](done/2026-07-11-ui-review-ia-usability.md): every session on the live task board shows as "Running" forever. Traced through `cmd/nomos/` and confirmed against the running database — this is not a frontend bug (the board correctly reflects real `agent_sessions.status` values). It's an agent-behavior gap: the model almost never calls the lifecycle tools (`set_goal` / `propose_plan` / `complete_task`) that the task-board feature (shipped today, [`done/2026-07-11-goal-oriented-chat-control-panel.md`](done/2026-07-11-goal-oriented-chat-control-panel.md)) depends on to know a task is finished. ## Evidence Queried the live nomos API directly (`curl localhost:8092/sessions` and per-session transcripts) against the running mac-mini stack: - **50/50 live sessions**: 49 `active`, 1 `planning`. Zero have ever reached `executing`, `awaiting_input`, `done`, or `failed`. - Across all 50 sessions: **`set_goal` called once. `propose_plan` called zero times. `complete_task` called zero times.** - The dominant pattern (43/50 sessions, 2-message transcripts) is a single quick exchange: operator asks something narrow ("what's the hostname of lxc:caddy?"), the model runs one read tool (`run hostname`), answers in plain text, and the turn ends — no lifecycle tool call at all. This is exactly the case [`nomos/SOUL.md:105-109`](../nomos/SOUL.md#L105) calls out by name ("a trivial read-only task... is a degenerate case... answer it and `complete_task` with a one-line summary") — the instruction exists and is explicit, and the model skips it anyway, consistently. - The one session that *did* call `set_goal` (a fleet health check) did substantial real research (`get_health_summary`, `get_state_snapshot`, `get_signal_history`, `list_lxcs`), gave the operator a full structured answer, and then also just stopped — no `propose_plan`, no `complete_task`. Status: stuck at `planning` since 2026-07-11T11:35, still showing "Running" on the board. This means the board's "N Running / 0 Done / 0 Failed" isn't a fluke or an edge case — it's the default outcome for essentially every task the system has ever run. The feature as designed (terminal state is 100% dependent on the model remembering to call one specific tool) doesn't hold up against real model behavior, even with an explicit prompt instruction already in place. ## Where this lives in the code `cmd/nomos/agent.go`'s `chatWith` has exactly one place a turn ends with a plain-text answer and no tool calls: ```go // agent.go:359-368 if len(msg.ToolCalls) == 0 { emit(agentEvent{Type: "text", Data: msg.Content, SessionID: sessionID}) emit(agentEvent{Type: "done", ...}) return } ``` This is reached for the trivial-Q&A case (a turn that made zero or a few read-only tool calls this iteration, then answered in text) and is where 43/50 of the stuck sessions are produced. There's a second, rarer exit at the step-limit fallback (`agent.go:487-494`, `finalSummary`) with the same gap. Neither exit currently checks whether the session ever reached a terminal state — the turn just ends, and `agent_sessions.status` is left wherever it was (usually `active`, its creation-time default, [`store.go:71,84`](../cmd/nomos/store.go#L71)). ## Design Two different failure shapes need two different fixes — collapsing them into one heuristic would either auto-close genuinely in-progress structured tasks or fail to catch the trivial-Q&A majority. **1. Trivial/no-lifecycle-tool sessions (the 43/50 case) — auto-complete inline, same turn.** If a turn ends with a plain-text response (`len(msg.ToolCalls) == 0`, the existing exit at `agent.go:359`) AND this session has never called `set_goal` in its history, that's strong evidence this was never meant to be a structured multi-step task — it's a one-shot question that got answered. Call `store.completeTask` server-side right there, before the `return`, with `outcome="success"` and a summary derived from the response text (first ~120 chars, same truncation pattern `buildContinuationNote` already uses at [`continue.go:203-205`](../cmd/nomos/continue.go#L203)). No LLM call needed — this is a mechanical default, not a judgment call, matching the "trivial task" case SOUL.md already describes. If the session *has* called `set_goal` (meaning the model explicitly framed this as a task, e.g. the fleet-health-check session), auto-completing on the very next plain-text turn is riskier — the model may reasonably expect to be asked something next. Skip the inline auto-complete for these; case 2 covers them. **2. Structured (goal/plan set) sessions that stall — idle sweep, not inline.** Extend the existing `runContinuationWorker` ticker ([`continue.go:41-57`](../cmd/nomos/continue.go#L41), already polling every 4s for a different purpose) with a second, coarser sweep — e.g. every 5 minutes — that finds sessions where: - `status` is `active`, `planning`, or `executing` (not already terminal or `awaiting_input`, which has its own resolution path), AND - `set_goal` was called (this is a real task, not case 1), AND - `last_active_at` is older than some idle threshold (start with 15 minutes — long enough that it's not still mid-turn, short enough that the board doesn't lie for hours). First idle hit: inject a system note next time nothing else touches the session ("[System: this task has been idle for N minutes with no `complete_task` call. If the goal is done, call it now with a summary. If you're genuinely still working, ignore this.]") the same way `buildContinuationNote` already injects notes into resumed sessions — reuse `resumeSession`'s live-persist pattern ([`continue.go:106-196`](../cmd/nomos/continue.go#L106)) so the nudge and the model's response show up in the transcript, not silently. If a second idle sweep finds the same session still not completed (i.e. the nudge didn't take), auto-complete it directly with `outcome="partial"` and a summary noting it was auto-closed after an unanswered nudge — same reasoning as `resumeSession`'s existing "give the task a real, operator-visible terminal state instead of leaving it silently stuck forever" logic at [`continue.go:179-193`](../cmd/nomos/continue.go#L179), which already does exactly this for a different failure mode (a resume that produces no response). This is the same architectural pattern, applied to a session that produces responses but never a terminal tool call. **3. Leave `ask_operator` and gated-execution flows alone.** Those already have real terminal signals (`awaiting_input` status, the continuation worker's assent-window logic) — this plan only targets sessions that fall through with no lifecycle signal at all. ## Fix plan 1. **Inline safety net (case 1)** — in `chatWith`'s plain-text exit (`agent.go:359`), check `set_goal` was never called for this session (cheap: track a bool while replaying `history` in the same function, no extra query — the loop at `agent.go:209-226` already walks every persisted message and could flag `sawSetGoal` while extracting tool calls). If not sawSetGoal, call `completeTask` before returning. 2. **Idle sweep (case 2)** — new ticker in `continue.go` (or extend the existing one with a slower secondary tick), a new store query (`store.staleGoalSessions(ctx, idleThreshold)` mirroring `pendingContinuations`'s shape), and reuse of `resumeSession`'s live-persist injection for the nudge. 3. **Second-strike auto-close (case 2, continued)** — track nudge count (a new `agent_sessions` column, e.g. `completion_nudges int default 0`, or reuse the existing `summary`/attributes json instead of a schema change if that's preferable) so the sweep can tell "never nudged" from "nudged once already, still stuck." 4. **Backfill** — the 50 already-stuck live sessions won't get fixed by new code alone (they're historical). One-time cleanup: run the same case-1/case-2 classification against existing rows once the code ships, so the board doesn't show 50 permanently-orphaned "Running" cards on top of new correctly-terminating ones. This should be a script, not a manual UPDATE — the classification logic will already exist in Go. ## Implementation order 1. Fix 1 (inline safety net) first — it's the highest-leverage, lowest-risk change (self-contained, no schema change, covers 43/50 of the evidence). 2. Fix 2+3 (idle sweep + second-strike) — needs the schema decision (new column vs. attribute) settled first; smaller blast radius than 1 but touches the ticker/worker machinery, deserves its own review pass. 3. Fix 4 (backfill) last, once 1-3 are deployed and verified live — running it before the code ships would just recreate the same gap for new sessions created in between. ## Verification - After fix 1: start a few trivial one-shot chats against the live agent (`hostname`-style questions), confirm each session reaches `status=done` immediately after the answer, via `curl localhost:8092/sessions/:id` or the task board. - After fix 2+3: manually let a goal-bearing session go idle past the threshold (or lower the threshold for a local test run), confirm the nudge appears in the transcript, then confirm second-strike auto-close fires if the nudge is ignored. - Re-run the same audit query used to find this bug (`curl localhost:8092/sessions` → status histogram) a day after deploy; the "stuck active/planning forever" count should track only genuinely in-flight tasks, not accumulate. ## Open questions - **Outcome for case-1 auto-complete**: always `"success"`, or worth a cheap heuristic (e.g. scan the final text for obvious failure language)? Recommend starting with always-`"success"` — SOUL.md's own trivial-task guidance doesn't distinguish, and a wrong "success" on a genuinely-failed one-shot lookup is low-stakes (the transcript still shows the real answer; nothing acts on the outcome besides the board's color). - **Idle threshold (15 min) and nudge-to-close gap**: arbitrary starting points, not measured against real task durations — worth revisiting after a week of the new sessions' real timing data exists. - **Schema change for nudge tracking**: a new column is simpler to query than packing state into existing JSON, but adds a migration — worth confirming that's acceptable before starting fix 2+3 (this plan defers that call to whoever implements it, per Implementation order above).