11 KiB
Task completion safety net: every live task is stuck "Running"
Status: In Progress — 2026-07-11. Fixes 1-3 implemented, built, tested
(go build ./..., go test ./cmd/nomos/...), and committed
(3b9c75f). Not yet deployed or verified live. Fix 4 (backfill of the 50
already-stuck live sessions) intentionally not started — per the
implementation order below, it needs 1-3 deployed and verified first.
Scope
Fix the root cause of a production-wide defect found while UI-testing
2026-07-11-ui-review-ia-usability.md:
every session on the live task board shows as "Running" forever. Traced
through cmd/nomos/ and confirmed against the running database — this is
not a frontend bug (the board correctly reflects real agent_sessions.status
values). It's an agent-behavior gap: the model almost never calls the
lifecycle tools (set_goal / propose_plan / complete_task) that the
task-board feature (shipped today,
done/2026-07-11-goal-oriented-chat-control-panel.md)
depends on to know a task is finished.
Evidence
Queried the live nomos API directly (curl localhost:8092/sessions and
per-session transcripts) against the running mac-mini stack:
- 50/50 live sessions: 49
active, 1planning. Zero have ever reachedexecuting,awaiting_input,done, orfailed. - Across all 50 sessions:
set_goalcalled once.propose_plancalled zero times.complete_taskcalled zero times. - The dominant pattern (43/50 sessions, 2-message transcripts) is a single
quick exchange: operator asks something narrow ("what's the hostname of
lxc:caddy?"), the model runs one read tool (
run hostname), answers in plain text, and the turn ends — no lifecycle tool call at all. This is exactly the casenomos/SOUL.md:105-109calls out by name ("a trivial read-only task... is a degenerate case... answer it andcomplete_taskwith a one-line summary") — the instruction exists and is explicit, and the model skips it anyway, consistently. - The one session that did call
set_goal(a fleet health check) did substantial real research (get_health_summary,get_state_snapshot,get_signal_history,list_lxcs), gave the operator a full structured answer, and then also just stopped — nopropose_plan, nocomplete_task. Status: stuck atplanningsince 2026-07-11T11:35, still showing "Running" on the board.
This means the board's "N Running / 0 Done / 0 Failed" isn't a fluke or an edge case — it's the default outcome for essentially every task the system has ever run. The feature as designed (terminal state is 100% dependent on the model remembering to call one specific tool) doesn't hold up against real model behavior, even with an explicit prompt instruction already in place.
Where this lives in the code
cmd/nomos/agent.go's chatWith has exactly one place a turn ends with a
plain-text answer and no tool calls:
// agent.go:359-368
if len(msg.ToolCalls) == 0 {
emit(agentEvent{Type: "text", Data: msg.Content, SessionID: sessionID})
emit(agentEvent{Type: "done", ...})
return
}
This is reached for the trivial-Q&A case (a turn that made zero or a few
read-only tool calls this iteration, then answered in text) and is where
43/50 of the stuck sessions are produced. There's a second, rarer exit at
the step-limit fallback (agent.go:487-494, finalSummary) with the same
gap.
Neither exit currently checks whether the session ever reached a terminal
state — the turn just ends, and agent_sessions.status is left wherever it
was (usually active, its creation-time default,
store.go:71,84).
Design
Two different failure shapes need two different fixes — collapsing them into one heuristic would either auto-close genuinely in-progress structured tasks or fail to catch the trivial-Q&A majority.
1. Trivial/no-lifecycle-tool sessions (the 43/50 case) — auto-complete
inline, same turn.
If a turn ends with a plain-text response (len(msg.ToolCalls) == 0, the
existing exit at agent.go:359) AND this session has never called
set_goal in its history, that's strong evidence this was never meant to
be a structured multi-step task — it's a one-shot question that got
answered. Call store.completeTask server-side right there, before the
return, with outcome="success" and a summary derived from the response
text (first ~120 chars, same truncation pattern
buildContinuationNote already uses at
continue.go:203-205). No LLM call needed
— this is a mechanical default, not a judgment call, matching the "trivial
task" case SOUL.md already describes.
If the session has called set_goal (meaning the model explicitly framed
this as a task, e.g. the fleet-health-check session), auto-completing on
the very next plain-text turn is riskier — the model may reasonably expect
to be asked something next. Skip the inline auto-complete for these; case 2
covers them.
2. Structured (goal/plan set) sessions that stall — idle sweep, not
inline.
Extend the existing runContinuationWorker ticker
(continue.go:41-57, already polling every
4s for a different purpose) with a second, coarser sweep — e.g. every 5
minutes — that finds sessions where:
statusisactive,planning, orexecuting(not already terminal orawaiting_input, which has its own resolution path), ANDset_goalwas called (this is a real task, not case 1), ANDlast_active_atis older than some idle threshold (start with 15 minutes — long enough that it's not still mid-turn, short enough that the board doesn't lie for hours).
First idle hit: inject a system note next time nothing else touches the
session ("[System: this task has been idle for N minutes with no
complete_task call. If the goal is done, call it now with a summary. If
you're genuinely still working, ignore this.]") the same way
buildContinuationNote already injects notes into resumed sessions — reuse
resumeSession's live-persist pattern
(continue.go:106-196) so the nudge and
the model's response show up in the transcript, not silently.
If a second idle sweep finds the same session still not completed (i.e.
the nudge didn't take), auto-complete it directly with
outcome="partial" and a summary noting it was auto-closed after an
unanswered nudge — same reasoning as resumeSession's existing
"give the task a real, operator-visible terminal state instead of leaving
it silently stuck forever" logic at
continue.go:179-193, which already does
exactly this for a different failure mode (a resume that produces no
response). This is the same architectural pattern, applied to a session
that produces responses but never a terminal tool call.
3. Leave ask_operator and gated-execution flows alone. Those already
have real terminal signals (awaiting_input status, the continuation
worker's assent-window logic) — this plan only targets sessions that fall
through with no lifecycle signal at all.
Fix plan
- Inline safety net (case 1) — in
chatWith's plain-text exit (agent.go:359), checkset_goalwas never called for this session (cheap: track a bool while replayinghistoryin the same function, no extra query — the loop atagent.go:209-226already walks every persisted message and could flagsawSetGoalwhile extracting tool calls). If not sawSetGoal, callcompleteTaskbefore returning. - Idle sweep (case 2) — new ticker in
continue.go(or extend the existing one with a slower secondary tick), a new store query (store.staleGoalSessions(ctx, idleThreshold)mirroringpendingContinuations's shape), and reuse ofresumeSession's live-persist injection for the nudge. - Second-strike auto-close (case 2, continued) — track nudge count (a
new
agent_sessionscolumn, e.g.completion_nudges int default 0, or reuse the existingsummary/attributes json instead of a schema change if that's preferable) so the sweep can tell "never nudged" from "nudged once already, still stuck." - Backfill — the 50 already-stuck live sessions won't get fixed by new code alone (they're historical). One-time cleanup: run the same case-1/case-2 classification against existing rows once the code ships, so the board doesn't show 50 permanently-orphaned "Running" cards on top of new correctly-terminating ones. This should be a script, not a manual UPDATE — the classification logic will already exist in Go.
Implementation order
- Fix 1 (inline safety net) first — it's the highest-leverage, lowest-risk change (self-contained, no schema change, covers 43/50 of the evidence).
- Fix 2+3 (idle sweep + second-strike) — needs the schema decision (new column vs. attribute) settled first; smaller blast radius than 1 but touches the ticker/worker machinery, deserves its own review pass.
- Fix 4 (backfill) last, once 1-3 are deployed and verified live — running it before the code ships would just recreate the same gap for new sessions created in between.
Verification
- After fix 1: start a few trivial one-shot chats against the live agent
(
hostname-style questions), confirm each session reachesstatus=doneimmediately after the answer, viacurl localhost:8092/sessions/:idor the task board. - After fix 2+3: manually let a goal-bearing session go idle past the threshold (or lower the threshold for a local test run), confirm the nudge appears in the transcript, then confirm second-strike auto-close fires if the nudge is ignored.
- Re-run the same audit query used to find this bug
(
curl localhost:8092/sessions→ status histogram) a day after deploy; the "stuck active/planning forever" count should track only genuinely in-flight tasks, not accumulate.
Open questions
- Outcome for case-1 auto-complete: always
"success", or worth a cheap heuristic (e.g. scan the final text for obvious failure language)? Recommend starting with always-"success"— SOUL.md's own trivial-task guidance doesn't distinguish, and a wrong "success" on a genuinely-failed one-shot lookup is low-stakes (the transcript still shows the real answer; nothing acts on the outcome besides the board's color). - Idle threshold (15 min) and nudge-to-close gap: arbitrary starting points, not measured against real task durations — worth revisiting after a week of the new sessions' real timing data exists.
- Schema change for nudge tracking: a new column is simpler to query than packing state into existing JSON, but adds a migration — worth confirming that's acceptable before starting fix 2+3 (this plan defers that call to whoever implements it, per Implementation order above).