docs: task completion safety net — every live task stuck Running

Traced during UI-review verification: 50/50 live sessions are stuck
active/planning, never done/failed. Root cause confirmed against the
running DB — set_goal called once, propose_plan and complete_task
called zero times across all 50 sessions. The model consistently
skips the terminal complete_task call despite SOUL.md explicitly
instructing it to, especially for trivial single-tool Q&A turns.
Plan proposes an inline safety net for the common case plus an idle
sweep for structured goal/plan sessions that stall.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-11 22:01:35 +02:00
parent fb4c76ba82
commit e3850f6820
2 changed files with 205 additions and 0 deletions

View File

@@ -0,0 +1,204 @@
# Task completion safety net: every live task is stuck "Running"
Status: Planned — 2026-07-11.
## Scope
Fix the root cause of a production-wide defect found while UI-testing
[`2026-07-11-ui-review-ia-usability.md`](done/2026-07-11-ui-review-ia-usability.md):
every session on the live task board shows as "Running" forever. Traced
through `cmd/nomos/` and confirmed against the running database — this is
not a frontend bug (the board correctly reflects real `agent_sessions.status`
values). It's an agent-behavior gap: the model almost never calls the
lifecycle tools (`set_goal` / `propose_plan` / `complete_task`) that the
task-board feature (shipped today,
[`done/2026-07-11-goal-oriented-chat-control-panel.md`](done/2026-07-11-goal-oriented-chat-control-panel.md))
depends on to know a task is finished.
## Evidence
Queried the live nomos API directly (`curl localhost:8092/sessions` and
per-session transcripts) against the running mac-mini stack:
- **50/50 live sessions**: 49 `active`, 1 `planning`. Zero have ever reached
`executing`, `awaiting_input`, `done`, or `failed`.
- Across all 50 sessions: **`set_goal` called once. `propose_plan` called
zero times. `complete_task` called zero times.**
- The dominant pattern (43/50 sessions, 2-message transcripts) is a single
quick exchange: operator asks something narrow ("what's the hostname of
lxc:caddy?"), the model runs one read tool (`run hostname`), answers in
plain text, and the turn ends — no lifecycle tool call at all. This is
exactly the case
[`nomos/SOUL.md:105-109`](../nomos/SOUL.md#L105) calls out by name
("a trivial read-only task... is a degenerate case... answer it and
`complete_task` with a one-line summary") — the instruction exists and is
explicit, and the model skips it anyway, consistently.
- The one session that *did* call `set_goal` (a fleet health check) did
substantial real research (`get_health_summary`, `get_state_snapshot`,
`get_signal_history`, `list_lxcs`), gave the operator a full structured
answer, and then also just stopped — no `propose_plan`, no
`complete_task`. Status: stuck at `planning` since 2026-07-11T11:35, still
showing "Running" on the board.
This means the board's "N Running / 0 Done / 0 Failed" isn't a fluke or an
edge case — it's the default outcome for essentially every task the system
has ever run. The feature as designed (terminal state is 100% dependent on
the model remembering to call one specific tool) doesn't hold up against
real model behavior, even with an explicit prompt instruction already in
place.
## Where this lives in the code
`cmd/nomos/agent.go`'s `chatWith` has exactly one place a turn ends with a
plain-text answer and no tool calls:
```go
// agent.go:359-368
if len(msg.ToolCalls) == 0 {
emit(agentEvent{Type: "text", Data: msg.Content, SessionID: sessionID})
emit(agentEvent{Type: "done", ...})
return
}
```
This is reached for the trivial-Q&A case (a turn that made zero or a few
read-only tool calls this iteration, then answered in text) and is where
43/50 of the stuck sessions are produced. There's a second, rarer exit at
the step-limit fallback (`agent.go:487-494`, `finalSummary`) with the same
gap.
Neither exit currently checks whether the session ever reached a terminal
state — the turn just ends, and `agent_sessions.status` is left wherever it
was (usually `active`, its creation-time default,
[`store.go:71,84`](../cmd/nomos/store.go#L71)).
## Design
Two different failure shapes need two different fixes — collapsing them
into one heuristic would either auto-close genuinely in-progress structured
tasks or fail to catch the trivial-Q&A majority.
**1. Trivial/no-lifecycle-tool sessions (the 43/50 case) — auto-complete
inline, same turn.**
If a turn ends with a plain-text response (`len(msg.ToolCalls) == 0`, the
existing exit at `agent.go:359`) AND this session has never called
`set_goal` in its history, that's strong evidence this was never meant to
be a structured multi-step task — it's a one-shot question that got
answered. Call `store.completeTask` server-side right there, before the
`return`, with `outcome="success"` and a summary derived from the response
text (first ~120 chars, same truncation pattern
`buildContinuationNote` already uses at
[`continue.go:203-205`](../cmd/nomos/continue.go#L203)). No LLM call needed
— this is a mechanical default, not a judgment call, matching the "trivial
task" case SOUL.md already describes.
If the session *has* called `set_goal` (meaning the model explicitly framed
this as a task, e.g. the fleet-health-check session), auto-completing on
the very next plain-text turn is riskier — the model may reasonably expect
to be asked something next. Skip the inline auto-complete for these; case 2
covers them.
**2. Structured (goal/plan set) sessions that stall — idle sweep, not
inline.**
Extend the existing `runContinuationWorker` ticker
([`continue.go:41-57`](../cmd/nomos/continue.go#L41), already polling every
4s for a different purpose) with a second, coarser sweep — e.g. every 5
minutes — that finds sessions where:
- `status` is `active`, `planning`, or `executing` (not already terminal or
`awaiting_input`, which has its own resolution path), AND
- `set_goal` was called (this is a real task, not case 1), AND
- `last_active_at` is older than some idle threshold (start with 15
minutes — long enough that it's not still mid-turn, short enough that the
board doesn't lie for hours).
First idle hit: inject a system note next time nothing else touches the
session ("[System: this task has been idle for N minutes with no
`complete_task` call. If the goal is done, call it now with a summary. If
you're genuinely still working, ignore this.]") the same way
`buildContinuationNote` already injects notes into resumed sessions — reuse
`resumeSession`'s live-persist pattern
([`continue.go:106-196`](../cmd/nomos/continue.go#L106)) so the nudge and
the model's response show up in the transcript, not silently.
If a second idle sweep finds the same session still not completed (i.e.
the nudge didn't take), auto-complete it directly with
`outcome="partial"` and a summary noting it was auto-closed after an
unanswered nudge — same reasoning as `resumeSession`'s existing
"give the task a real, operator-visible terminal state instead of leaving
it silently stuck forever" logic at
[`continue.go:179-193`](../cmd/nomos/continue.go#L179), which already does
exactly this for a different failure mode (a resume that produces no
response). This is the same architectural pattern, applied to a session
that produces responses but never a terminal tool call.
**3. Leave `ask_operator` and gated-execution flows alone.** Those already
have real terminal signals (`awaiting_input` status, the continuation
worker's assent-window logic) — this plan only targets sessions that fall
through with no lifecycle signal at all.
## Fix plan
1. **Inline safety net (case 1)** — in `chatWith`'s plain-text exit
(`agent.go:359`), check `set_goal` was never called for this session
(cheap: track a bool while replaying `history` in the same function, no
extra query — the loop at `agent.go:209-226` already walks every
persisted message and could flag `sawSetGoal` while extracting tool
calls). If not sawSetGoal, call `completeTask` before returning.
2. **Idle sweep (case 2)** — new ticker in `continue.go` (or extend the
existing one with a slower secondary tick), a new store query
(`store.staleGoalSessions(ctx, idleThreshold)` mirroring
`pendingContinuations`'s shape), and reuse of `resumeSession`'s
live-persist injection for the nudge.
3. **Second-strike auto-close (case 2, continued)** — track nudge count (a
new `agent_sessions` column, e.g. `completion_nudges int default 0`, or
reuse the existing `summary`/attributes json instead of a schema change
if that's preferable) so the sweep can tell "never nudged" from "nudged
once already, still stuck."
4. **Backfill** — the 50 already-stuck live sessions won't get fixed by new
code alone (they're historical). One-time cleanup: run the same
case-1/case-2 classification against existing rows once the code ships,
so the board doesn't show 50 permanently-orphaned "Running" cards on top
of new correctly-terminating ones. This should be a script, not a manual
UPDATE — the classification logic will already exist in Go.
## Implementation order
1. Fix 1 (inline safety net) first — it's the highest-leverage, lowest-risk
change (self-contained, no schema change, covers 43/50 of the evidence).
2. Fix 2+3 (idle sweep + second-strike) — needs the schema decision
(new column vs. attribute) settled first; smaller blast radius than 1
but touches the ticker/worker machinery, deserves its own review pass.
3. Fix 4 (backfill) last, once 1-3 are deployed and verified live — running
it before the code ships would just recreate the same gap for new
sessions created in between.
## Verification
- After fix 1: start a few trivial one-shot chats against the live agent
(`hostname`-style questions), confirm each session reaches `status=done`
immediately after the answer, via `curl localhost:8092/sessions/:id` or
the task board.
- After fix 2+3: manually let a goal-bearing session go idle past the
threshold (or lower the threshold for a local test run), confirm the
nudge appears in the transcript, then confirm second-strike auto-close
fires if the nudge is ignored.
- Re-run the same audit query used to find this bug
(`curl localhost:8092/sessions` → status histogram) a day after deploy;
the "stuck active/planning forever" count should track only genuinely
in-flight tasks, not accumulate.
## Open questions
- **Outcome for case-1 auto-complete**: always `"success"`, or worth a
cheap heuristic (e.g. scan the final text for obvious failure language)?
Recommend starting with always-`"success"` — SOUL.md's own trivial-task
guidance doesn't distinguish, and a wrong "success" on a genuinely-failed
one-shot lookup is low-stakes (the transcript still shows the real
answer; nothing acts on the outcome besides the board's color).
- **Idle threshold (15 min) and nudge-to-close gap**: arbitrary starting
points, not measured against real task durations — worth revisiting after
a week of the new sessions' real timing data exists.
- **Schema change for nudge tracking**: a new column is simpler to query
than packing state into existing JSON, but adds a migration — worth
confirming that's acceptable before starting fix 2+3 (this plan defers
that call to whoever implements it, per Implementation order above).

View File

@@ -14,6 +14,7 @@ went sideways, open an investigation.
| 2026-07-08 | [Liveness, drift, and UX cohesion](2026-07-08-liveness-drift-and-ux-cohesion.md) | In Progress — Phase 5 deferred | | 2026-07-08 | [Liveness, drift, and UX cohesion](2026-07-08-liveness-drift-and-ux-cohesion.md) | In Progress — Phase 5 deferred |
| 2026-07-10 | [General gated execution: unlimited actions, gated by risk](2026-07-10-general-gated-execution.md) | In Progress — enum retirement + auto-act revival still open | | 2026-07-10 | [General gated execution: unlimited actions, gated by risk](2026-07-10-general-gated-execution.md) | In Progress — enum retirement + auto-act revival still open |
| 2026-07-11 | [Nomos agent code review: gaps and improvement plan](2026-07-11-nomos-agent-code-review.md) | In Progress — only C1 (unauthenticated nomos gateway) still open, deferred | | 2026-07-11 | [Nomos agent code review: gaps and improvement plan](2026-07-11-nomos-agent-code-review.md) | In Progress — only C1 (unauthenticated nomos gateway) still open, deferred |
| 2026-07-11 | [Task completion safety net: every live task is stuck "Running"](2026-07-11-task-completion-safety-net.md) | Planned — not started |
## Done ## Done