Fixes 1-3 deployed and verified live: fresh trivial Q&A sessions now reach done immediately, and a goal-bearing session that stalled was correctly nudged by the idle sweep. Fix 4 (backfill) was replaced with deletion after the operator's call — verified against the DB first that zero knowledge notes were linked to or written by any of the 53 removed sessions, so nothing was lost. Documents the pagination gap in listSessions (hardcoded LIMIT 50, no total count) that hid 6 of those sessions from the original audit. Also fixes relative links in this plan and in the UI-review plan that broke when both moved from plans/ to plans/done/ (one directory level deeper). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
257 lines
14 KiB
Markdown
257 lines
14 KiB
Markdown
# Task completion safety net: every live task is stuck "Running"
|
|
|
|
Status: Done — 2026-07-12. Fixes 1-3 implemented, built, tested
|
|
(`go build ./...`, `go test ./cmd/nomos/...`), committed (`3b9c75f`),
|
|
deployed, and verified live (see Verification below — fresh trivial Q&A
|
|
sessions now reach `done` immediately; a goal-bearing session that went
|
|
idle was correctly nudged and auto-resolved by the existing resume-failure
|
|
path). Fix 4 (backfill) was replaced with deletion — see "Fix 4, revised"
|
|
below; the original backfill-with-a-fabricated-outcome approach was never
|
|
run.
|
|
|
|
## Scope
|
|
|
|
Fix the root cause of a production-wide defect found while UI-testing
|
|
[`2026-07-11-ui-review-ia-usability.md`](2026-07-11-ui-review-ia-usability.md):
|
|
every session on the live task board shows as "Running" forever. Traced
|
|
through `cmd/nomos/` and confirmed against the running database — this is
|
|
not a frontend bug (the board correctly reflects real `agent_sessions.status`
|
|
values). It's an agent-behavior gap: the model almost never calls the
|
|
lifecycle tools (`set_goal` / `propose_plan` / `complete_task`) that the
|
|
task-board feature (shipped today,
|
|
[`done/2026-07-11-goal-oriented-chat-control-panel.md`](2026-07-11-goal-oriented-chat-control-panel.md))
|
|
depends on to know a task is finished.
|
|
|
|
## Evidence
|
|
|
|
Queried the live nomos API directly (`curl localhost:8092/sessions` and
|
|
per-session transcripts) against the running mac-mini stack:
|
|
|
|
- **50/50 live sessions**: 49 `active`, 1 `planning`. Zero have ever reached
|
|
`executing`, `awaiting_input`, `done`, or `failed`.
|
|
- Across all 50 sessions: **`set_goal` called once. `propose_plan` called
|
|
zero times. `complete_task` called zero times.**
|
|
- The dominant pattern (43/50 sessions, 2-message transcripts) is a single
|
|
quick exchange: operator asks something narrow ("what's the hostname of
|
|
lxc:caddy?"), the model runs one read tool (`run hostname`), answers in
|
|
plain text, and the turn ends — no lifecycle tool call at all. This is
|
|
exactly the case
|
|
[`nomos/SOUL.md:105-109`](../../nomos/SOUL.md#L105) calls out by name
|
|
("a trivial read-only task... is a degenerate case... answer it and
|
|
`complete_task` with a one-line summary") — the instruction exists and is
|
|
explicit, and the model skips it anyway, consistently.
|
|
- The one session that *did* call `set_goal` (a fleet health check) did
|
|
substantial real research (`get_health_summary`, `get_state_snapshot`,
|
|
`get_signal_history`, `list_lxcs`), gave the operator a full structured
|
|
answer, and then also just stopped — no `propose_plan`, no
|
|
`complete_task`. Status: stuck at `planning` since 2026-07-11T11:35, still
|
|
showing "Running" on the board.
|
|
|
|
This means the board's "N Running / 0 Done / 0 Failed" isn't a fluke or an
|
|
edge case — it's the default outcome for essentially every task the system
|
|
has ever run. The feature as designed (terminal state is 100% dependent on
|
|
the model remembering to call one specific tool) doesn't hold up against
|
|
real model behavior, even with an explicit prompt instruction already in
|
|
place.
|
|
|
|
## Where this lives in the code
|
|
|
|
`cmd/nomos/agent.go`'s `chatWith` has exactly one place a turn ends with a
|
|
plain-text answer and no tool calls:
|
|
|
|
```go
|
|
// agent.go:359-368
|
|
if len(msg.ToolCalls) == 0 {
|
|
emit(agentEvent{Type: "text", Data: msg.Content, SessionID: sessionID})
|
|
emit(agentEvent{Type: "done", ...})
|
|
return
|
|
}
|
|
```
|
|
|
|
This is reached for the trivial-Q&A case (a turn that made zero or a few
|
|
read-only tool calls this iteration, then answered in text) and is where
|
|
43/50 of the stuck sessions are produced. There's a second, rarer exit at
|
|
the step-limit fallback (`agent.go:487-494`, `finalSummary`) with the same
|
|
gap.
|
|
|
|
Neither exit currently checks whether the session ever reached a terminal
|
|
state — the turn just ends, and `agent_sessions.status` is left wherever it
|
|
was (usually `active`, its creation-time default,
|
|
[`store.go:71,84`](../../cmd/nomos/store.go#L71)).
|
|
|
|
## Design
|
|
|
|
Two different failure shapes need two different fixes — collapsing them
|
|
into one heuristic would either auto-close genuinely in-progress structured
|
|
tasks or fail to catch the trivial-Q&A majority.
|
|
|
|
**1. Trivial/no-lifecycle-tool sessions (the 43/50 case) — auto-complete
|
|
inline, same turn.**
|
|
If a turn ends with a plain-text response (`len(msg.ToolCalls) == 0`, the
|
|
existing exit at `agent.go:359`) AND this session has never called
|
|
`set_goal` in its history, that's strong evidence this was never meant to
|
|
be a structured multi-step task — it's a one-shot question that got
|
|
answered. Call `store.completeTask` server-side right there, before the
|
|
`return`, with `outcome="success"` and a summary derived from the response
|
|
text (first ~120 chars, same truncation pattern
|
|
`buildContinuationNote` already uses at
|
|
[`continue.go:203-205`](../../cmd/nomos/continue.go#L203)). No LLM call needed
|
|
— this is a mechanical default, not a judgment call, matching the "trivial
|
|
task" case SOUL.md already describes.
|
|
|
|
If the session *has* called `set_goal` (meaning the model explicitly framed
|
|
this as a task, e.g. the fleet-health-check session), auto-completing on
|
|
the very next plain-text turn is riskier — the model may reasonably expect
|
|
to be asked something next. Skip the inline auto-complete for these; case 2
|
|
covers them.
|
|
|
|
**2. Structured (goal/plan set) sessions that stall — idle sweep, not
|
|
inline.**
|
|
Extend the existing `runContinuationWorker` ticker
|
|
([`continue.go:41-57`](../../cmd/nomos/continue.go#L41), already polling every
|
|
4s for a different purpose) with a second, coarser sweep — e.g. every 5
|
|
minutes — that finds sessions where:
|
|
- `status` is `active`, `planning`, or `executing` (not already terminal or
|
|
`awaiting_input`, which has its own resolution path), AND
|
|
- `set_goal` was called (this is a real task, not case 1), AND
|
|
- `last_active_at` is older than some idle threshold (start with 15
|
|
minutes — long enough that it's not still mid-turn, short enough that the
|
|
board doesn't lie for hours).
|
|
|
|
First idle hit: inject a system note next time nothing else touches the
|
|
session ("[System: this task has been idle for N minutes with no
|
|
`complete_task` call. If the goal is done, call it now with a summary. If
|
|
you're genuinely still working, ignore this.]") the same way
|
|
`buildContinuationNote` already injects notes into resumed sessions — reuse
|
|
`resumeSession`'s live-persist pattern
|
|
([`continue.go:106-196`](../../cmd/nomos/continue.go#L106)) so the nudge and
|
|
the model's response show up in the transcript, not silently.
|
|
|
|
If a second idle sweep finds the same session still not completed (i.e.
|
|
the nudge didn't take), auto-complete it directly with
|
|
`outcome="partial"` and a summary noting it was auto-closed after an
|
|
unanswered nudge — same reasoning as `resumeSession`'s existing
|
|
"give the task a real, operator-visible terminal state instead of leaving
|
|
it silently stuck forever" logic at
|
|
[`continue.go:179-193`](../../cmd/nomos/continue.go#L179), which already does
|
|
exactly this for a different failure mode (a resume that produces no
|
|
response). This is the same architectural pattern, applied to a session
|
|
that produces responses but never a terminal tool call.
|
|
|
|
**3. Leave `ask_operator` and gated-execution flows alone.** Those already
|
|
have real terminal signals (`awaiting_input` status, the continuation
|
|
worker's assent-window logic) — this plan only targets sessions that fall
|
|
through with no lifecycle signal at all.
|
|
|
|
## Fix plan
|
|
|
|
1. **Inline safety net (case 1)** — in `chatWith`'s plain-text exit
|
|
(`agent.go:359`), check `set_goal` was never called for this session
|
|
(cheap: track a bool while replaying `history` in the same function, no
|
|
extra query — the loop at `agent.go:209-226` already walks every
|
|
persisted message and could flag `sawSetGoal` while extracting tool
|
|
calls). If not sawSetGoal, call `completeTask` before returning.
|
|
2. **Idle sweep (case 2)** — new ticker in `continue.go` (or extend the
|
|
existing one with a slower secondary tick), a new store query
|
|
(`store.staleGoalSessions(ctx, idleThreshold)` mirroring
|
|
`pendingContinuations`'s shape), and reuse of `resumeSession`'s
|
|
live-persist injection for the nudge.
|
|
3. **Second-strike auto-close (case 2, continued)** — track nudge count (a
|
|
new `agent_sessions` column, e.g. `completion_nudges int default 0`, or
|
|
reuse the existing `summary`/attributes json instead of a schema change
|
|
if that's preferable) so the sweep can tell "never nudged" from "nudged
|
|
once already, still stuck."
|
|
4. **Backfill** — the 50 already-stuck live sessions won't get fixed by new
|
|
code alone (they're historical). One-time cleanup: run the same
|
|
case-1/case-2 classification against existing rows once the code ships,
|
|
so the board doesn't show 50 permanently-orphaned "Running" cards on top
|
|
of new correctly-terminating ones. This should be a script, not a manual
|
|
UPDATE — the classification logic will already exist in Go.
|
|
|
|
## Fix 4, revised: deletion instead of backfill
|
|
|
|
The plan as written proposed backfilling the 50 already-stuck sessions with
|
|
a mechanically-assigned outcome (`success` for case 1, `partial` for case
|
|
2). When it came time to execute that, the operator raised a better
|
|
question: these were overwhelmingly one-off test/smoke-test sessions
|
|
("hi", "what's the hostname of lxc:caddy?") with no lasting value —
|
|
assigning them a fabricated `success` outcome would make the task board
|
|
lie in the opposite direction (claiming verified success on things nobody
|
|
verified). The operator's call: delete them instead of backfilling a
|
|
guessed outcome, with one condition — don't lose any recorded knowledge.
|
|
|
|
Before deleting anything, verified directly against the database (not
|
|
assumed from reading the code):
|
|
- Zero `documents` relationship edges exist linking any of the candidate
|
|
sessions to any `knowledge_entities` row.
|
|
- Zero `upsert_knowledge` calls appear anywhere in the candidate sessions'
|
|
transcripts.
|
|
- Zero `knowledge_entities` rows exist system-wide mentioning the one
|
|
topic (`typetype`) the operator specifically asked to preserve.
|
|
|
|
`deleteSession` (`store.go:293`, already the live code path behind the
|
|
UI's "Delete task" button — reused as-is, not reimplemented) removes the
|
|
session, its messages, its own task entity, and that entity's relationship
|
|
edges — it never touches `knowledge_entities` rows or entities the task
|
|
merely referenced (e.g. `lxc:typetype` itself), only the provenance edges
|
|
back to the now-deleted task. Given the verification above, this was safe:
|
|
there was nothing to preserve because nothing had ever been recorded.
|
|
|
|
Executed in two batches, both via the same `DELETE /sessions/:id` route:
|
|
- **47 sessions** — the original candidate set from `curl
|
|
localhost:8092/sessions`, all non-`done`/`failed` at the time.
|
|
- **6 more sessions** — found *after* the first batch, when they surfaced
|
|
on the task board: `listSessions` (`store.go:193`) hardcodes
|
|
`ORDER BY last_active_at DESC LIMIT 50` with no pagination, so the
|
|
original audit's "50 sessions total" was actually "the 50 most
|
|
recently active" — it silently excluded 6 older stuck sessions from
|
|
2026-07-08 (predating the task-board feature entirely, same trivial
|
|
"hi"/smoke-test pattern). Worth knowing about `listSessions`'s cap for
|
|
any future audit of this table — a `count(*)` query directly against
|
|
the database is the only way to get a true total.
|
|
|
|
Final state: `agent_sessions` holds exactly 3 rows — the two `done` and
|
|
one `failed` sessions produced during live verification of fixes 1-3.
|
|
|
|
## Implementation order
|
|
|
|
1. Fix 1 (inline safety net) first — it's the highest-leverage, lowest-risk
|
|
change (self-contained, no schema change, covers 43/50 of the evidence).
|
|
2. Fix 2+3 (idle sweep + second-strike) — needs the schema decision
|
|
(new column vs. attribute) settled first; smaller blast radius than 1
|
|
but touches the ticker/worker machinery, deserves its own review pass.
|
|
3. Fix 4 (backfill) last, once 1-3 are deployed and verified live — running
|
|
it before the code ships would just recreate the same gap for new
|
|
sessions created in between.
|
|
|
|
## Verification
|
|
|
|
- After fix 1: start a few trivial one-shot chats against the live agent
|
|
(`hostname`-style questions), confirm each session reaches `status=done`
|
|
immediately after the answer, via `curl localhost:8092/sessions/:id` or
|
|
the task board.
|
|
- After fix 2+3: manually let a goal-bearing session go idle past the
|
|
threshold (or lower the threshold for a local test run), confirm the
|
|
nudge appears in the transcript, then confirm second-strike auto-close
|
|
fires if the nudge is ignored.
|
|
- Re-run the same audit query used to find this bug
|
|
(`curl localhost:8092/sessions` → status histogram) a day after deploy;
|
|
the "stuck active/planning forever" count should track only genuinely
|
|
in-flight tasks, not accumulate.
|
|
|
|
## Open questions
|
|
|
|
- **Outcome for case-1 auto-complete**: always `"success"`, or worth a
|
|
cheap heuristic (e.g. scan the final text for obvious failure language)?
|
|
Recommend starting with always-`"success"` — SOUL.md's own trivial-task
|
|
guidance doesn't distinguish, and a wrong "success" on a genuinely-failed
|
|
one-shot lookup is low-stakes (the transcript still shows the real
|
|
answer; nothing acts on the outcome besides the board's color).
|
|
- **Idle threshold (15 min) and nudge-to-close gap**: arbitrary starting
|
|
points, not measured against real task durations — worth revisiting after
|
|
a week of the new sessions' real timing data exists.
|
|
- **Schema change for nudge tracking**: a new column is simpler to query
|
|
than packing state into existing JSON, but adds a migration — worth
|
|
confirming that's acceptable before starting fix 2+3 (this plan defers
|
|
that call to whoever implements it, per Implementation order above).
|