feat: event-driven auto-continuation — agent runs an approved plan to completion
The root cause behind "the agent stops at the first error and doesn't recover":
provisioning executions run ASYNCHRONOUSLY (pct_create fires the SSH work in a
goroutine and returns "running" immediately), so the agent's turn ENDS before
the result exists. The agent literally isn't running when the step fails — it
can't react to a failure it never observes. The only thing that fed results
back was the operator typing "continue" after every async step: the human was
the event loop. (In the flagged 18-message session the operator typed
continue/proceed/?? eight times while the agent correctly diagnosed each failure
but couldn't advance a step on its own.)
This makes the system the event loop instead:
- migrations/017: nomos_plan_executions links each gated execution to the chat
session that started it.
- cmd/nomos: after a tool result, any "execution <uuid>" it started is linked
to the session. A background worker (continue.go) polls for those executions
reaching a terminal state and — while the agent has an open assent window (an
approved plan is in flight) — re-invokes the agent with the result
("execution X completed/failed: <result>"), so it proceeds to the next step
or diagnoses+fixes the failure, with no operator tick. Guarded against loops
(mark-continued before running) and bounded by the 30-min window.
- chatWith(): chat() variant that injects the finished-execution note after
replayed history without persisting a fake user turn.
- DecideApproval: approving a step by ANY route (button or chat-assent) now
opens the assent window, so auto-continuation works regardless of how the
operator approved — previously only typing "go ahead" opened it.
- SOUL: the agent is told it will be auto-re-invoked when async steps finish —
don't poll get_execution_status, don't wait for "continue"; end the turn and
keep going step by step until the goal is verified or a genuine blocker.
This is the root fix, not another per-command patch: you can't enumerate every
failure of an unbounded action space, but you can give the agent a loop that
observes each result and adapts — because "do anything" always includes "the
first attempt failed."
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -193,6 +193,19 @@ operator if:
|
||||
Do NOT stop after every step waiting for "continue". The operator approved
|
||||
the plan — execute it end to end.
|
||||
|
||||
**Automatic continuation — you are re-invoked when async steps finish.** Some
|
||||
steps (`pct_create`, `apt_upgrade`) run asynchronously: the tool returns
|
||||
"execution <id> running" immediately, and the actual work (which can take
|
||||
minutes) finishes later. **You do NOT need to poll `get_execution_status` in a
|
||||
loop, and you do NOT need the operator to say "continue".** When such a step
|
||||
finishes, the system automatically re-invokes you with a
|
||||
`[System: execution <id> finished with status=…]` note carrying the result.
|
||||
So: after you launch an async step, briefly say what you're doing and END your
|
||||
turn — you will be woken up with the result and should then proceed to the next
|
||||
step (on success) or diagnose and fix (on failure). Keep going, step by step,
|
||||
until the whole goal is verified working — the loop only ends when you report
|
||||
completion or hit a genuine blocker.
|
||||
|
||||
**When a step fails:** diagnose the error, try an alternative approach, and
|
||||
continue. For example, if `docker: command not found` appears, install Docker
|
||||
CE via `get.docker.com` and retry. If a package is missing, install it. If a
|
||||
|
||||
Reference in New Issue
Block a user