The root cause behind "the agent stops at the first error and doesn't recover":
provisioning executions run ASYNCHRONOUSLY (pct_create fires the SSH work in a
goroutine and returns "running" immediately), so the agent's turn ENDS before
the result exists. The agent literally isn't running when the step fails — it
can't react to a failure it never observes. The only thing that fed results
back was the operator typing "continue" after every async step: the human was
the event loop. (In the flagged 18-message session the operator typed
continue/proceed/?? eight times while the agent correctly diagnosed each failure
but couldn't advance a step on its own.)
This makes the system the event loop instead:
- migrations/017: nomos_plan_executions links each gated execution to the chat
session that started it.
- cmd/nomos: after a tool result, any "execution <uuid>" it started is linked
to the session. A background worker (continue.go) polls for those executions
reaching a terminal state and — while the agent has an open assent window (an
approved plan is in flight) — re-invokes the agent with the result
("execution X completed/failed: <result>"), so it proceeds to the next step
or diagnoses+fixes the failure, with no operator tick. Guarded against loops
(mark-continued before running) and bounded by the 30-min window.
- chatWith(): chat() variant that injects the finished-execution note after
replayed history without persisting a fake user turn.
- DecideApproval: approving a step by ANY route (button or chat-assent) now
opens the assent window, so auto-continuation works regardless of how the
operator approved — previously only typing "go ahead" opened it.
- SOUL: the agent is told it will be auto-re-invoked when async steps finish —
don't poll get_execution_status, don't wait for "continue"; end the turn and
keep going step by step until the goal is verified or a genuine blocker.
This is the root fix, not another per-command patch: you can't enumerate every
failure of an unbounded action space, but you can give the agent a loop that
observes each result and adapts — because "do anything" always includes "the
first attempt failed."
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
26 lines
1.3 KiB
SQL
26 lines
1.3 KiB
SQL
-- 017_nomos_plan_executions.up.sql
|
|
-- Links a gated execution back to the chat session that initiated it, so the
|
|
-- Nomos auto-continuation worker can re-invoke the agent for that session when
|
|
-- the (asynchronous) execution finishes. This is the "the system is the event
|
|
-- loop, not the human" foundation: the human no longer types "continue" after
|
|
-- every async step — the worker feeds each execution's result back into the
|
|
-- agent automatically.
|
|
--
|
|
-- Owned by the nomos process. execution_id references the execution entity by
|
|
-- UUID but intentionally without a hard FK — nomos records the link from the
|
|
-- tool-result text it gets back, and we don't want a race between the API
|
|
-- creating the execution entity and nomos linking it to break the insert.
|
|
CREATE TABLE IF NOT EXISTS nomos_plan_executions (
|
|
execution_id UUID PRIMARY KEY,
|
|
session_id UUID NOT NULL,
|
|
-- when the worker fed this execution's result back to the agent (NULL = not yet)
|
|
continued_at TIMESTAMPTZ,
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
|
|
);
|
|
|
|
-- Worker query: find terminal executions not yet fed back. Partial index on the
|
|
-- not-yet-continued rows keeps the poll cheap as history accumulates.
|
|
CREATE INDEX IF NOT EXISTS idx_nomos_plan_exec_pending
|
|
ON nomos_plan_executions (created_at)
|
|
WHERE continued_at IS NULL;
|