Reframes the chat surface as a board of tasks: each task carries a goal, a plan approved once, a lifecycle status, an outcome, and a knowledge loop that links learnings to the involved entities (and the task entity itself) via relationships so future tasks compound. Supersedes the sidebar-only framing and the free-form chat portion of the control-room web UI plan. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
16 KiB
2026-07-11 — Tasks: the chat page as goal-structured autonomous work
Status: Planned — supersedes the sidebar-only framing; this is the definitive direction for the chat surface. Any earlier plan detail that conflicts with "the chat page is a board of tasks" is overridden by this doc (notably the Chat portion of control-room-webui, which described chat as a free-form session list).
The vision (operator, distilled)
Structure the whole chat page as tasks. A task is a card — you see its status (running / completed / failed), its description. A task is a goal: "install the service", "give me the key status of X". The agent takes the goal, finds what it needs, proposes a plan, the operator approves it once — that single approval is the only one needed — and the agent then executes the whole plan autonomously until the goal is achieved. Every task has a completion status: successful or not, and its learnings move to knowledge, attached via relationships to the entities that were involved, so future tasks — successful or unsuccessful — make the agent better over time. Inside a task is the conversation (tools, thinking, questions if needed); the sidebar shows the live context: which entities the agent is exploring, the steps and their status, whether the task succeeded, and the knowledge it recorded — all populated in real time as the agent works.
Three pillars: task as the unit, one approval → autonomous execution, a knowledge loop that compounds.
The reframe
Today a "session" is a title + a flat message list (migrations/015); a "plan" is prose the model types; there is no goal, status, outcome, or step object. We elevate the session into a task:
- A task = a session with a goal, a plan, a lifecycle status, and an outcome. One task per chat. The chat page becomes a task board of status cards; opening a card shows the task: conversation in the center, live context in the sidebar.
- The plan is approved once. Machinery already exists — the assent window + event-driven auto-continuation shipped in autonomous-plan-execution (continue.go, assent.go) already turn a single approval into an autonomy grant the agent runs to completion. This plan gives that flow a structured surface: the one thing the operator approves is a named, stepped plan, and progress is visible.
- On completion the task deposits knowledge, linked by relationships to the
entities involved and to the task itself, tagged success/failure — and
future tasks read it back at planning time. The substrate exists:
upsert_knowledgewrites a knowledge doc-entity and adocumentsrelationship (server.go:1519);get_entity_knowledgereads it (server.go:170). We add task-linkage, an outcome flavor, and retrieval-at-planning.
Builds on / aligns with
- general-gated-execution — the
classifier +
runprimitive is the execution substrate; a plan step is just a described unit of work mapping to arun/request_executioncall. No fixed step enum. - autonomous-plan-execution — the single-approval autonomy window + auto-continuation loop.
- The knowledge tools + relationships graph (
upsert_knowledge,get_entity_knowledge,get_relations, the temporalrelationshipstable).
Data model (migration 018_tasks.up.sql)
Elevate the session into a task; add plan steps, questions, and the task→knowledge linkage.
ALTER TABLE agent_sessions
ADD COLUMN goal TEXT NOT NULL DEFAULT '',
ADD COLUMN status TEXT NOT NULL DEFAULT 'active',
-- active | planning | awaiting_approval | executing
-- | awaiting_input | done | failed | abandoned
ADD COLUMN outcome TEXT, -- success | failure | partial (NULL until done)
ADD COLUMN summary TEXT NOT NULL DEFAULT '', -- one-line result, shown on the card
ADD COLUMN entity_id UUID; -- the task's OWN entity (type 'task'), for
-- knowledge/relationship linkage (see below)
CREATE TABLE session_plan_steps (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
session_id UUID NOT NULL REFERENCES agent_sessions(id) ON DELETE CASCADE,
seq INT NOT NULL,
title TEXT NOT NULL,
detail TEXT NOT NULL DEFAULT '',
status TEXT NOT NULL DEFAULT 'pending',
-- pending | running | done | failed | skipped | blocked
execution_id UUID,
target_slug TEXT,
started_at TIMESTAMPTZ,
finished_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_plan_steps_session ON session_plan_steps(session_id, seq);
CREATE TABLE session_questions (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
session_id UUID NOT NULL REFERENCES agent_sessions(id) ON DELETE CASCADE,
prompt TEXT NOT NULL,
context JSONB NOT NULL DEFAULT '{}', -- { entities:[], options:[], why:"" }
status TEXT NOT NULL DEFAULT 'open', -- open | answered | dismissed
answer TEXT,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
answered_at TIMESTAMPTZ
);
CREATE INDEX idx_questions_session_open
ON session_questions(session_id) WHERE status = 'open';
The task as an entity. Each task gets a row in entities (type task,
slug task:<short-id>), stored in agent_sessions.entity_id. This is what
makes the knowledge loop use the existing graph machinery unchanged:
knowledge and involved-entity links hang off the task entity via
relationships, exactly like any other entity.
Task lifecycle
created → planning → awaiting_approval → executing ⇄ awaiting_input → done
│ (outcome:
└──────────────→ failed success/
failure/
partial)
- planning: agent calls
get_entity_knowledgeon the target(s) first (prior learnings), thenset_goal+propose_plan. - awaiting_approval: the plan is the single approval gate. Operator approves → opens the assent window (existing) → executing.
- executing: steps flip pending→running→done via
update_plan_stepand the execution→step auto-close (below); the agent runs autonomously (auto-continuation) with no per-step re-approval. - awaiting_input: only when the agent hits a real decision →
ask_operator; answering resumes execution. - done/failed: agent sets
outcome+summaryand deposits knowledge.
Single approval → autonomous execution
Already the behaviour of the assent window + auto-continuation. This plan makes the approved object a structured plan rather than an individual command: approving the plan (one click / one "go ahead") authorizes every read-only + config-mutation step in it. Destructive steps still require typed confirmation unless named in the approved plan (the existing pre-authorized destructive-step rule). Nothing new in the execution engine — we're giving it a legible unit to approve and to show progress against.
The knowledge loop (capture → link → retrieve)
Capture (task end). On done/failed, the agent (nudged by SOUL, enforced
by a server-side fallback) calls upsert_knowledge with the concrete learning —
what worked, what didn't, the gotcha — and we link the resulting knowledge
doc-entity to:
- the entities involved (already supported via
about), and - the task entity (
agent_sessions.entity_id), via a newoutcome_of/produced_byrelationship, tagged{"outcome":"success|failure"}.
Link. Involved entities are captured cheaply: every entity.touched (below)
records a relationships edge task —involved→ entity. So a task's entity
neighborhood is its involved-entity set — queryable with the existing
get_relations.
Retrieve (task start). At planning, before proposing, the agent pulls
prior knowledge for the target entities (get_entity_knowledge) — which now
surfaces both successful and failed prior tasks (the outcome tag lets it weight
"last time apt install docker.io failed on Debian, used get.docker.com
instead"). This is the compounding: each task's outcome becomes the next task's
prior. SOUL makes this the first planning move.
UI
Task board (replaces the raw session rail / empty chat state)
Sessions.svelte / SessionRail.svelte become a board of task cards. Each card:
- goal as the title, one-line
summary, - a status pill (running ◐ / awaiting you / done ✓ / failed ✗) with the
step progress (
4/6), - outcome color on completion, knowledge-count badge (♦ 2 learned),
- click → open the task.
Grouped/filterable by status (Running, Needs input, Done, Failed). "New task" replaces "new chat" — the empty state asks for a goal.
Task detail = conversation + live context sidebar
Center column: the existing chat transcript (tools, thinking, questions inline) — unchanged rendering (Chat.svelte).
Right sidebar becomes TaskContextPanel.svelte, populated in real time, top
to bottom:
- GoalHeader — goal + status pill + outcome (once done); editable goal.
- PlanProgress — ordered steps, live status icons,
4/6bar, click a step → scroll chat to its tool call / open its execution output. - OperatorQuestion — pinned structured card when a question is open: prompt,
why, context-entity chips (→ EntitySheet), option buttons or free-text. Answering POSTs the answer and resumes the agent. Same card also renders inline in the transcript at the point it was raised. (The operator's "structured component with relevant context.") - LiveEntityPanel — the SessionGraph
upgraded from passive to live:
entity.touched→ the node pulses + "now touchinglxc:foo";health.changed→ recolor + transienthealthy→degradeddiff badge. - Outcome & Knowledge — on completion: success/failure banner, the
summary, and the knowledge notes recorded (links to the knowledge entities), i.e. the SessionDigest evolved into a task-outcome card.
Real-time event contract (global /events/stream)
The panel is driven by the always-on events stream,
not the per-turn chat SSE — so it stays live during server-side
auto-continuation (when no chat turn is open) and survives a tab reload. New
types, each carrying correlation_id = session_id:
| type | data |
|---|---|
task.status |
{ status, outcome?, summary? } |
goal.set |
{ goal } |
plan.proposed |
{ steps:[{seq,title,detail,target_slug}] } |
plan.step.started / plan.step.finished |
{ step_id, seq, status, execution_id? } |
question.raised / question.answered |
{ question_id, prompt?, context?, answer? } |
entity.touched |
{ slug, tool } |
knowledge.recorded |
{ title, about, outcome } |
entity.touched is emitted from the withActivityLogging wrapper
(server.go:832) — it wraps every tool call, so
touched-entity tracking needs zero agent changes; it also writes the
task —involved→ entity relationship. health.changed already exists.
Agent surface (new MCP tools + SOUL)
Thin declarations that write the tables/relationships and publish the event in-process (event and row commit together):
set_goal(goal)propose_plan(steps:[{title,detail?,target_slug?}])update_plan_step(seq,status,execution_id?)— plus the execution's terminal status auto-closes its linked step where phase3.go finalizes executions (belt and suspenders).ask_operator(prompt,options?,context_entities?,why?)— creates the question, status→awaiting_input, ends the turn; answer resumes via the existing assent/continuation path.complete_task(outcome,summary)— sets outcome/summary, status→done/failed; server enforces "a completed task must have deposited ≥1 knowledge note" (fallback: auto-summarize into one if the model forgot).
SOUL: "Every task has a goal. First, read prior knowledge for the target
entities (get_entity_knowledge) — learn from past tasks, successful or not.
Then set_goal + propose_plan. Execute autonomously after approval, marking
steps. Ask via ask_operator only for real decisions. When the goal is
verified, complete_task with the outcome and record what you learned."
Implementation order
- Migration
018+ task-entity creation (ataskentity per session) + store methods. Sessions gain goal/status/outcome/summary; no behaviour change. entity.touched+task —involved→ entityfromwithActivityLogging— cheapest live win; graph starts pulsing, involved-set is captured for free.- Knowledge loop close:
complete_task+ retrieval-at-planning in SOUL + outcome-taggedoutcome_oflink. Makes tasks compound. set_goal/propose_plan/update_plan_step+ execution→step auto-close.ask_operatorend-to-end (tool → question → pinned card inline+panel → answer resumes).- UI: TaskContextPanel (GoalHeader, PlanProgress, OperatorQuestion,
LiveEntityPanel, Outcome/Knowledge) +
workspace.tsstore + REST hydration (GET /sessions/{id}/{plan,questions}). - UI: Task board — session rail/list → status-card board, "new task" flow.
Each step ships value: 2 = live entity awareness; 3 = compounding knowledge; 4-5 = plan progress + interactive questions; 6-7 = the full task surface.
Verification
- Run "deploy TypeType as an LXC on strong" as a task. Expect: at planning the
agent reads prior knowledge for
host:strong;propose_planrenders steps; operator approves once; steps flip live; the touched node pulses; a mid-flow ambiguity surfaces as anask_operatorcard answered in the panel; on successcomplete_tasksets outcome=success, deposits a knowledge note linked tolxc:typetype,host:strong, and the task entity. - Start a second task touching
host:strong; confirm the first task's knowledge surfaces at planning (get_entity_knowledge) — the compounding loop. - Reload the tab mid-execution → panel rehydrates from REST and keeps updating from the global stream (proves it isn't chat-SSE-bound).
- Board shows the task moving Running → Done with the right outcome color and
knowledge badge.
SELECT status, outcome FROM agent_sessionsshows a real lifecycle, not allactive. get_relationson the task entity returns its involved entities + produced knowledge.
Open questions
- One task per session vs sequential tasks in a chat — v1: one task = one session (matches "each chat is a goal"). A new goal starts a new task. Multi-task threads are a later extension.
- Failure knowledge weighting — do we just tag
outcome:failureand let the model judge, or add explicit "avoid this" surfacing at planning? Lean tag-only first; revisit if the agent repeats known-failed approaches. complete_taskenforcement — hard-require a knowledge note (block completion) or soft (auto-generate a stub)? Lean soft, so a trivial "give me status" task isn't forced to invent a learning.- Board vs thread for very short tasks ("key status of X") — a status query is a degenerate task (no plan, instant done). Render it as a lightweight card that never shows an approval, so the board isn't cluttered with heavyweight chrome for one-shot questions.