fix stuck 'Agent is thinking' + flip activity to old-to-new
- AgentIndicator now shows only during stream or when running tools exist (not on session status=executing which never cleared) - Activity timeline: oldest-first ordering (reads top-to-bottom naturally) - Removed unused liveStatus derivation and currentTask import from Chat - Plan: Phase A+B+C for activity gaps + plan-approve-once policy
This commit is contained in:
159
plans/2026-07-14-activity-gaps.md
Normal file
159
plans/2026-07-14-activity-gaps.md
Normal file
@@ -0,0 +1,159 @@
|
||||
# 2026-07-14 — Session review + Activity timeline refinements
|
||||
|
||||
**Status:** Planned
|
||||
|
||||
## Session analysis: `722d8878` (2026-07-14T10:45)
|
||||
|
||||
"Fleet-wide audit: check all services, identify what needs updating"
|
||||
|
||||
### What happened
|
||||
- User asked for fleet audit → agent called `set_goal` (session → executing)
|
||||
- 82 tool calls in ONE turn: 58 `run`, 8 `get_entity`, 4 `get_relations`, etc.
|
||||
- **No `propose_plan`** — no plan steps, entered executing directly
|
||||
- **No `complete_task`** — agent produced a final text response but never closed the task
|
||||
- Session stuck at `executing` with a text conclusion but no terminal state
|
||||
|
||||
### Gaps found
|
||||
|
||||
| # | Issue | Root cause |
|
||||
|---|---|---|
|
||||
| 1 | "Agent is thinking" stuck at end | `liveStatus === 'executing'` is true even after the turn ends. AgentIndicator shows generic message because there are no running activity entries to describe. |
|
||||
| 2 | No plan proposed | Model skipped `propose_plan` — possibly because it found prior knowledge and skipped to execution. `setGoal` now sets `executing` directly (our fix), which makes `propose_plan` optional — but the sidebar "Plan" section shows "No plan yet" permanently. |
|
||||
| 3 | Task never completed | Model produced final text but never called `complete_task`. The idle sweep (2 min now) will nudge, then auto-close. |
|
||||
| 4 | Approvals invisible in Activity | 58 `run` calls, many requiring approval. These show in chat via InlineApproval but NOT in the Activity timeline. The operator has to switch to Ops page to track approvals. |
|
||||
| 5 | Activity is newest-first | Timeline shows newest at top. Feels unnatural for a sequential log — bottom-scrolling with newest at bottom is more intuitive for "watching" what the agent does. |
|
||||
| 6 | Knowledge recorded not shown | The agent recorded knowledge but it doesn't appear in Activity if it came from prior knowledge entries via `get_knowledge_content`. |
|
||||
|
||||
---
|
||||
|
||||
## Plan
|
||||
|
||||
### 1. Fix "Agent is thinking" stuck indicator
|
||||
|
||||
**Root cause:** AgentIndicator shows when `active={$streaming || liveStatus === 'executing'}`. After the turn completes, `liveStatus` is still `'executing'` and there are no running entries, so the label falls through to the generic `"Agent is thinking…"` fallback.
|
||||
|
||||
**Fix:** Change the active condition to only show when there's actual work:
|
||||
```ts
|
||||
active={$streaming || $activityLog.some((e) => e.status === 'running')}
|
||||
```
|
||||
This way it shows during streaming AND when there are running tools (auto-continuation), but NOT when the session is just "executing" with no active work.
|
||||
|
||||
Also: when the last assistant message has text AND no pending tool calls, auto-hide the indicator. The `liveStatus === 'executing'` check is too broad — it covers the entire session lifetime.
|
||||
|
||||
### 2. Approvals in Activity timeline
|
||||
|
||||
**What:** InlineApproval cards show in chat but not in Activity. Every `run` that queues an execution with "requires approval" should appear as an entry in the Activity timeline.
|
||||
|
||||
**How:**
|
||||
- In `activity.ts`, detect tool results containing "requires approval" + execution ID
|
||||
- Add `type: 'approval_pending'` entries with the execution ID, target, action, and status
|
||||
- Poll execution status and update the entry (pending → approved → running → completed/failed)
|
||||
- The InlineApproval component stays in chat for the Approve/Deny buttons
|
||||
- Activity shows the full lifecycle: approval requested → approved → running → done
|
||||
|
||||
### 3. Old-to-new ordering
|
||||
|
||||
**Fix:** Remove `.sort((a, b) => b.timestamp - a.timestamp)` → change to `.sort((a, b) => a.timestamp - b.timestamp)` or no sort at all (entries are already added in chronological order).
|
||||
|
||||
This means the timeline reads top-to-bottom as the session unfolds. Currently `newest at top` means the "Goal" and "Plan" entries appear at the bottom, which is confusing.
|
||||
|
||||
### 4. Knowledge detection
|
||||
|
||||
**Fix:** Extend the knowledge detection in `activity.ts` to also catch `upsert_knowledge` calls from `tool_use` events (not just `tool_result`), so the entry appears as "running" while recording and then "done" when the result comes back.
|
||||
|
||||
### 5. Auto-hide indicator when turn ends with text
|
||||
|
||||
**Fix:** Detect when the last assistant message has text content AND there are no pending tool_use entries without matching tool_result. In that case, the turn is complete — don't show the indicator.
|
||||
|
||||
---
|
||||
|
||||
## Plan-approve-once (new policy)
|
||||
|
||||
### Problem
|
||||
|
||||
Today: agent calls `propose_plan` + `run` × 10 in the same turn. Each `run`
|
||||
queues an individual approval. Operator sees 10 "requires approval" cards.
|
||||
After operator types "yes", each one is individually approved, THEN the
|
||||
assent window opens and future calls auto-run.
|
||||
|
||||
The operator shouldn't see per-action approvals when they already approved
|
||||
the plan. The plan IS the approval. Individual actions within an approved
|
||||
plan should auto-execute.
|
||||
|
||||
### Target
|
||||
|
||||
```
|
||||
User: "audit fleet"
|
||||
Agent: "Here's my plan: 1. List LXCs 2. Check apt on each 3. Report" ← proposes plan
|
||||
[Proposed plan: 3 steps] [Approve plan?]
|
||||
User: "approved"
|
||||
Agent: ◉ Listing containers… ← auto-runs
|
||||
◉ Checking apt on lxc:jellyfin… ← auto-runs
|
||||
...
|
||||
"Done. 19 LXCs have pending updates."
|
||||
```
|
||||
|
||||
One approval for the plan. All actions within it auto-execute. No per-action
|
||||
approval cards. Only re-approve when the agent calls `propose_plan` again
|
||||
(significant plan change).
|
||||
|
||||
### How (server-side)
|
||||
|
||||
The classification logic in `internal/mcp/server.go:run()` needs to know whether
|
||||
a plan-approval-assent-window is active for this session. Currently it checks
|
||||
`autonomy_settings` for the assent window key. The change: when `propose_plan`
|
||||
is called, pre-activate the window with a "plan-proposed" state. When the
|
||||
operator approves, transition to "plan-active". `run` calls within an active
|
||||
plan window auto-execute at `config_mutation` level.
|
||||
|
||||
Key change in `store.go:proposePlan()`:
|
||||
```go
|
||||
// Pre-record a plan-proposed window so that run calls know a plan is pending approval.
|
||||
// Once approved, this becomes the full assent window.
|
||||
key := planWindowKey(agentID, sessionID)
|
||||
s.pool.Exec(ctx, `INSERT INTO autonomy_settings (key, value) VALUES ($1, 'proposed')
|
||||
ON CONFLICT (key) DO UPDATE SET value = 'proposed'`, key)
|
||||
```
|
||||
|
||||
Then in the `run` handler, check for `plan-active` OR `assent-active`:
|
||||
```go
|
||||
// If a plan window is active, this run call is part of the approved plan
|
||||
// and config_mutation commands auto-execute without individual approval.
|
||||
if planWindowActive(ctx, agentID, sessionID) { ... }
|
||||
```
|
||||
|
||||
### How (frontend)
|
||||
|
||||
- Instead of 10 individual InlineApproval cards, show ONE "Approve plan?" card
|
||||
- When approved, all queued `run` calls from the plan turn auto-grant
|
||||
- Activity timeline shows plan approval as one entry: "✓ Plan approved — 12 actions"
|
||||
- Subsequent `run` calls show with an "auto (plan)" badge instead of approval cards
|
||||
|
||||
---
|
||||
|
||||
## Implementation plan (consolidated)
|
||||
|
||||
### Phase A — Quick fixes (today)
|
||||
|
||||
| # | Fix |
|
||||
|---|---|
|
||||
| A1 | AgentIndicator: only show when `$streaming || hasRunningActivity` (not on `liveStatus === 'executing'`) |
|
||||
| A2 | Activity timeline: old-to-new ordering |
|
||||
| A3 | Knowledge entries: detect `tool_use` events for running state |
|
||||
|
||||
### Phase B — Approvals in Activity
|
||||
|
||||
| # | Fix |
|
||||
|---|---|
|
||||
| B1 | `activityLog`: add `approval_pending` / `approval_granted` / `execution_running` / `execution_done` lifecycle entries |
|
||||
| B2 | Poll execution status and update Activity entries inline |
|
||||
| B3 | InlineApproval stays in chat (needs operator action), but lifecycle tracked in Activity |
|
||||
|
||||
### Phase C — Plan-approve-once (policy change)
|
||||
|
||||
| # | Fix |
|
||||
|---|---|
|
||||
| C1 | Backend: plan window in `autonomy_settings`, checked by `run` handler |
|
||||
| C2 | Frontend: single "Approve plan?" card instead of per-action cards |
|
||||
| C3 | Backend: when plan is approved, auto-grant all pending executions from the plan turn |
|
||||
| C4 | Activity: plan approval as single timeline entry with action count badge |
|
||||
@@ -127,8 +127,8 @@ export const activityLog = derived([messages, planSteps, currentTask], ([$msgs,
|
||||
})
|
||||
}
|
||||
|
||||
// Sort newest first
|
||||
entries.sort((a, b) => b.timestamp - a.timestamp)
|
||||
// Sort oldest first
|
||||
entries.sort((a, b) => a.timestamp - b.timestamp)
|
||||
|
||||
return entries
|
||||
})
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
<script lang="ts">
|
||||
import { messages, streaming, connectionState, currentSession, sendMessage, cancelStream, reconnect, error, chatErrors, dismissError } from '$lib/stores/chat'
|
||||
import { activityLog } from '$lib/stores/activity'
|
||||
import { currentTask } from '$lib/stores/workspace'
|
||||
import SessionRail from '$lib/components/SessionRail.svelte'
|
||||
import TaskContextPanel from '$lib/components/TaskContextPanel.svelte'
|
||||
import InlineApproval from '$lib/components/InlineApproval.svelte'
|
||||
@@ -103,8 +102,6 @@
|
||||
if ($streaming) return
|
||||
sendMessage(q)
|
||||
}
|
||||
|
||||
const liveStatus = $derived($currentTask?.status ?? ($currentSession ? 'active' : null))
|
||||
</script>
|
||||
|
||||
<div class="flex h-full min-h-0">
|
||||
@@ -152,7 +149,7 @@
|
||||
</div>
|
||||
{/each}
|
||||
<AgentIndicator
|
||||
active={$streaming || liveStatus === 'executing'}
|
||||
active={$streaming || $activityLog.some((e) => e.status === 'running')}
|
||||
lastActivity={$activityLog.find((e) => e.status === 'running') ?? null}
|
||||
error={$error}
|
||||
/>
|
||||
|
||||
Reference in New Issue
Block a user