1540f74342d04d490b81e12b24d8131df734ce38
582 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| 1540f74342 |
feat(audit): read-only knowledge-graph drift report + skill
Adds audit_knowledge_graph (MCP tool) and GET /api/v1/audit/drift (endpoint) backed by a shared internal/audit package. One pass surfaces the structural gaps an operator otherwise finds by accident: orphan check entities, checks targeting deprecated/destroyed entities, probes stuck down/unknown, unmonitored declared types, and live edges pointing at destroyed targets. Each finding carries a suggested remediation runbook. Read-only and safe to run unattended. Ships the knowledge-graph-audit skill (SKILL.md + seeded runbook) that interprets the report and routes findings to the lifecycle runbooks. |
|||
| c7729b2ef6 |
fix(scheduler): stop monitoring deprecated/destroyed targets
ListEnabledCheckDefs now LEFT JOINs the target entity and excludes rows whose target is deprecated or destroyed, so retired things (secrets-issuance, homelab-mcp, the dead secrets ingress route) stop generating permanent false alarms instead of waiting for an operator to disable the check_def by hand. coverageSweep's None() branch previously did nothing, so a type changed from declared monitoring to `monitoring: none` (dns-zone) left its open `unmonitored` signals lingering forever — a None() entity never gains a check, so the hasCheck resolution path never fired. It now resolves those signals. |
|||
| b8b4aa2aee |
feat(remote): route LXC/VM checks through the Proxmox host, not direct SSH
The scheduler SSHed each guest directly and assumed a deployed probe script plus working root SSH at the guest's address — false for headless (nfs-export), keyless (teddycloud), mesh-only (rclone), and macOS (mac-mini) targets, which left 49 enabled checks stuck "down" on a healthy fleet. Extract the MCP run tool's resolveExecTarget into a shared internal/remote package and make it the single execution path for both the scheduler and MCP. LXC/VM checks now host-hop via pct exec / qm guest exec through the owning Proxmox host (no per-guest lan_ip, sshd, or authorized key needed); hosts and workstations resolve their address and user live, so mac-mini's `user: dtoro` is honored without a re-seed. Address preference now prefers public_ipv4 over mesh, so netbird-vps is probeable from the scheduler container. cpu_check.sh gains a real Darwin branch (it reported cpu_pct 0 before). checkdefaults.resolveSSHUser reads the top-level `user` attribute too. A machine-target resolution failure is now logged before falling back to baked config, so a broken probe-config is distinguishable from a real outage. |
|||
| c10f6920cd |
fix: blast radius walks dependency direction, and reachability survives no ICMP
Two things the entity window redesign surfaced but deliberately left alone. **blast_radius answered the wrong question.** It walked source→target for every relationship type, but which end of an edge is the dependent differs per type: "machine hosts container" means the target breaks, while "service depends-on service" and "ingress routes-to service" mean the SOURCE breaks. Walking everything forwards was right for hosts/provides and backwards for everything else — and swept in 2,800+ documents/involves/targets edges of pure bookkeeping, so the result contained tasks and executions that cannot break. Direction is now declared per relationship type in seeds/ontology.yaml (blast_direction: forward | backward | none), the same shape as the entity types' monitoring: declaration, and defaults to none so an undeclared edge contributes nothing rather than a confidently wrong answer. It also needed a modelling fix: `routes-to` names an ingress's BACKEND, so nothing recorded that all 21 public hostnames are terminated by caddy. A `served-by` edge type now says so. pool:ludo-lvm 2 -> 23 (every container storing on it, then their services) lxc:caddy 4 -> 22 (service:caddy, then all 21 ingress routes) service:authentik 7 (what authenticates via it) **Every ping check was reporting down.** Not a host:strong false positive: all seven, including ws:mac-mini — the Docker host itself. The scheduler runs in Docker on macOS, whose VM does not route ICMP to the LAN; loopback pings succeed and every LAN ping fails. Under health aggregation each broken probe dragged its entity to down. The question the check exists to answer is "is it reachable", and ICMP is only one way to ask it. checkPing now falls back to a TCP connect before concluding anything, which restores an honest verdict for the four hosts that are genuinely up while leaving the genuinely unreachable ones down. TestBlastRadiusTerminatesOnCycles asserted the old direction (caddy=1, authentik=2 — the cycle walked the wrong way); it now asserts the corrected depths, and its exact-node-count check is relaxed because walking the right way also surfaces the seed's own real dependents, which are correct answers. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| ad29295c93 |
feat(web): make the entity window a triage surface, not a data dump
The window rendered the same 13 collapsible sections for every entity, sorted
only by "does it have content". Audit trail carried the same visual weight as
Health, and the window answered "what data do we hold about X?" rather than
"what do I need to know, and what should I do?".
Measured against prod: host:hubris has 223 relations, 1,601 events, 2.7M metric
samples and 148 executions; an ingress route has three facts. Both got 13
identical headers. Expanding a host put ~540 interactive elements on screen.
- **A verdict header that never collapses.** Not just "down" but *why*:
"ping failing · 5 of 6 checks passing". That line did not previously exist
and could not have — checks rendered as configuration, never as results.
- **Sections composed per type.** A document has no checks, metrics or blast
radius; a signal or execution is a record, not a thing. Infrastructure gets
Status/Impact/Activity/Metrics/Reference, knowledge types lead with Content,
records get a minimal view. Unknown types fall back to infrastructure so a
new entity type is never a blank window.
- **Status replaces Monitoring**, showing each check's own verdict and when it
last ran — the section that answers the header's "why".
- **Impact** finally calls /entities/{id}/blast-radius. The endpoint has existed
since the first API and had no frontend caller anywhere, despite
.agents/OIKOS.md naming blast radius as the reason the ontology exists. Its
outgoing-edges-only limitation is stated in the UI rather than hidden.
- **Activity merges four lists** (executions, signals, events, agent activity)
that were telling one story in four places.
- **Relations cap at 8 with a drill-in** — 540 interactive elements down to 126.
- **Ask Nomos** opens a task pre-scoped to what you are looking at, seeded with
the verdict just computed, via an optional draft threaded through
openNewTaskWindow -> NewTaskChat -> ChatThread.
Requires exposing check_defs.last_health/last_run_at through the API (the
columns landed with the health-aggregation work but were never surfaced).
Adding a fourth enum containing "unknown" made oapi-codegen disambiguate all
enum constants by type prefix, so metrics.go moves to gen.TrendDirection*.
Verdict derivation and type->section composition live in $lib/entityView.ts as
pure functions with 15 unit tests, including the host:strong case that
motivated this.
Co-Authored-By: Claude <noreply@anthropic.com>
|
|||
| 6ca6d5b352 |
fix(scheduler): derive entity health from all its checks, not the last one
host:strong logged 226 health.changed events in one hour, oscillating down/healthy while the host was fine throughout. host:hubris did it 126 times. runCheck wrote entity_status.health on every check completion, so an entity's health was simply whichever of its checks finished most recently. A host with six checks reported whichever facet happened to be sampled last, and one failing probe alternating with five passing ones flapped forever. resolveSignal forced "healthy" too, a second path by which one passing probe erased another probe's genuine failure. On this fleet the trigger is a known false positive: the scheduler's network vantage point cannot ICMP host:strong, so its ping check fails while every ssh-script check succeeds. Under last-writer-wins that single probe declared the whole host down, twice a minute. Each check now records its own verdict (check_defs.last_health, migration 027) and the entity's health is the worst across its enabled checks. A failing probe now degrades the entity honestly and *stably*, without erasing what the other five report, and health.changed fires only when that aggregate actually moves. Checks that have never run are ignored rather than counted as unknown, so adding a check cannot drag a known-good entity down before it has a verdict. Also declares service:oikos in the seed. The previous commit re-pointed the mcp ingress at it, but the entity only ever existed in the production database — so a fresh seed (a new install, or a DR restore) failed on an unresolvable edge. Caught by seeding an empty database rather than a copy of prod, which is the only way that class of bug shows up. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| af450dac2a |
fix(web): break the effect feedback loop, and stop serving HTML as JavaScript
Two unrelated console errors. effect_update_depth_exceeded — mine, from the previous commit. The live-update effects both read and wrote the same state: FleetMap's health patch builds a new `graph` object every run, and EntityDetailContent's refreshExecutions() assigns a fresh `executions` array. Svelte tracked those reads, so each write re-triggered the effect, which wrote again, until it gave up. The effects now depend on liveEvents alone and do their work inside untrack(). Applied to all four live effects, including the two that happened to settle on their own — relying on "applyHealthEvent returns the same reference when nothing changed" to break a feedback loop is far too subtle to leave implicit. SyntaxError: expected expression, got '<' — pre-existing, and unrelated to the live-update work. index.html loads /wails/runtime.js unconditionally; that file only exists inside the Wails desktop wrapper, which serves the same dist/ from its own asset handler. In a browser it is missing, and the SPA fallback answered it with index.html — so the browser parsed "<!doctype html>" as JavaScript on every single page load. The web Caddyfile now returns a real 404 for /wails/*, and more generally serves asset extensions without the SPA fallback: a missing .js or .css answered with HTML is always a confusing parse error rather than an honest 404. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 6ed9dc39e8 |
fix(compose): wait for the API to be healthy before starting nomos
nomos declared `depends_on: api: condition: service_started`, which only waits for the container to exist. It came up while the API was still binding :8090, failed its MCP initialize with "connection refused", exited 1, and crash-looped for ~25 seconds on every single deploy. It always recovered on its own, which is precisely why it went unnoticed. service_healthy waits for the API to answer, so this needs api to declare a healthcheck — wget is BusyBox's, already present in the alpine runtime image, so nothing new is installed. /healthz pings the database, so "healthy" means genuinely able to serve rather than merely listening. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| cc8eae4979 |
perf(web): patch health in place instead of refetching, and reconnect the SSE stream
Refetching everything on a health event was wasteful and churned the UI: one container going degraded pulled down the entire fleet entity list (plus its parent-grouping pass), or the whole fleet graph, to learn something the event had already delivered. health.changed / health.stale carry the new value in their payload, so the views that hold the entity just patch it: - Fleet table: patch the row. Only entity.* changes which entities exist, so only that still refetches. - Fleet map: patch the node AND graph.health[id] — healthOf() reads the side map in preference to the node's own field, so patching only the nodes would have left the rendered colour unchanged. - Entity detail: patch the open entity. Signals still need a read (the event says one was raised, not what the list now contains) but only the signals, not the entity and checks alongside them. Shared in $lib/health.ts, which returns the original array when an event does not apply so unrelated rows keep their identity and do not re-render. Note it matches on entity_id, never data.slug: the scheduler emits health.changed with entity_id = the observed entity but slug = the *check's* slug. Separately, events.ts had no reconnect. onerror was empty on the assumption the browser retries, but EventSource only does that for a transient failure -- once it reaches CLOSED (an HTTP error on connect, e.g. the API restarting during a deploy) it stays closed forever. A single blip silently froze every live surface in the app with nothing on screen to say so. Now reconnects with capped exponential backoff, and exports eventsConnected so a future indicator can show when the stream is down. Verified against live prod: flipping lxc:apps health recoloured the map node and moved its counts (30 healthy -> 29, 9 down -> 10) with ZERO network requests. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 4f706fa65f |
fix(web): keep health and status live everywhere they are shown
The SSE stream already carried health.changed, health.stale, signal.raised, signal.resolved and coverage.unmonitored, but two of the places that render health never listened for them. - The Fleet table refreshed only on entity.*, so its Health column sat at whatever it was when the page mounted while the map view beside it — which did listen — updated live. Health arrives on its own events, not entity.*. Coalesced on a 400ms timer because health.stale fires once per entity during a sweep, and refetching the whole fleet per event would mean a burst of identical requests. - The entity detail window loaded health, signals and monitoring once on open and never again, so a window left on screen kept showing the health it had at mount. That is the same staleness this whole change set has been about, reproduced one window at a time. Now scoped by entity_id, and re-reads only what a health or signal event can actually change rather than re-running the full 11-request load(). Verified against live prod: flipping lxc:apps healthy -> degraded -> healthy updated the Fleet table and an open detail window together, without a reload. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 50e899e5ee |
fix(checks): stop process_check.sh emitting invalid JSON, and mint one kind
`systemctl is-active` prints the state AND exits non-zero when a unit is not active, so `... || echo unknown` appended a second line: STATE became "inactive\nunknown" and the script emitted a raw newline inside a JSON string. The scheduler rejected all 14 process checks with "invalid character '\n' in string literal". Latent since the script was written — process checks never actually ran, because checkdefaults wrote an `args` config the ssh-script checker ignored. Passing args through finally executed them and exposed it. - head -1 keeps the state, and the fallback only fires on empty output. - Quotes are stripped from both the unit name and the state; either would break the hand-built JSON just as thoroughly. - signalKind is now the constant "process" rather than "$SERVICE". Emitting the service name minted a distinct signal kind per service (kind=paperless, kind=qbit, …) — nothing an approval_rule can match, and it makes "how many process checks are failing?" unanswerable. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 42751623ea |
fix(seeds): relax documents cardinality, re-point the mcp ingress
Rehearsing the deploy against a full copy of prod surfaced 40+ cardinality violations that would have failed the seed. Since api/scheduler/notifier all depend on `seed: service_completed_successfully`, and this change alters the seed files (so the content hash changes and a full re-ingest runs for the first time in months), that failure would have stopped those services from starting at all. None of them are new. The foreign-key bug in checkdefaults was aborting the ingest earlier, during entity ingest, so ValidateCardinality at the end never got the chance to run. Fixing the first failure revealed the next. - `documents` was declared many-to-one, meaning a document may document at most one entity. Nomos has been writing docs that cover several (a fleet-wide apt audit documents every host it touched) for months, which is reasonable — the ontology was the strict one. Now many-to-many. - The mcp ingress still routed to service:homelab-mcp, which prod marks deprecated: the Python MCP server on apps/105 was stopped at the Go cutover. Nomos re-pointed it at service:oikos on 2026-07-12 and was right; the seed was stale, and re-asserting the old edge alongside the new one is what made it a violation. Remaining after this: one genuine drift, `hosts target=lxc:caddy (2 edges)`, which needs a prod data fix rather than a code change — see the follow-up. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 98e19bb14a | Merge remote-tracking branch 'origin/main' into claude/coolify-oikos-comparison-d8c2d3 | |||
| d7b526a112 |
fix(scheduler): honour check_defs.interval_s, and renumber migrations off main
ListEnabledCheckDefs selected interval_s but never filtered on it, so every enabled check ran on every 30s pass and the declared per-check intervals were decorative. Invisible at 17 enabled checks; at ~180 it would have meant ~126 SSH connections every 30s (~363k/day) and `apt update` on every machine every 30 seconds — 14,400 mirror hits a day to answer a question that changes daily. - check_defs.last_run_at (migration 026) + a due-ness predicate in the query. A column rather than scheduler memory because this control plane restarts on every deploy, and an in-memory map would re-fire every check on each restart. - runCheck stamps last_run_at before processing the result, so a permanently failing check backs off to its interval instead of re-running every pass. - updates and backup-freshness drop to daily. Both answer questions whose answers change about once a day; 60s was just the shared ssh-script default. - last_run_at is seeded to a random offset within the interval so checks created by the same seed do not stay in lockstep — otherwise ~165 probes land in the same instant each minute instead of spread across it. Deliberately not in the upsert's DO UPDATE: a re-seed must not re-herd them. Steady state becomes ~180k SSH/day (down from ~363k) and 5 apt runs/day (down from 14,400), with each 60s check landing at its own point in the minute. Also renumbers 022→023, 023→024, 024→025: origin/main added its own 022_knowledge_revisions, and prod has already applied version 22. Left colliding, prod would have skipped the monitoring_spec migration entirely and then failed the seed on a missing column. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 1dca2cfd7a |
feat(observability): restore monitoring coverage, make gaps visible, stream executions
Monitoring coverage was 3 of 89 active entities. Three bugs, each hidden by discarded errors in checkdefaults: - writeCheck generated a fresh uuid, inserted the check entity ON CONFLICT (slug) DO NOTHING, then wrote a check_defs row referencing it. On any re-seed the slug already existed, the entity insert no-oped, and the FK violated — aborting the ingest transaction and surfacing as an unrelated failure several entities later. Re-seeding has been broken since; prod's coverage was frozen at its first successful seed. This is what TestSeedIngestIdempotentAndNoDuplicateEdges had been reporting. - shortSlug truncated to the last 8 chars, so all 21 ingress routes collapsed to ".network" and overwrote each other; service:jellyfin collided with lxc:jellyfin. - The ssh-script checker never read the `args` config checkdefaults wrote, so process_check.sh always ran without its unit name and returned "unknown". Coverage is now 75/89. Monitoring is declared per entity type in seeds/ontology.yaml and resolved through the is-a hierarchy, so a type can say it warrants nothing (site, lan, mesh, cluster) and never be reported as a gap. coverageSweep raises an `unmonitored` signal only where a type declares monitoring it lacks — 8 real gaps, no false positives. Also: - entity_types.attribute_schema was never ingested: the seed loader read "attribute_schema" but the YAML says "attributes", so all 60 types stored JSON null. - ListExecutions ignored its declared target/action/correlation_id filters and paginated on a non-unique target slug, dropping and repeating rows. - started_at was captured but only written at terminal state, so a running execution reported NULL for its whole life. The three MCP auto-run copies wrote no timing at all; they are now one autoRun helper. - SSH output was buffered to completion and discarded entirely on timeout. Both sshExec copies now stream through a shared execlog sink into execution_logs, and keep partial output when a command is cancelled. - executions.correlation_id was a random per-execution uuid that correlated nothing; it is now the chat session id, which is what lets the chat tail live output. - reversible_low had no auto-run branch despite policy declaring it unattended. Since computeCommandRisk never returns it, the class only arises when an agent declares it over a read_only command — so gating it penalised candor without adding safety. - backup-target gains a backup-freshness checker (portable find -mmin, since the first target is on macOS), resolving its host by walking backs-up-to backwards. The pre-deploy pg_dump is now a tracked backup target. UI: an Executions section on entity detail with live output tailing, and streamed output under a running `run` call in the chat timeline. Migrations 022-024. Ops.svelte and context.ts exclude execution.output from their refetch triggers, which would otherwise fire once a second per command. Co-Authored-By: Claude <noreply@anthropic.com> |
|||
| 7e1ccad5f4 |
fix(web): replace generic spinners with content-shaped loading skeletons
The six loading states across the Knowledge wiki (initial app load, the reader's note/history fetches, and the four Cleanup tabs) all showed a centered spinner with no relation to what was about to render — costing a full reflow the instant real content landed. Replaces each with a skeleton shaped like its actual content (tree rows, reader header + prose, revision list + diff, cluster cards, table rows, flat lists) using the existing shadcn Skeleton primitive already used elsewhere. Verified each of the six by temporarily injecting a delay into fetchWithAuth and screenshotting the transient state. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 89312a9ce4 |
feat(web): redesign Knowledge as an editable wiki
Replaces the read-only stats dashboard with a three-pane wiki: a navigator tree (group by folder/type/tag/entity), a reader/editor with bare-slug auto-linking and revision history + diff, and a context rail for backlinks and related notes. Adds a Cleanup mode for the drift tools (duplicates, tag manager, orphans, trash) and a Cmd+K quick-open. Also: - Adds a real landing view (hero count, KPI row, Nomos-share meter, recently-updated, busiest tags) in place of the old "Select a note" empty state, and extends the design pass across the tree/reader/rail (kind icons instead of repeated text badges, accent-bar selection, constrained prose measure). - Guards every note-selection path behind a confirm when there's an unsaved edit in progress, so switching notes can no longer silently discard a draft. - Extracts the markdown-rendering CSS duplicated across ChatThread, EntityDetailContent, and the new WikiReader into a shared .markdown-body class in app.css, with ChatThread keeping only its decorative deltas. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| ce0e4142ff |
feat(api): add knowledge base write path, revision history, and drift tooling
The Knowledge page was read-only from the HTTP API — the only writer was the agent's MCP upsert_knowledge tool. Adds create/update/soft-delete/ restore/trash endpoints, a DB-trigger-backed revision history (catches both the web UI and the MCP tool), and maintenance endpoints: duplicate detection (pg_trgm + complete-linkage clustering), tag rename/normalize, orphan detection, and merge. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 873b00ac42 |
style(web): fix prettier config, format entire web/ tree
.prettierrc.json was missing "semi": false, so prettier wanted to add semicolons to a codebase written without them (763 semicolon-free statements vs. 150 with, in hand-written .ts; zero hand-written .svelte files use them at all). That's why prettier --check failed on 249 files — not because the code was unformatted, but because the config didn't match the actual house style. Added "semi": false; left printWidth/etc as configured (printWidth barely moves the failure count: 218/213/212 files at 100/120/140). Ran `prettier --write .` with the corrected config. Verified semantics-preserving before and after: - eslint: 142 problems both before and after, byte-identical - build passes, 38/38 tests pass - token-stream diff (whitespace/semicolons/quotes normalized) on all 218 changed files: only 52 had any remaining token change, all either trailing-comma removal (matching trailingComma: "none") or import/ ternary reflow — no semantic changes - live smoke test: Knowledge, Tasks, Fleet map, and a chat window (AgentTrace, markdown, Scope graph, activity rail) all render correctly, no console errors Most of the diff is shadcn/ui vendor files (lib/components/ui/) moving from the CLI's own style (double quotes, tabs, semicolons) to house style; re-running `shadcn-svelte add` on a component will need a follow-up format pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
|||
| b345783eef |
fix(web): resolve two eslint errors in Knowledge.svelte
- drop the unused KnowledgeItem type import
- the svelte/no-at-html-tags disable comment sat on the wrong line: the
multi-line Card.Description opening tag meant "next line" wasn't the
line with {@html}, so it never suppressed. Reformatted so the {@html}
is on its own line, directly after the disable comment. Sanitization
(DOMPurify with ALLOWED_TAGS: ['b']) is unchanged — this was a false
positive, verified live (search results still render <b> highlights,
no script/attribute injection).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
|||
| 29d5cb8b85 |
fix(web): drop the taskbar theme label, align window padding to p-2
- Taskbar: the theme toggle is icon-only now, matching the Settings button beside it. The theme name moves into the title/aria-label so an icon-only control still has an accessible name and the current theme stays discoverable on hover. - Pages: Fleet, Knowledge, Learning, Ops and Signals used p-4 (or p-4 md:p-6) while Tasks used p-2, so windows didn't line up. All now p-2. App Store and Settings are deliberately untouched — they have no root padding, using a full-bleed header whose border spans the window; insetting them would break that. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
|||
| 8f440c5ad5 |
feat(web): collapse chat tool calls into one agent trace, reverse the activity rail
The chat rendered one card per tool call, so a 20-call turn buried the answer under 20 stacked cards. Merge them with the "thinking" indicator into a single collapsible strip above the answer: - collapsed: the live activity while running, a count once finished - expanded: the turn's work in humanized language (reuses the activity log's toolActivityLabel, so ten identical "run · target: host:strong" rows now read as what they actually did) - per row: the raw args/result, one more click in Also flip the Activity rail to newest-first with the current step on top: - follow-mode/auto-scroll re-anchored to the top to match, or it would jump to the oldest entry on every new event - pending plan steps park at the tail rather than sorting above the running step and pushing it off the top; the goal anchors the bottom Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
|||
| 4e4e2c169c |
feat(web): replace 3D graph with fleet map, add desktop background patterns, rename Knowledge Base to Fleet
- FleetMap: service-centric host -> container -> service graph replacing the WebGL 3D force graph, with health coloring, hover-to-trace blast radius, and click-to-open - Desktop background: configurable CSS pattern picker in Settings -> Appearance (8 patterns, color/fill/opacity/fade/size/rotation), replacing the hardcoded ambient graph background - Fix missing data-orientation/data-disabled Tailwind custom variants so the shadcn Slider's track actually renders - Rename "Knowledge Base" app to "Fleet"; scope its table to the same fleet entities as the graph (compute-entity descendants + service) instead of all entities - Remove dead code: EntityGraph, GraphBackground, categories.ts, MultiSelectFilter (all superseded by the above) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
|||
| c151a66627 |
fix(web): window/table polish — opaque windows, sortable columns, sticky-header scrollbar, viewport-clamped windows
- Windows: make floating windows fully opaque (drop backdrop-blur/color-mix
transparency), add margin around windows, reduce Overview padding to p-2.
- DataTable: fix Toolbar always rendering an empty padded bar (children
slot was always truthy regardless of actual content); split header/body
into separate tables so the scrollbar no longer overlaps the sticky
header; make sort work for derived/synthetic columns by sorting on the
column's accessor instead of a nonexistent row key.
- Overview: enable sorting on Status and Task columns via accessors.
- TaskContextPanel: give the Activity pane more height by default (Scope
30% / Activity 70%), fixing that the saved split sizes were never
actually applied to the bound Pane sizes.
- windows.ts: clamp new/resized windows to the desktop viewport so
content-heavy entity windows can't grow taller than the visible screen;
fixes a bad defaultSize.height ('30vh', an invalid non-numeric value)
that had silently left window height unconstrained.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
|||
| 1c12d40712 |
feat(web): adopt shadcn context-menu for mascot + desktop right-click menus
Problem: the mascot's right-click menu was non-interactive — RadialMenu's root div rendered inside MascotLayer's pointer-events-none root (and the new DockedLayer wrapper compounded it) without re-enabling pointer-events, so clicks passed straight through. The desktop right-click menu was a hand-rolled positioned div, inconsistent with the rest of the UI. Change: both menus now use the shadcn-svelte context-menu primitive (bits-ui, portaled to <body>). - Mascot: MascotMenu.svelte renders the action tree recursively — children become ContextMenu.Sub (native hover sub-menu navigation, replacing the manual breadcrumb stack), leaves become ContextMenu.Item with onSelect. MascotLayer wraps <Mascot> in a ContextMenu.Trigger; visibility predicates read reactively off ctx.model so items appear/disappear live. Removed the manual menuPos/openMenu/closeMenu machinery. RadialMenu.svelte deleted. - Desktop: the surface's bare-desktop hit area is now a ContextMenu.Trigger layer (absolute inset-0, pointer-events-auto) placed before the icons/windows in the DOM. The DOM-structure gate (icons/windows are pointer-events-auto siblings that paint on top and intercept their own right-clicks; bare desktop falls through to the trigger) replaces the old fragile e.currentTarget === e.target check. Left-click blur moved onto the trigger; Undo/Redo disabled state snapshotted via onOpenChange (canUndo/canRedo are wmkit methods). Risk: the blocker that made the mascot menu non-interactive in the first place — Mascot.svelte's handleContextMenu called e.stopPropagation(), which would have prevented a ContextMenu.Trigger wrapper from ever seeing the right-click. Removed that handler; bits-ui now owns right-click on the mascot, left-click drag/pet passes through. The context-menu content portals to <body>, escaping the pointer-events-none mascot and docked layers entirely — the structural fix, not just a component swap. Verification: vitest 38/38; svelte-check + tsc clean for changed files; eslint clean (the shadcn-generated ui/context-menu/* files carry the same baseline custom_element_props_identifier warnings as the rest of the ui/ folder, not from this change); vite build green; runtime confirmed — right-click mascot opens the action tree with hover sub-menus, right-click bare desktop opens Cascade/Tile/Show/Reset/ Undo/Redo, right-click on an icon or window does not. |
|||
| 482c7f3448 |
feat(web): app-registry architecture — OS + Apps, lazy loading, installable apps
Problem: the frontend had an implicit OS+Apps metaphor (desktop, floating windows, an app registry) but the contract was informal — the mascot was hardcoded into the shell, all apps were statically imported into one 800KB bundle, and there was no install/uninstall path. Change: three phases landed. - Phase 1 (contract + docked kind): AppDef extended with docked/noIcon and optional geometry; the mascot registered as a docked app via a generic DockedLayer that replaces the hardcoded <MascotLayer />; openAppWindow branches on docked → toggleDocked; persisted docked visibility store (absent key = visible, no APPS import to avoid a static cycle). - Phase 2 (lazy loading): AppDef.component is now a dynamic-import loader; LazyApp renders with a loading skeleton; Vite code-splits each app (main bundle 800KB→485KB); the LazyMascot wrapper is gone since the lazy loader breaks the import cycle directly. - Phase 3 (installable apps, local bundles): AppManifest + catalog + installApp/uninstallApp + localStorage persistence; reactive apps store (built-in + installed) and derived appById; App Store page; Notes demo app; icons.ts and WindowLayer's orphan-close react to registration so installs appear without a reload. - Structure: data-table casing unified to PascalCase; the mislabeled DataTable.svelte.ts (pure types, not runes) renamed to types.ts; LazyApp colocated with its desktop-shell consumers; app-store moved under lib/ so the dependency direction is consistent. Risk: the app registry is now a reactive store, not a static array, so every consumer (Desktop, DockedLayer, Taskbar, icons, windows) reads from derived stores. Two static-cycle traps are documented in docs/mbse/components.md §9: docked.ts must not import APPS (it would fire a TDZ at init via the apps.ts→pages→windows.ts→here path), and apps.ts must not statically import the mascot (the lazy loader defers its module graph). Remote bundle loading, the /api/v1/apps endpoint, and permission enforcement are deliberately NOT in this commit — they are security-critical and deferred to Phase 4 with an ADR. Verification: vitest 38/38; svelte-check + tsc clean for changed files; eslint clean; vite build green; runtime smoke confirmed (install Notes → icon appears → open → uninstall → icon + window gone; survives reload). docs/mbse/components.md Component 9 and the plan updated. Plan: plans/2026-07-21-frontend-os-apps-architecture.md |
|||
| 50aed11cc4 |
feat(web): adopt @vincjo/datatables for all tables, standardize shared components
- Add DataTable.svelte: declarative columns, built-in sorting, sticky headers, text truncation, column alignment, configurable widths, optional pagination/search - 12 built-in renderers: BadgeRenderer, StatusBadgeRenderer (unified risk/severity/ execution/state/type variant mapping), HealthDotRenderer, RelativeTimeRenderer, DateRenderer, DurationRenderer, StatusDotRenderer, SignalActions, ApprovalActions, ActivityAction, ActivityCancel - Migrate Overview (task board), Signals, Ops (3 tables) to DataTable - Refactor EntityTable treegrid to use shared SortHeader, EmptyState, HealthDotRenderer - Create shared components: EmptyState, StatusBadge, FilterTabs - Clean up Knowledge.svelte: replace inline relTime() and typeVariant() with shared utils - Add width, align, truncate column props; table-fixed layout; rounded-xl borders - Bump version to 0.11.0 |
|||
| ccbf6a8aac |
fix(nomos): sessions blocked on an execution approval now show "Needs input" and never idle-close
A config_mutation/destructive run() queued for approval never touched agent_sessions.status — only ask_operator did that, setting awaiting_input. So a task blocked on an execution approval was indistinguishable from one still genuinely working: the frontend's "Needs input" bucket only checks status===awaiting_input (never lit up for these), and the idle-sweep safety net only excludes awaiting_input from its stale-task query, so after ~30 minutes idle it would nudge the agent and then auto-close the task with outcome=partial while the approval was still sitting there undecided. classifyAndGate now flips the session into awaiting_input the moment an execution is queued (internal/mcp/server.go), and DecideApproval flips it back to executing once the approval is approved, denied, or revoked (internal/httpapi/approvals.go) — mirroring askOperator / answerQuestion's existing pattern for session_questions. Both emit task.status so the board updates live. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| ef2956619f |
fix(web): pin the tasks table header while scrolling
sticky top-0 on each <th> (not the <thead> itself — more consistent sticky support across browsers for table headers) plus a background so scrolled rows don't show through underneath it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| dffe01fb02 |
fix(web): use the terracotta accent instead of green for done checkmarks
text-primary instead of text-success — keeps the "done" state on-brand with the rest of the UI (buttons, focus rings) rather than introducing a separate green that only really worked well on the dark theme. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 0f9e366ad5 |
fix(web): tool-call checkmarks nearly invisible on the terracotta (light) theme
text-success/50 and /60 washed out to almost nothing against the light theme's cream card background — full-opacity text-success still reads as a calm, muted green (not alarming) but is actually visible on both themes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 052230209c |
fix(web): task launcher textarea no longer grows while typing; less transparent windows on Firefox
The launcher's textarea inherited the base Textarea component's default field-sizing:content (auto-grow to fit typed content) — ChatThread's input already overrides this with field-sizing-fixed, but the desktop launcher never did, so the box would jump taller the moment you started typing. Also bumps the floating-window frosted-glass opacity from 70% to 85%: backdrop-filter's blur strength isn't consistent across engines, and Firefox blurs noticeably less than Chromium at the same radius, making the Chromium-tuned opacity look far too see-through there. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| e5a81241b7 |
feat(web): operator questions inline in chat, mascot reactions scoped to the focused task
The pending operator-question card now renders inline in ChatThread (the newest thing in the conversation) instead of in the context rail — it's part of the chat, not a separate side panel, and the panel's hasContext gate no longer needs to special-case it. The desktop mascot's reactions are now entirely about whichever task window has focus, not fleet-wide events: thinking/talking is a new continuous `busy` behavior that tracks the focused session's own streaming state (thinking before any text arrives, talking once it does — using the previously-unwired peep/talk sprite), eureka fires with the actual knowledge title that was recorded, happy fires with the task's own completion summary, and alarmed now means "this task needs your OK" (an operator question was raised) rather than a fleet-wide critical/signal event. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 6b6bfe1fd8 |
feat(web): new task opens straight into chat, context rail waits for content, frosted windows
New Task now opens directly as an empty ChatThread (NewTaskChat) instead of a separate compose screen, sized like a real task window. The Scope/Activity context rail in a task window no longer renders until there's actually something to show (touched entities, activity, or an open question), avoiding an empty-placeholder sidebar on every new task. Also fixes the chat input defaulting to several lines tall on window open, centers the empty-chat greeting vertically, and gives floating windows the same frosted-glass look as the desktop's task launcher card. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| ce34cfeac7 |
feat(web+nomos): fix chat streaming reactivity, unified activity timeline, tool cards in chat
- fix(web): Svelte 5 identity-based reactivity broke text_delta streaming — immutable message objects in all three chat handlers so text streams live - feat(web): streaming cursor + inline status indicator merged into message flow - feat(web): expandable inline tool call cards in chat thread - feat(web): merge Plan + Event log into one backbone Activity timeline — filled status nodes, branch stubs, auto-scroll follow mode, per-session activityLog, compact for the rail - fix(nomos): add X-Accel-Buffering:no to /chat SSE (proxy buffering) - fix(nomos): plan step auto-close SQL param bug (store.go) - polish: timestamps, role labels, code copy button, table overflow, min window size, delete AgentIndicator/ActivityTimeline dead code |
|||
| 55b93c59ef | fix(web): remove redundant GraphBackground from Tasks page | |||
| eb3d2de1ca | feat(web): mascot physics juice (bounce, skid, spring squash, hop) + typewriter speech bubble | |||
| d82095213a |
fix(web): mascot physics, drag reliability, and speech-bubble polish
Audits and fixes ground-teleport/flat-fall/toss-momentum physics bugs, fixes drag getting stuck via missing pointercancel handling, replaces sprite-based speech bubbles with real HTML text/emoji bubbles, adds drag-onto-icon "investigate" reactions and idle chatter, merges the name badge and reaction bubble into one floating element, and caps the bubble to one line with a teleprompter-style auto-scroll instead of ellipsizing overflow text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 7b1dfbc8aa |
feat(web): desktop mascot ("Cluck") — egg/chick/adult tamagotchi that roams the desktop, reacts to chat/events, walks on top of windows
Implements plans/2026-07-20-desktop-mascot.md. New code under web/src/lib/mascot/ (types/sprites/render/state/behavior/actions/ stimuli + Mascot/MascotLayer/RadialMenu/NameDialog components) plus CC0 sprite sheets at web/public/mascot/ (chicken + Onocentaur egg pack + reaction bubbles). MascotLayer is inserted into Desktop.svelte after WindowLayer; <2-line integration. Tamagotchi: egg -> chick -> adult lifecycle persisted to localStorage['oikos-mascot'] (debounced 300ms). Egg hatches on first naming (no timed incubation per implementation deviation). Chick/adult wander, peck, sleep, blink autonomously via a weighted-random FSM; the chicken walks above windows (ground line = highest window top edge beneath its x, recomputed each tick from wmState; rides the ground when the window beneath is dragged). Interaction: draggable with flutter-fall physics on release mid-air; plain click = pet (heart bubble + happy anim); right-click opens a rounded-button radial menu (Interact/Care/Identity/Debug nested groups) mirroring the desktop's own right-click menu styling; auto-flips above/ left near screen edges. Awareness: stimulus bus subscribes to chat.ts streaming, activity.ts activityLog (knowledge-entry diff), events.ts liveEvents (critical/ signal -> alarmed, execution -> happy), with priority+cooldown gating. Egg-stage reactions are suppressed. Reaction bubbles are anti-aliased. Sprite loop runs at ~60fps via setTimeout (not rAF) per GraphBackground convention, dt clamped to 100ms; position via transform: translate3d + will-change: transform for compositor-friendly motion. Z-index ordering: WindowLayer z-40 < MascotLayer z-[45] < desktop context menu z-50 < RadialMenu/NameDialog z-[60]. Docs: plan + docs/mascot/README.md (MBSE subsystem model) updated to Implemented with a deviations note covering hatch-on-naming, PNG-sheet art, button-column radial menu, 60fps loop, egg-reaction suppression, and window-walking ground model. VERSION bumped 0.7.13 -> 0.8.0. |
|||
| f1cdf4ea13 | chore: trigger webhook redelivery (verify ALLOWED_HOST_LIST fix) | |||
| e055a7c6ce |
feat(nomos): session-review improvements (P0/P1/P2 from 2026-07-20 audit)
Classifier now unwraps pct exec / qm guest exec / bash -c / sh -c / sudo
and env-var assignments before classification, so read-only inspection
wrapped in pct exec no longer escalates to config_mutation. curl GET
(default method, no -d/-F/-T/-o/>) is read-only. Eliminates the three
duplicate rclone sessions (a51e2086, 8acea2e3, cb8c8a4a) that bounced
off the classifier for the same goal.
New classify_command MCP tool: command-scoped preflight that returns the
exact risk class run would assign. Documented in SOUL.md with guidance
to pre-classify before run when the verdict is uncertain.
set_goal surfaces prior partial/failed sessions from the last 24h so the
agent picks up the thread instead of rediscovering it.
completeTask auto-closes in-flight plan steps (pending/running -> done
on success, skipped on partial/failure), so one-step plans no longer
need the per-step running->done dance right before completion.
Migration 021 adds blocker + closed_at to agent_sessions. completeTask
sets closed_at once and derives a structured blocker reason
(approval_timeout, user_abandoned, classifier_overreach, model_refusal,
tool_error, ...) from the last assistant message.
/sessions list now carries message_count, tool_call_count,
duration_seconds (server-side aggregates — no more N+1 transcript
fetches to audit a fleet). GET /sessions/{id} returns both metadata
and messages. New query params filter + paginate: outcome, status,
entity_id, blocker, since (RFC3339 or Go duration), cursor, limit.
Titles now prefer the goal when set; sessions without a goal fall back
to the first assistant text.
New GET /sessions/{id}/tool_calls flat view for audit scripts.
Plan: plans/2026-07-20-session-review-ten-sessions.md. VERSION 0.7.12 -> 0.7.13.
|
|||
| 9f4d645d06 |
docs: plan a pixel-art desktop mascot ("Cluck")
Design-only (no code yet): an MBSE subsystem model for a chicken mascot that roams the desktop shell, is draggable, opens a Sims-style nested radial menu, and has a tamagotchi lifecycle (egg -> chick -> adult) that reacts to real app activity (chat streaming, knowledge-graph writes, signals). Everything (animations, autonomous behaviors, menu actions, environment reactions) is scoped as a data-driven registry for easy extension. - docs/mascot/README.md: subsystem Model conforming to docs/mbse's Holt-based Framework — mission/boundary, requirements, structural view (module registry map), behavioral view (behavior FSM + lifecycle state machines + a stimulus sequence diagram), interfaces view (which web stores it observes, read-only), extension guide, verification view. - plans/2026-07-20-desktop-mascot.md: the concrete file-by-file implementation plan for web/src/lib/mascot/ derived from the model, with an ordered build sequence and a manual browser verification checklist. - Indexed both in docs/index.md and plans/index.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| aee458ce83 |
feat(web): add windowed Settings app, separate from initial Config screen
The taskbar's gear icon reopened the full-page "Connect to Oikos" screen even once already connected. Split that: Config.svelte stays as the first-run/unconfigured screen; a new Settings app (windowed, like Tasks or Operations) now handles in-session changes, with a section list (Connection, Appearance) built to grow — future settings are one more entry, not a new screen. - pages/Settings.svelte: Connection (server URL/token/Authentik, reusing config.ts + oidc.ts) and Appearance (Terracotta/Carbon picker) sections. - apps.ts: registered as a normal desktop app. - Taskbar's gear button now opens the Settings window; removed the onOpenConnection prop threaded through App -> Desktop -> Taskbar, since Settings' "Forget saved connection" (clear config + reload) replaces it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 58a11ca872 |
feat(web): resizable panels via svelte-splitpanes + Claude-style composer
Replace hand-rolled pointer-resize logic (TaskContextPanel's 3-way vertical split, SessionChatWindow's rail, ChatThread's message/input split) with svelte-splitpanes, themed onto the app's existing border/primary tokens. - TaskContextPanel: Scope/Plan/Event-log sections collapse to a fixed header height and restore their last size on reopen. - ChatThread: input area is now a separate resizable pane, clamped to a measured one-line minimum and a 45% max, instead of a fixed max-h textarea. - Send button restyled to sit inside the input's corner (Claude-style), swapping the up-arrow for a corner-down-left return icon. - Adds a $app/environment shim + optimizeDeps exclude, since svelte-splitpanes assumes SvelteKit and this is a plain Vite app. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| aed068de12 |
feat(web): redesign UI as an OS-style desktop shell
Replace the sidebar + hash-routed page shell with a desktop metaphor:
draggable app icons, apps opening as floating wmkit windows, a centered
"What should Nomos do?" task launcher, and a bottom taskbar showing all
open windows plus a system tray.
- New app registry ($lib/apps.ts) — adding an app is one entry, nothing
else to touch.
- New desktop shell components (Desktop, WindowLayer, DesktopIcon,
Taskbar, TaskLauncher) under $lib/components/desktop-shell/.
- Icon positions are a persisted, collision-avoiding grid ($lib/stores/icons.ts).
- Window layout persists across reloads (wmkit persist), with
drag-to-maximize, F6 window cycling, and now a right-click desktop menu
(cascade/tile/show desktop/reset icons) plus Cmd/Ctrl+Z undo/redo for
window moves, resizes, and closes.
- Taskbar buttons get a hover-close and self-correct their title once a
new task's real goal is known.
- New task windows (desktop launcher and the Tasks app's "New task"
button) open as a window, not a dialog, and hand off to the real
session window once the backend assigns an id.
- Fixed a real gap along the way: GET /sessions/{id} couldn't tell
"session deleted" from "session has no messages yet" (both returned
200 with an empty list) — cmd/nomos/main.go now checks existence and
404s, so a stale/persisted task window shows "Task not found" instead
of a misleadingly empty, live-looking chat.
- Test coverage for the new pure logic (icon placement/collision
avoidance, app registry id helpers) plus a vitest matchMedia polyfill
needed to import anything touching the theme store.
Deletes the now-superseded sidebar shell, MinimizedWindowsBar, and the
standalone Chat/EntityDetail pages (folded into the window layer).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
|||
| 8657ac5669 |
feat(web): open tasks/sessions as floating windows with independent live chat
Clicking a task now opens it as a wmkit floating window (like entity windows already do) instead of navigating away from wherever you were. Several task windows can be open and actively streaming at once, each fully independent — no "which one's on screen" guard needed, since each window owns its own store bundle: - chat.ts: chatFor(sessionId)/loadSessionChat/sendSessionMessage give each window its own messages/streaming/connectionState, alongside the existing singleton path the main Chat page still uses unchanged. - workspace.ts: same split for plan/questions/touched/health-diffs (workspaceFor/startSessionWorkspace), each with its own live-event watermark since several windows can watch the same event stream. - activity.ts: activityLogFor(sessionId) mirrors the global derivation. SessionGraph.svelte, OperatorQuestion.svelte, and ActivityTimeline.svelte were converted from store-importing to prop-driven (matching the new ChatThread.svelte, extracted from Chat.svelte's transcript/input so both the main page and task windows share one implementation instead of duplicating markup/styling) so each can render either the global "current session" or a specific window's session. Also: minimized-window taskbar chips now cap at a max width with middle-ellipsis truncation instead of growing unbounded, and the window header's title/action-button row is fixed to genuinely match heights (not just share a center point) for more robust alignment. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| e28e0e9ea3 |
feat(web): unify Knowledge Base filtering into one type multiselect
Replace the Fleet/Network/Identity/Knowledge category tabs (which scoped entity fetches server-side) with a single "Types" multiselect shared by both the table and graph views — both now fetch the whole entity set (paginated via the new fetchAllEntities) and filter client-side, defaulting to fleet's types. Table and graph also share one search/highlight field instead of two separately-labeled ones. Along the way, fixed a real bug the wider entity set exposed: the treegrid's parent/child grouping fired one fetchGraph call per candidate root entity, fine for the old ~50-entity fleet scope but an ERR_INSUFFICIENT_RESOURCES flood once scoped to the full ~1700-entity set. Replaced with a single whole-graph fetch, deriving parent/child pairs from its edges client-side. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 6051fb4845 |
feat(web): curved edges + unique SVG ids for concurrent graph views
Quadratic-bezier edges instead of straight lines, and drop the auto-refit-on-load that caused a jarring zoom/pan snap once the force simulation settled. Also namespace each graph's dot-grid pattern id with a per-instance uuid — multiple SessionGraph instances can now be mounted at once (one per open task window), and duplicate SVG ids silently blanked out every graph's background but the first. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
|||
| 544afae77f |
feat(nomos): retry cap, vm: targets, inspect_path, goal supersession, runbooks
Session-review implementation for the three sessions audited in
plans/2026-07-18-session-review-three-sessions.md. v0.7.11 → v0.7.12.
P0.1 — retry cap + investigate-before-retry (cmd/nomos/retrycap.go,
agent.go): after 3 identical failing run calls in a single turn, refuse
to dispatch the call again and return a directive to investigate *why*
(ps/strace/lsof) or surface the blocker. Per-turn scope so a fresh turn
after the operator responds can retry once more. Session 1e9c7691's 20+
identical chown retries (knfsd held a kernel lock on the exported NFS
dir) is the direct motivation.
P0.2 + P1.8 + P2.10 — SOUL.md guidance: hung command is not a failed
command (investigate before retry); ask before proposing a multi-step
migration; multi-goal sessions summarize the arc not just the last goal.
P1.3 — two new runbook entities in seeds/knowledge.yaml:
- nfs-exported-dir-mutation-hang (the knfsd fchownat lock procedure:
killall → exportfs -u → mutate → exportfs -a → verify)
- netbird-mgmt-oidc-race-after-upgrade (docker restart netbird-mgmt
after ~30s for the traefik/authentik OIDC race)
P1.4 — setGoal emits task.superseded event when prior goal is overwritten
by a different goal (store.go, TestSetGoal_SupersededEvent). Session
55927f0a had two set_goal calls with the first silently abandoned.
P1.5 — inspect_path MCP tool: runs mount/df/ls/stat for one path across
up to 8 targets in one parallel call, replacing the 15+ run-call
fact-gathering fan-out sessions 1 and 2 each spent on cross-target path
tracing (tools.go, server.go: inspectPathAcrossTargets, inspectOneTarget).
P1.6 — vm: target support in run via qm guest exec (no more SSH-hop
with nested quoting). Extracted shared resolveProxmoxHostSlug for
LXC + VM, with hosts-relationship fallback when attributes.host is
absent (server.go, tools.go). Session 55927f0a's SSH-hop workarounds
for vm:zimaos are the direct motivation.
Deferred (documented in plan): P1.7 (approval window auto-extend on
timeout) and P2.9 (long-running command PENDING detection) — both
addressed at lower cost by the retry cap. Session 3's poll-after-timeout
pattern already works; the cap protects against the failure mode.
|
|||
| bd44626532 |
feat(web): floating entity-detail windows (wmkit), replacing sidebar/sheet
Every place that showed entity detail (Knowledge Base's right sidebar, the EntitySheet drawer used by Knowledge and the chat session graph, the standalone /entity/:slug page) now opens the entity in its own floating, draggable, resizable window instead — several can be open side by side, and clicking a relation inside one opens another, building up a stack. Windows are managed by one global wmkit instance (new $lib/stores/windows.ts + $lib/components/EntityDesktop.svelte, mounted once in App.svelte), themed with the app's own card/border/ring tokens rather than wmkit's bundled themes (app.css). - Delete EntitySheet.svelte (redundant) and the KnowledgeBase resizable detail pane; row/graph-node click handlers now call openEntityWindow(slug) instead of setting local sidebar state. - SessionGraph (chat's "Scope" mini-graph): clicking a node opens its window directly instead of a click-through mini-detail panel with its own resize handle and "Full detail" button — that whole subsystem is now dead and removed. Node highlight ring is kept (still useful to see what you last opened) and now clears itself via an effect watching the shared window-manager store, so closing a window drops the highlight instead of leaving it pointing at nothing — same fix applied to Knowledge Base's row highlight. - Compact the entity-detail panel's padding (container + each DetailSection) now that it's typically viewed in a small window rather than a full-height sidebar. - Fix KnowledgeBase's browse pane losing its flex-1/min-w-0 (and thus full width) when the wrapping single-child div around it was removed along with the old detail-pane split. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |