docs: fix plan/repo drift, retire dead Goose+Nomos and Caveman tooling
Documentation and repo-hygiene pass following the client/server split:
Plan drift (audited all other active plans against current code):
- oikos-gaps-and-improvements.md: mark Section C and D.5 resolved (both
described cmd/hermes, renamed to cmd/nomos with a real LLM loop since);
refresh ~10 stale file:line citations; fix tool-count (33, not 28).
- liveness-drift-and-ux-cohesion.md: fix stale default-model claim (now
deepseek-v4-pro since 2026-07-10) and "not yet deployed" status.
- nomos-agent-code-review.md: fix C1's citation (one unauthenticated route
to nomos now, not two, after the client/server split).
- wails-desktop-app.md: record the production deploy outcome.
Repo structure: added missing directories to README/CONTRIBUTING layout
tables (checks/, tools/, cmd/webhook/, docs/operations/), fixed a broken
link, added ADR 0015 documenting the auth/CORS/client-split model (there
wasn't one despite CONTRIBUTING's own process requiring it), normalized
ADR 0013/0014's format drift, added an Authentication section to
AGENTS.md/CLIENTS.md (every example call was missing the now-required
bearer header).
Retired the Goose+Nomos workstation flow (bootstrap.sh --with-nomos,
tools/setup-nomos-soul.sh, .agents/operations/nomos-agent.md) and the
Caveman auto-install tooling (tools/setup-caveman.sh, tools/caveman/) —
both superseded by the production containerized Nomos agent, which has
never used either. Kept .agents/shared/caveman.md itself (the terse
writing-style convention agents still follow by reading it).
Deleted the orphaned legacy Python oikos/ directory — nothing imports it,
and bin/homelab (the CLI it was kept for) no longer exists in the repo.
Rewrote .agents/operations/agent-enrollment.md (365 -> ~110 lines) and
commands.md to match the current architecture instead of the retired
`homelab` CLI; migrated the still-true networking prerequisites (Netbird,
split-horizon DNS, SSH key distribution) into the knowledge base as a
runbook via upsert_knowledge rather than duplicating them in markdown.
Updated all 10 .agents/skills/ runbooks referencing the dead CLI with
their real MCP tool / REST API equivalents, or flagged them as needing
verification where no equivalent is confirmed yet.
Two real bugs found and fixed, not just docs:
- The tools/setup-*.sh auto-setup glob was tools/*.setup.sh in THREE
places (tools/post-pull.sh, bootstrap.sh, and internal/httpapi/impl.go's
GetClientContext handler) since the mechanism's introduction on
2026-06-02 — never matched any real filename, so no client has ever
picked up an auto-setup script via git-pull or the context-poller sync.
Fixed all three; the Go server-side fix is the one that actually matters
since it's what the current context-poller mechanism depends on.
- bootstrap.sh removed dead vestigial --gitea-token/--gitea-user flags
(parsed, never consumed) left over from an earlier clone-based model.
Also flagged, not fixed (documented as an open gap in
client-enrollment/SKILL.md): bootstrap.sh tells a freshly-enrolled client
to call POST /api/v1/clients/{slug}/activate to finish enrollment, but
that route doesn't exist in api/openapi.yaml — EnrollClient sets entities
to provisioning and nothing currently transitions them to active.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,10 +1,33 @@
|
||||
# 2026-07-12 — Wails desktop application
|
||||
|
||||
**Status:** In Progress — Phase 0 (0.1-0.4, 0.6) done and verified live
|
||||
(browser: cross-origin static SPA + API on different ports, CORS, bearer
|
||||
auth, SSE query-token auth, localStorage persistence across reload — see
|
||||
"Plan review" for the gaps found and fixed along the way). Phase 1 (Wails
|
||||
shell) not started.
|
||||
**Status:** In Progress — Phase 0 (0.1-0.4, 0.6) done, verified live in a
|
||||
local browser test, and **deployed to production** (mac-mini, commit
|
||||
`0c0f35a`, 2026-07-12). Phase 1 (Wails shell) not started.
|
||||
|
||||
**Production deploy (2026-07-12):** merged to `main`, picked up by the
|
||||
2-minute deploy poller (`scripts/deploy.sh`: pg_dump backup → rebuild →
|
||||
rolling restart → health check), `healthy after 1s`. Verified post-deploy:
|
||||
unauthenticated `/api/v1/*` now 401s (the dev-open bypass was live in
|
||||
production before this — `OIKOS_ENV=dev` with no token set — so this closed
|
||||
a real, currently-exploitable hole, not just future prep); `/healthz` stayed
|
||||
open; nomos reconnected its MCP session with the new
|
||||
`OIKOS_MCP_BEARER_TOKEN` and a real tool call round-tripped end to end
|
||||
(`get_health_summary` via `/query`). A real random token was generated and
|
||||
added to mac-mini's `.env` (not committed — gitignored) before deploy, so
|
||||
the `${OIKOS_MCP_BEARER_TOKEN:-dev-token}` fallback in `docker-compose.yml`
|
||||
never activated with the weak literal default.
|
||||
|
||||
**Deliberately not done as part of this deploy** (out of scope — a different
|
||||
host/repo than "mac-mini", not touched): the Caddy LXC (121) and
|
||||
`dtoro/caddy-conf`. Checked the real production Caddyfile directly — there is
|
||||
**no `oikos.hubris.network` site block at all yet**, so the Authentik-bypass
|
||||
risk (gap 1 below) doesn't apply yet; there's no public UI exposed to break.
|
||||
`mcp.hubris.network` exists but still reverse-proxies to the old
|
||||
pre-consolidation service on LXC 105 (`192.168.8.205:9810`), unrelated to
|
||||
this stack — stale, but pre-existing and out of scope here. Exposing
|
||||
`oikos.hubris.network` publicly (with the `@api` bypass this plan's
|
||||
Caddyfile.oikos reference copy already has) is unstarted follow-up work, not
|
||||
a regression from this deploy.
|
||||
|
||||
## Plan review — gaps found before starting Phase 0
|
||||
|
||||
|
||||
Reference in New Issue
Block a user