dtoro c5ffaec85b fix(agent): panic recovery on every background goroutine (B1+B2)
Fixes B1 and B2 of plans/2026-07-11-nomos-agent-code-review.md together,
since the right granularity for B1 in the auto-continuation worker turned
out to require B2's restructuring anyway (see below).

B1: grep -rn "recover()" cmd/nomos/ internal/mcp/ internal/httpapi/ returned
nothing before this — every explicitly-spawned goroutine (continuation
worker, resumed chat turns, async execution dispatch, the SSE listener, two
duplicate sshExec implementations' output-collector goroutines) crashed the
whole process on an unhandled panic, not just that one goroutine. More
consequential post-concurrency: more simultaneous unattended background work
means more surface area for one bad input to end every running task.

New internal/safego package: Go(label, fn) launches fn in a goroutine with a
recover-and-log wrapper. Applied at every bare `go` spawn site across the
three packages. Two sites needed bespoke handling instead of the generic
helper because their callers block on a channel and a silent recover would
just make them hang until timeout: sshExec's output-collector goroutine (two
near-identical copies, internal/mcp/server.go and internal/httpapi/phase3.go)
and httpapi's ListenAndServe goroutine — both now recover AND send a
synthetic error result so the waiting select unblocks immediately instead of
waiting out the full timeout.

httpapi's sseListener got extra treatment: its per-notification handling was
extracted into handleNotification with its own recover, so a panic decoding
ONE malformed pg_notify payload can't kill the listener goroutine for every
connected SSE client — the outer goroutine spawn only needs to guard the
connection setup/reconnect code around it.

B2: cmd/nomos/continue.go's processContinuations used to run every pending
continuation SEQUENTIALLY in a plain for loop, in the SAME goroutine as the
ticker — meaning (a) task B's continuation waited for task A's full (up to
10-minute) resumed turn to finish first, undercutting this session's earlier
concurrency work on exactly the path autonomous tasks depend on most, and
(b) an unrecovered panic anywhere in that call chain didn't just crash the
process (B1) — even WITH B1's recovery wrapped only at the top-level worker
spawn, the panic would still unwind the ENTIRE ticker-loop goroutine,
silently ending auto-continuation for every task until nomos restarted.
Fixed by spawning each pending item via safego.Go individually: real
parallelism, and a bad item can now only ever take down its own goroutine.

Added internal/safego/safego_test.go: TestGo_RecoversPanic is the concrete
proof — a deliberate panic inside Go() that would otherwise crash the whole
test binary; reaching the assertion after it IS the evidence recovery works.

Verified live against the rebuilt containers: full chat turn round-tripped
correctly (hostname lookup, 2 iterations, normal completion) — no regression
from threading safego.Go through the tool-dispatch/continuation paths.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 20:05:19 +02:00

Oikos

Agentic homelab operating system written in Go. Single binary (cmd/oikos), Docker-deployed on mac-mini, with a standalone Nomos MCP agent gateway (cmd/nomos). Manages the hubris Proxmox homelab autonomously — observes state, classifies actions against policy, executes approved procedures over SSH, learns from outcomes, and escalates when uncertain.

For agents running on enrolled clients: start with AGENTS.md. For client machines: see CLIENTS.md. For developers: see CONTRIBUTING.md.

Quick start

# Dev stack (postgres + api + scheduler + notifier)
docker compose --profile dev up -d

# Full stack (adds Nomos agent gateway)
docker compose --profile full up -d

# Build standalone binary
go build -o bin/oikos -tags timetzdata ./cmd/oikos

# Run all roles in one process (dev mode)
OIKOS_DATABASE_URL="postgres://oikos:oikos_dev@localhost:5432/oikos?sslmode=disable" \
  go run ./cmd/oikos all

Architecture

                  ┌──────────────────────────────────┐
                  │         mac-mini (Docker)         │
                  │                                   │
  Workstation ─── │  nomos (8092) ──MCP── api (8090) │
  (mesh)          │    MCP gateway      REST + MCP    │
                  │                                   │
                  │  scheduler ── notifier ── postgres │
                  │  (observe)    (Matrix)   (Timescale)│
                  └──────────────────────────────────┘
Component Port Role
oikos api 8090 REST API + MCP server (15 tools)
oikos scheduler Probe runner, signal lifecycle, metrics
oikos notifier Approval tokens, Matrix alerts
nomos serve 8092 MCP client gateway, query routing

Phases

Phase Status Description
1 — Ontology + DB TimescaleDB, migrations, seeds, blast_radius
2 — API OpenAPI-first REST + MCP, auth, SSE, audit
3 — Control loop Scheduler, actuator, learning, classifier, notifier
4 — Nomos agent Standalone MCP client gateway, agent activity
5 — Secrets Infisical backend + SOPS fallback, rotation runbooks
6 — Deploy CI pipeline, cutover checklist, watchdog, rollback

Full plan: plans/2026-07-06-consolidate-oikos-control-plane-onto-mac-mini.md.

Operations

API endpoints

curl http://localhost:8090/api/v1/entities?type=service  # fleet
curl http://localhost:8090/api/v1/health                  # fleet health
curl http://localhost:8090/api/v1/agent-activity          # agent log

Nomos queries

# Structured tool call
curl -X POST localhost:8092/query -H "Content-Type: application/json" \
  -d '{"tool":"get_blast_radius","args":{"entity_id":"service:authentik"}}'

# Natural language
curl -X POST localhost:8092/query -H "Content-Type: application/json" \
  -d '{"query":"what depends on authentik?"}'

CLI

oikos migrate     # apply DB migrations
oikos seed        # ingest ontology/inventory/policy seeds
oikos export      # export DB state to YAML
oikos api         # serve REST + MCP
oikos scheduler   # run observe loop
oikos notifier    # run notification loop
oikos all         # all roles in one process
oikos secret list # enumerate SOPS secrets
oikos secret migrate  # SOPS → Infisical

Repo layout

cmd/oikos/          Go entry point — single binary
cmd/nomos/          Nomos MCP client gateway
internal/           Go packages (httpapi, mcp, scheduler, actuator, learning,
                    notifier, policy, secrets, db, config, ontology, domain,
                    knowledge)
api/openapi.yaml    API contract (OpenAPI 3.1)
migrations/         Forward-only SQL migrations (TimescaleDB)
seeds/              Bootstrap YAML (ontology, inventory, policy, knowledge)
compose/            Dockerfiles + Caddy config
scripts/            Deploy, watchdog, verification, rollback
nomos/              Nomos config, persona, skills
.agents/            Agent instruction files, shared conventions, skills
archive/            Historical reference (legacy wiki, plans, SOPS backups)
plans/              Design documents (active + done)
docs/adr/           Architecture decision records

For agents

See AGENTS.md for the full orientation. Quick reference:

  • Source of truth: DB (runtime) then seeds (bootstrap). Old wiki is archived at archive/knowledge/ — use MCP search_knowledge instead.
  • Mutations: classify against policy, request approval for destructive/config_mutation
  • Secrets: Infisical (primary) or SOPS (fallback) — never hardcode
Description
Agentic OS for running a Homelab
Readme 37 MiB
Languages
Go 53.1%
Svelte 25.7%
TypeScript 14%
Shell 3.8%
Python 1.7%
Other 1.5%