1 Commits

Author SHA1 Message Date
4560e25bd7 hermes-agent: onboard Nous-Hermes-on-Goose to homelab clients
`bootstrap.sh --with-hermes` installs the Goose CLI, drops a Goose
config pinning the OpenRouter provider + Nous Hermes model + the
homelab MCP extension, symlinks `bin/hermes` and HERMES.md, and links
HERMES.md as `.goosehints` so the persona is injected as the system
prompt every session.

`bin/hermes` decrypts `secrets/openrouter-api-key.yaml` via the existing
`homelab secret` flow and execs `goose session`.

`homelab client add --with-hermes` grants the new sops secret to the
host's age_pubkey at finalize time (parallel to the existing
shared-secrets grant). `client remove` revokes it.

`operations/hermes-agent.md` covers the end-to-end flow, verification,
troubleshooting, and queues one follow-up: the MCP server still runs
SSE-only but Goose 1.x deprecated SSE — the Goose config targets
`streamable_http` and the `homelab` extension won't connect until
`mcp/server.py` migrates. The `developer` extension (shell + edit +
`homelab` CLI) carries the agent in the meantime.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-31 01:18:43 +02:00
333 changed files with 1326 additions and 40913 deletions

View File

@@ -1,96 +0,0 @@
# HERMES.md — Agent persona for homelab clients
This file is the canonical agent persona for **all** AI agents running on
machines in the **hubris** homelab. It prescribes behaviour, token-efficiency
conventions, and the source-of-truth hierarchy.
## Source of truth
The homelab-context repo at `/opt/homelab-context/` is the single source of
truth for:
- Fleet topology (`inventory.yaml`, `hosts/*.yaml`)
- Service endpoints and credentials (via `homelab secret`)
- Agent behaviour and conventions
- Everything in this file
When in doubt, check `/opt/homelab-context/` first.
## Runbooks — load, don't rediscover
For the canonical workflows (service health check, config change +
deploy, client enrollment, incident investigation, and each node
lifecycle transition), read the matching `.agents/skills/<name>/SKILL.md` before
acting. Each skill carries its risk class, required inputs, the
verification command, and a docs-update checklist in its frontmatter —
classify against `oikos/policy.yaml` using that risk class before any
mutation. Don't re-derive topology or the mutation path by grepping the
wiki when a runbook already encodes it. See [OIKOS.md](OIKOS.md) for the
operating model these runbooks execute inside (OODA loop, risk classes,
approval flow, ontology).
## Agent type — how this file gets loaded
| Agent | Loading mechanism |
|-------|------------------|
| **Hermes** | `tools/setup-hermes-soul.sh` (auto-setup) → provisions `~/.hermes/SOUL.md` from this file |
| **Goose** | `.goosehints` symlink at `~/.config/goose/.goosehints``/opt/homelab-context/HERMES.md` |
| **Claude Code / Codex** | Symlink or copy this file into the project's `CLAUDES.md` / `.claude` instructions |
**Do not edit SOUL.md or .goosehints directly.** Edit this file in the
homelab-context repo instead. Changes propagate to all clients on the next
sync (`sudo homelab sync`).
---
## Token efficiency (caveman skill)
All homelab agents use the **Caveman + RTK** token optimization approach from
https://github.com/adityahimaone/hermes-agent-rtk-caveman.
### Before running any CLI command, ask:
1. **Is there a caveman wrapper equivalent?** Use the wrapper for token-efficient
output. Available wrappers (installed at `~/bin/caveman_wrapper.sh`):
- `~/bin/caveman_wrapper.sh git-status` — compact git status
- `~/bin/caveman_wrapper.sh git-log [n]` — compact git log
- `~/bin/caveman_wrapper.sh lint [target]` — compact lint results
- `~/bin/caveman_wrapper.sh test-results [cmd]` — compact test results
2. **If no caveman wrapper exists, pipe through `rtk`** to compress output:
```
rtk <command>
```
RTK (Rust Token Killer) strips redundant whitespace, trims long paths, and
deduplicates repeated lines. This reduces token usage by 60-90% on CLI
operations.
3. **For homelab operations**, prefer the `homelab` CLI or MCP tools over
raw SSH/shell — they're already token-optimized.
### Templates
Caveman templates live at `~/templates/`:
- `git_status.txt` — compact git status format
- `git_log.txt` — compact git log format
- `lint_results.txt` — compact ESLint format
- `test_results.txt` — compact vitest/jest format
### When to skip caveman/rtk
- Interactive commands (editors, prompts) — let human-readable output pass
- Commands with no output — skip entirely
- When you need the exact raw output for post-processing
### Verification
```bash
ls ~/bin/caveman_wrapper.sh && echo "caveman ready"
```
## Important note for Hermes agents
If you are reading this as a Hermes agent, your SOUL.md was auto-provisioned
by `tools/setup-hermes-soul.sh`. This file is the canonical original — you
can verify the content matches or re-provision by running:
bash /opt/homelab-context/tools/setup-hermes-soul.sh

View File

@@ -1,227 +0,0 @@
# Oikos — the operating model
Oikos (Greek: *household*) is the agent operating system layered on this
repo. It is not new infrastructure: `inventory.yaml` is the kernel data
structure, the `homelab` CLI and MCP server are the syscall surface, and
this page defines the rules everything above them follows.
Read this after [AGENTS.md](../AGENTS.md). Machine-readable companions:
[oikos/ontology.yaml](../oikos/ontology.yaml) (systems model),
[oikos/policy.yaml](../oikos/policy.yaml) (risk & approval).
## The kernel loop: OODA
Every Oikos activity — scheduled probe, agent task, operator request — is
one pass through **Observe → Orient → Decide → Act**:
1. **Observe** — probes, drift detectors, and agent findings produce
**Signals** (structured records, not loose messages): pending updates,
high temperature, low disk, service down, cert expiry, stale backup,
inventory drift.
2. **Orient** — walk the ontology graph: what entity is affected, what
depends on it (blast radius), its lifecycle state, whether a runbook
matches, what the ledger says about past attempts.
3. **Decide** — the classifier scores **risk class × blast radius ×
confidence** and routes:
- **auto-act**: within autonomy policy, high confidence, contained radius
- **escalate**: operator approval via Matrix (✅/❌ reaction) or the
Oikos Console's `/approvals` page (destructive actions additionally
need a typed confirmation phrase either way)
- **queue**: informational — console + reports
The classifier can only *lower* autonomy relative to policy, never raise
it. When in doubt, escalate.
4. **Act** — execute through `homelab` commands or runbooks (never ad-hoc
SSH), then **verify** with the action's verification command, write a
**ledger** entry, resolve the Signal, and update docs in the same session.
## Primitives
| Primitive | What it is | Lives in |
|---|---|---|
| Host / Service | topology entities | `inventory.yaml` (+ generated `hosts/*.yaml`) |
| Secret | SOPS+age encrypted value, per-client recipients | `secrets/` + `.sops.yaml` |
| Runbook | executable workflow with risk class + verification | `.agents/skills/<name>/SKILL.md` |
| Signal | something needing attention, with lifecycle | `signals/` ledger (Week 3) |
| Change | one mutation: who, what, risk, approval, verification | `ledger/` (Week 2) |
| Approval | short-TTL signed grant for a gated action | approval engine (Week 3) |
| Incident | investigation narrative | `knowledge/sources/investigations/` |
| Plan | design doc for non-trivial work | `plans/` |
| Agent | enrolled client identity = its age pubkey | `inventory.yaml` + `.sops.yaml` |
## Risk classes (enforced, not advisory)
From [oikos/policy.yaml](../oikos/policy.yaml):
- **read_only** — status, logs, docs, inventory. Unattended.
- **reversible_low** — restart, cache clear, sync pull. Unattended + ledger.
- **config_mutation** — tracked-config edits (commit+push, never local),
deploys, upgrades, DNS/ingress changes. Operator approval.
- **destructive** — destroy, format, wipe, rotate, revoke. Approval +
typed confirmation phrase.
Lifecycle gates modify these: `provisioning` nodes are freely mutable
(nothing depends on them); `deprecated` nodes accept no new dependents;
anything touching a `destroyed` node is drift.
## The systems model
Eight domains — physical, compute, network, storage, software,
identity & access, operations, external — cover everything in the lab;
entities are connected by typed edges (`hosts`, `provides`, `mounts`,
`stores-on`, `routes-to`, `can-decrypt`, `depends-on`, `backs-up-to`, …)
defined in [oikos/ontology.yaml](../oikos/ontology.yaml). Rule of
completeness: **if it can break, be changed, or hold data, it has an
entity and edges.** Blast-radius questions ("what breaks if strong goes
down?") are graph walks, not doc archaeology.
Nodes move through an explicit lifecycle —
`planned → provisioning → active → migrating → deprecated → destroyed`
stored as `state:` in inventory (absent = active). Destroyed nodes live in
the `archaeology:` section. Each transition is a runbook checklist;
deprecation completes only when inbound edges reach zero.
Generated views: [infrastructure/topology.md](../knowledge/wiki/infrastructure/topology.md)
(Mermaid, regenerated from inventory) and the live, clickable version at
`oikos.hubris.network/graph` once the Console is deployed.
## Conventions carried forward
- Inventory is the truth; live state wins over narrative docs.
- Prefer `homelab` CLI and MCP over ad-hoc SSH.
- Meaningful changes update docs in the same session.
- Secrets are decrypted locally via per-client keys; never into docs/comments.
- Tracked configs change by commit + push, not local edits.
- Netbird is the preferred mesh path for new traffic.
- Agents are terse ([caveman.md](shared/caveman.md)), verify claims, and fix
collateral drift when found.
## Build status (30-day roadmap, started 2026-07-05)
- **Week 1**: policy, ontology, service contract, archaeology, topology
generator, this brief. Shipped.
- **Week 2**: context cards, `homelab service <name> …`, change ledger,
`node relations`, runbooks. Shipped.
- **Week 3**: ops scheduler + state cache (`homelab service <name> health`
is cache-first, `--live` forces a probe), drift detectors, signal engine
(`homelab signal …`), decision classifier (`homelab decide …`), approval
engine (`homelab approval …` — shared-HMAC grants; Matrix delivery is
Hermes's existing `@dtoro:avispero` send path, not a new bot, see
`oikos/approve.py`), daily brief + weekly report (`oikos/report.py`).
Shipped, except: Prometheus is still `planned` (see
[plans/2026-07-05-oikos-prometheus-lxc.md](../plans/2026-07-05-oikos-prometheus-lxc.md)) —
trend signals (disk-full prediction, temp creep) wait on that LXC; the
scheduler's disk check today is point-in-time only, and CPU/NVMe
temperature isn't probed at all yet (no confirmed sensor path on
hubris/strong). DNS-vs-inventory and generic tracked-config-cleanliness
drift checks are also deferred (see `oikos/drift.py` docstring).
- **Week 4**: Oikos Console v0 shipped — signals landing page, service
grid + detail, node/blast-radius view, live Mermaid graph, drift view,
approvals queue (approve/deny, destructive confirmation-phrase
enforced), daily/weekly reports. Server-rendered FastAPI + Jinja2, no
SPA build chain, tested end-to-end against live production data (see
`oikos/console/`). Deploys as a third webhook on `dtoro/Homelab-Docs`
(`/opt/oikos-console`, port :9831) — see
[oikos/console/deploy/README.md](../oikos/console/deploy/README.md) for
the Caddy route and Gitea webhook registration this repo can't do for
itself. Approval grants are now single-use (a second `check_grant` call
for the same request fails even within the TTL) and already exact-bound
to request id + entity + action.
**Not shipped as originally planned:** per-agent *age-key-signed*
request authentication — age has no signing primitive (it's an
encryption-only keypair format), so "age-key-signed" wasn't
buildable as stated. The real alternative (SSH-key signing via
`ssh-keygen -Y sign`/`-Y verify`, using each host's already-provisioned
SSH key) is real and buildable, but needs SSH public keys recorded in
inventory first — not there today. Moved to the 60/90-day backlog.
Authentik step-up re-auth on the approve/deny route is documented but
needs a live Authentik instance to configure — also backlog.
Docs pass done (this file, AGENTS.md, operations/commands.md); found
and fixed two more stale references while at it (DNS section still
pointed at destroyed LXC 124/dnsmasq instead of Technitium on 107, and
a `claudio-monitor` reference that's been deprecated since 2026-06-04).
### Real drift found while building Week 3 (unresolved, needs operator action)
The drift detectors surfaced genuine, currently-true findings on first
run against production — recorded here rather than silently fixed, since
each is a `config_mutation`/`destructive`-class decision:
- `republic-laptop` has no `age_pubkey:` in `inventory.yaml`, but its real
age key is granted on nearly every shared secret in `.sops.yaml`
(`age1vf8h7...`) — the enrollment write-back to inventory never
happened. Fix: `homelab client add republic-laptop --finalize-pubkey
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6`.
- `grimmory` has an `age_pubkey` in inventory but is missing from
`secrets/hello.yaml`'s recipient list — incomplete enrollment the
other direction. Fix: re-run `homelab client add grimmory
--finalize-pubkey <its key>`.
- `pve_id 131` exists live on hubris (`pct list`) with no inventory entry
— investigate before assuming it's a stale ID (see the Prometheus LXC
plan doc above, which flags this explicitly).
- Three `lifecycle-pve-id-reuse` info findings (100, 106, 107 each shared
between an active host and an archaeology entry) — expected/benign ID
reuse after destroy, no action needed.
## 60/90-day backlog
Derived from gaps observed while building the 30-day roadmap, not
guesswork. Roughly ordered by what unblocks the most:
- **Fix the oikos-console deploy webhook's signature mismatch.** Console
is live on apps (105) via a manual `deploy.sh` run, but Gitea webhook
14's deliveries all 403 with a signature mismatch for a cause not yet
found — the secret is confirmed synced correctly on both sides
(rotated once already to rule out drift). Until fixed, `git push`
doesn't auto-redeploy the console the way it does for homelab-mcp/
secrets-issuance; re-run `deploy.sh` on apps manually after changes.
See [oikos/console/deploy/README.md](../oikos/console/deploy/README.md).
- **SSH-key-signed approval requests.** Replaces the design note in
Week 4: age keys can't sign (encryption-only format), so per-agent
request authentication needs `ssh-keygen -Y sign`/`-Y verify` against
each host's existing SSH key. Blocked on a schema gap: inventory
doesn't record SSH public keys today, only ports/users. First step is
populating that field on enrollment, then wiring `oikos/approve.py` to
require and verify a signature over the request payload.
- **Authentik step-up re-auth** on the Console's `/approvals` POST route
— needs a live Authentik `PromptStage`/reauth flow scoped to that path;
not configurable without a running instance to test against.
- **Prometheus provisioning** (see
[plans/2026-07-05-oikos-prometheus-lxc.md](../plans/2026-07-05-oikos-prometheus-lxc.md))
— unblocks trend signals (disk-full prediction, temp creep) and real
sparklines in the Console; investigate the undocumented `pve_id 131`
on hubris first.
- **CPU/NVMe temperature probing** in the scheduler — needs a confirmed
sensor path on hubris and strong (lm-sensors vs vendor tool) before a
real check can be written; guessing one risks a probe that silently
never fires.
- **DNS-vs-inventory drift check** — compare Technitium zone records
against `services.*.url`/`public_host`; not implemented (`oikos/drift.py`
has no Technitium API wiring yet).
- **Generic tracked-config-cleanliness drift check** — today only caddy's
`/etc/caddy` git-checkout path is hardcoded in `oikos/drift.py`; every
other service with a `config_repo` needs its local checkout path
recorded (a `mutation_path`-style field, same gap Week 1's service
contract flagged but didn't backfill) before this generalizes.
- **Per-service policy overrides** (`oikos/policy.yaml`
`service_overrides`) — schema is ready (caddy/dns already use it);
populate more as specific services turn out to need non-default risk
classes.
- **Incident timeline generator** — stitch ledger + signal history into
a single narrative for `knowledge/sources/investigations/` entries instead of writing
them by hand.
- **Secret access audit** — who-can-decrypt-what report from
`.sops.yaml` + inventory `age_pubkey`s, extending what
`oikos/drift.py`'s SOPS check already partially does.
- **Restore drills** — exercise `backs-up-to` (once populated) by
actually restoring from a backup target on a schedule, not just
checking freshness.
- **Multi-agent delegation model** — more than one agent acting
concurrently; needs the ledger's `agent` field to carry real identity
(age pubkey, not just hostname) consistently, which it mostly does
already but hasn't been stress-tested with concurrent writers.
- **Grafana** — only if the Console's own Prometheus-backed sparklines
turn out to be insufficient once Prometheus ships.
- **"Generalize later" extraction** — the original decision was personal-
first, generalize-later (see Week 1). Once patterns stabilize, extract
a config-driven Oikos core with no `hubris.network`/`hubris`/`strong`
hardcoding, so it's installable on a different homelab.

View File

@@ -1,48 +0,0 @@
# Knowledge domain — schema
The knowledge domain is the durable, authoritative current-state documentation of the homelab: one
page per node and per cross-cutting system, synthesized from live state and evidence. It answers
"what exists and how does it work right now."
It follows the [LLM Wiki layer model](../../shared/llm-wiki.md) and the
[writing-style](../../shared/writing-style.md) and [page-templates](../../shared/page-templates.md)
rules.
## The narrative / substrate split
The knowledge wiki is **narrative**. It sits alongside a **machine-readable substrate** that it
describes but never contains. The split is load-bearing: several programs read the substrate at
fixed paths, so the wiki reorganization never moves it.
| Layer | Location | Consumed by |
|-------|----------|-------------|
| Substrate — source of truth | `inventory.yaml` (root) | MCP server, `homelab` CLI, `oikos/` scheduler/drift/relations/gen-topology |
| Substrate — generated host records | `hosts/*.yaml` (root) | `mcp/server.py` (`HOSTS_DIR`), `bin/homelab`; written by `mcp/build_host_files.py` |
| Substrate — kernel + context cards | `oikos/` (code, `oikos/cards/`, `oikos/state.json`) | MCP `explain`, scheduler |
| Narrative — synthesized wiki | `knowledge/wiki/{hosts,containers,vms,infrastructure}/` | humans, agents via MCP `get_page` / `search_docs` |
| Evidence — immutable sources | `knowledge/sources/` (references + investigations) | synthesis into wiki pages |
## Wiki pages
- **Node pages** (`knowledge/wiki/containers/<id>-<name>.md`, `.../vms/<id>-<name>.md`,
`.../hosts/<name>.md`) follow the container/host template in
[page-templates.md](../../shared/page-templates.md): opening definition, `## At a glance`,
`## Role`, service/port map, storage, auto-deploy, `## Related`, `## Changelog`.
- **Cross-cutting pages** (`knowledge/wiki/infrastructure/<topic>.md`) follow the cross-cutting
template: `## Why`, `## Components`, `## How to apply`, `## Gotchas`, `## Related`, `## Changelog`.
- Each `inventory.yaml` host entry carries a `doc_page:` field pointing at its narrative page.
Changing where a page lives means updating that field (read by `bin/homelab`).
## The two logs
- The per-page **`## Changelog`** records infrastructure changes and is machine-parsed
(`get_changelog`, the Oikos ledger). Keep the `### YYYY-MM-DD — title` shape.
- **`knowledge/log.md`** is append-only and records *documentation-maintenance* operations only
(restructures, source ingests, lint sweeps): `## [YYYY-MM-DD] <op> | <summary>`. It never
duplicates the Oikos change ledger (`oikos/ledger.py`).
## Same-session update rule
A change to a node updates every page that references it in the same session — the node page, the
section `README.md` table, the root `README.md`, the Caddy/DNS/ingress pages, the host page, and
`inventory.yaml`. See [page-templates.md](../../shared/page-templates.md#same-session-update-rule).

View File

@@ -1,55 +0,0 @@
# Operations domain — schema
The operations domain holds the procedural and time-stamped documentation: runbooks (repeatable
procedures), investigations (incident evidence), and plans (design docs for non-trivial work). It
follows [writing-style](../../shared/writing-style.md); runbooks and plans use the imperative voice
exception.
Where each kind lives: runbooks are skills under [`.agents/skills/`](../../skills/); operator
reference (command cheatsheet, enrollment, Hermes agent) lives in
[`.agents/operations/`](../../operations/); investigations are sources under
`knowledge/sources/investigations/`; plans stay in the repo-root `plans/` folder (below).
## Plans always live in `plans/`
**Any plan or design doc for the Homelab is written into the repo `plans/` folder as
`plans/YYYY-MM-DD-slug.md` — never a scratch path, an agent-private plan location, or a chat
message.** An agent drafting a plan:
1. Writes the file under `plans/` using the plan template in [page-templates.md](../../shared/page-templates.md).
2. Lists it in `plans/index.md`.
3. On completion, moves it to `plans/done/` and updates the index status.
This is the single source for homelab design intent; keeping it in-repo means the plan is
versioned, reviewable, and reachable by MCP `get_page`/`search_docs` like any other doc.
## Runbooks
Repeatable procedures are skills — one folder per skill at `.agents/skills/<name>/SKILL.md`, with
YAML front-matter that the Oikos policy and lifecycle machinery reads:
```yaml
---
name: <name>
risk_class: read_only | reversible_low | config_mutation | destructive
inputs: [<param>, ...]
verification: "<shell expression that proves success>"
docs_update_checklist: [<doc artifacts to update>]
transition: "<from> -> <to>" # only for lifecycle runbooks
---
```
`risk_class` values and the lifecycle `transition` states must match
[`oikos/policy.yaml`](../../../oikos/policy.yaml) and [`oikos/ontology.yaml`](../../../oikos/ontology.yaml).
## Investigations
Incident records live in `knowledge/sources/investigations/YYYY-MM-DD-slug.md` and are **evidence sources** — written
once at incident time, then linked from the changelogs of the nodes they implicate. Sections:
`## Summary`, `## Timeline`, `## Root cause`, `## Mitigations applied`, `## Open questions`. Resolved
incidents move to `knowledge/sources/investigations/archive/`.
## The operations log
`plans/log.md` and `knowledge/log.md` are append-only records of documentation operations on
those areas (`## [YYYY-MM-DD] <op> | <summary>`), distinct from the Oikos change ledger.

View File

@@ -1,91 +0,0 @@
# Operations cheatsheet
Run from the [hubris host](../../knowledge/wiki/hosts/hubris.md) as root. When working from `/root` on Linux you're already on hubris — don't `ssh hubris` / `ping hubris`.
## Proxmox CLI
| Command | Use |
| --- | --- |
| `pct list` / `qm list` | List LXC containers / VMs |
| `pct config <id>` / `qm config <id>` | Container / VM config |
| `pct exec <id> -- <cmd>` | Run command inside an LXC without entering it (no initgroups — see [media permissions](../../knowledge/wiki/infrastructure/media-permissions.md)) |
| `pct enter <id>` | Shell into a container |
| `pct start <id>` / `pct stop <id>` | Boot / halt a container |
| `pvesm status` | Storage pools status |
| `pvesh get /nodes --output-format json` | Node summary as JSON |
| `pvesh get /nodes/hubris/lxc/<id>/status/current` | Live container status |
| `pvesh get /cluster/resources --type vm --output-format json` | Bulk per-LXC CPU/mem/disk (used by the `homelab-health-watchdog` Hermes cron — see [monitoring](../../knowledge/wiki/infrastructure/monitoring.md); the old `claudio-monitor` this once fed is deprecated) |
| `pveversion` | PVE version |
| `journalctl -u pve-cluster -n 100` | PVE service logs |
## Storage
- Shared mount: `/mnt/library` (ext4 on lvmthin `library`).
- Bind into a container: `pct set <id> -mp<N> /mnt/library/<sub>,mp=/data`
- For the standard whole-tree mount: `pct set <id> -mp0 /mnt/library,mp=/mnt/library`. See [media permissions](../../knowledge/wiki/infrastructure/media-permissions.md) for the GID-10000 onboarding recipe.
## Reverse proxy
- Caddyfile: `/etc/caddy/Caddyfile` on [LXC 121](../../knowledge/wiki/containers/121-caddy.md).
- **CRITICAL:** This file is tracked in `dtoro/caddy-conf` (https://git.hubris.network/dtoro/caddy-conf). Never edit it directly on the LXC — commit + push to the repo instead. Caddy auto-deploys on push (see [auto-deploy](../../knowledge/wiki/infrastructure/auto-deploy.md)). If you edit directly, the change will be lost on the next pull and agents won't know about it.
- Hot reload: `pct exec 121 -- systemctl reload caddy`.
- Validate: `pct exec 121 -- caddy validate --config /etc/caddy/Caddyfile`.
- Git workflow shortcut: `pct exec 121 -- "cd /etc/caddy && git add Caddyfile && git commit -m '...' && git push"`.
## DNS
- Split-horizon authority: [Technitium DNS](https://technitium.com) on [dns (107)](../../knowledge/wiki/containers/107-dns.md) at `192.168.8.2:53`. Web UI at `http://192.168.8.2`. (Formerly dnsmasq on the now-destroyed LXC 124 — decommissioned 2026-06-04.)
- Add/edit records in the Technitium UI; the NetBird managed zone sync (`scripts/dns-sync.py` cron on 107) picks changes up within ~10 minutes.
- Verify: `dig @192.168.8.2 +short <host>.hubris.network`.
- See [DNS](../../knowledge/wiki/infrastructure/dns.md).
## Web access
- `https://proxmox.hubris.network` or `https://192.168.8.77:8006` — Proxmox UI
## Telemetry quick checks
- `ras-mc-ctl --summary` — summary of any RAS events (memory / PCIe AER / thermal) since boot
- `ras-mc-ctl --errors` — full event log
- `cat /sys/devices/system/cpu/cpu0/cpufreq/energy_performance_preference` — should be `balance_power`
- `cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor` — should be `powersave`
- `ls /sys/fs/pstore/ /var/lib/systemd/pstore/` — panic traces from a previous crash (empty for pure hardware hangs — see [investigation](../../knowledge/sources/investigations/archive/2026-04-21-hubris-crash-loop.md))
## Fleet apt operations
Two `homelab` subcommands wrap the common patterns; both fan out to hubris + every LXC.
| Command | What it does |
| --- | --- |
| `homelab apt-audit [--target HOST]` | Per-host table: dpkg-interrupted state, holds, upgradable count, non-apt binaries in system paths, DNS health. Exits nonzero if any host has dpkg-interrupted state. |
| `homelab apt-upgrade --target HOST` | Launch `apt update && apt upgrade` inside a transient `systemd-run --collect` unit on the target. Survives ssh teardown. Apt configured with `Acquire::Retries=3` + `ForceIPv4=true`. |
| `homelab apt-upgrade --all` | Same, fanned out across the standard targets. |
| `homelab apt-upgrade ... --status` | Show running unit + tail `/var/log/homelab-apt-upgrade.log` on each target. |
| `homelab apt-upgrade ... --safe` | Take a pre-upgrade snapshot per LXC first (`pct snapshot``vzdump` fallback for bind-mounted LXCs). Refuses if any snapshot fails unless `--force`. |
| `homelab apt-upgrade ... --force` | Skip both the dpkg-audit gate and snapshot-failure refusal. |
PVE/kernel deferral on hubris: `homelab apt-upgrade --target hubris` will try every upgrade, including kernel + `pve-*`. To skip those, `apt-mark hold` the relevant packages on hubris first; `homelab apt-audit` shows held packages so you can confirm.
## Oikos (agent OS layer)
See [OIKOS.md](../OIKOS.md) for the operating model. Quick reference:
| Command | What it does |
| --- | --- |
| `homelab service <name> explain\|health\|docs\|log\|actions\|history` | Service Console v0 — context card, cached health (`--live` to force a probe), docs, logs, safe actions + risk class, ledger history |
| `homelab node <name> relations` | Ontology blast-radius query: what this host/service impacts, is affected by, and its full transitive blast radius |
| `homelab change preflight <service>` | Dry-run report before mutating: risk class, current health, config repo, verification command |
| `homelab decide <action> <entity>` | Decision classifier: risk × blast radius × confidence → auto-act or escalate |
| `homelab signal list\|raise\|ack\|resolve\|mute` | The attention layer — pending updates, thresholds, drift, anything needing attention |
| `homelab approval request\|list\|reply\|check` | Escalate-route grants (Matrix-delivered via Hermes, or the Oikos Console's `/approvals` page) |
| `homelab restart <service> [--approval-id <id>]` | `--approval-id` is required whenever the service's risk class needs approval (e.g. `caddy`, `dns`) — refuses mechanically without a valid grant |
Oikos Console (read-mostly dashboard): `oikos.hubris.network` once deployed — see [oikos/console/deploy/README.md](../../oikos/console/deploy/README.md).
## Related
- [Hubris host](../../knowledge/wiki/hosts/hubris.md)
- [Containers index](../../knowledge/wiki/containers/index.md)
- [DNS](../../knowledge/wiki/infrastructure/dns.md)
- [Monitoring](../../knowledge/wiki/infrastructure/monitoring.md)
- [Auto-deploy](../../knowledge/wiki/infrastructure/auto-deploy.md)
- [Runbook: dpkg-interrupted recovery](../skills/runbook-dpkg-interrupted/SKILL.md) — what to do when apt got killed mid-transaction

View File

@@ -1,33 +0,0 @@
# CAVEMAN.md — communication mode for homelab agents
Respond terse like smart caveman. All technical substance stay. Only fluff die.
## Rules
Drop: articles (a/an/the), filler (just/really/basically/actually/simply), pleasantries (sure/certainly/of course/happy to), hedging. Fragments OK. Short synonyms (big not extensive, fix not "implement a solution for"). Technical terms exact. Code blocks unchanged. Errors quoted exact.
Pattern: `[thing] [action] [reason]. [next step].`
Not: "Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by..."
Yes: "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:"
## Levels
- **lite** — no filler/hedging. Keep articles + full sentences. Professional but tight.
- **full** (default) — drop articles, fragments OK, short synonyms. Classic caveman.
- **ultra** — abbreviate prose words (DB/auth/config/req/res), strip conjunctions, arrows (X → Y). Code symbols/API names/errors: never abbreviate.
Switch: `/caveman lite|full|ultra`. Stop: "normal mode".
## Auto-Clarity
Drop caveman for: security warnings, irreversible actions, multi-step sequences where fragments risk misread, user confused/repeating. Resume after clear part done.
## Boundaries
Code/commits/PRs: write normal. "stop caveman" or "normal mode": revert. Level persist until changed or session end.
---
Source: https://github.com/JuliusBrussee/caveman
Copy to `~/.hermes/skills/` for Hermes Agent, or `~/.claude/projects/<name>/SKILL.md` for Claude Code.

View File

@@ -1,41 +0,0 @@
# LLM Wiki — the documentation contract
How the narrative documentation in this repo is organized. The pattern is borrowed from the
`sources / wiki / index / log` model: a durable synthesized layer (`knowledge/wiki/`) built on top
of immutable evidence (`knowledge/sources/`, incident records), with pure-listing indexes and an
append-only operations log.
This contract governs the **narrative layer only**. The machine-readable substrate — `inventory.yaml`,
generated `hosts/*.yaml`, `oikos/`, `mcp/`, `secrets/`, `bin/` — is not part of the wiki and never
moves under it. See [the knowledge schema](../domains/knowledge/schema.md) for the split.
## Layers
- **Sources** are immutable raw material: incident records (`knowledge/sources/investigations/`), external reference
docs (`knowledge/sources/references/`), and the live system itself (`pct config`, `docker inspect`).
Read them; do not rewrite them into other sources.
- **Wiki** (`knowledge/wiki/`) is the synthesized, authoritative current-state layer: one page per
node (`containers/`, `vms/`, host narratives) and per cross-cutting system (`infrastructure/`). A
reader understands the topic from the wiki page without reading the sources.
- **Index** (`index.md` / folder `README.md`) is a pure listing — every page in scope with a
one-line summary, and nothing else. Anything the section wants to say up front goes into a page
the index lists, not into the index.
- **Log** (`log.md`) is append-only, recording *doc-maintenance operations* (restructures, source
ingests, lint sweeps) in single-line format: `## [YYYY-MM-DD] <op> | <summary>`.
## Two logs, kept distinct
- **`## Changelog`** on each node/topic page records *infrastructure* changes to that node. It is
machine-parsed (`get_changelog`, the Oikos ledger) — keep the `### YYYY-MM-DD — title` shape.
- **`log.md`** per area records *documentation* operations only. It never duplicates the Oikos
change ledger (`oikos/ledger.py`), which stays authoritative for infra changes with
who/what/risk/approval/verification.
## Rules
- Wiki pages stay short and focused. A page past ~300 lines splits.
- Pages stay flat under `wiki/<section>/` until there are enough to warrant a sub-group.
- Every page follows [writing-style.md](writing-style.md).
- Plans and design docs always live in the repo `plans/` folder (`plans/YYYY-MM-DD-slug.md`),
listed in `plans/index.md`, moved to `plans/done/` on completion — never a scratch path or a chat
message. See [the operations schema](../domains/operations/schema.md).

View File

@@ -1,174 +0,0 @@
# Page templates for the Homelab Wiki
The structural templates for each page type. Prose voice, vocabulary, and cross-reference rules live
in [writing-style.md](writing-style.md); the layer model (sources / wiki / index / log) lives in
[llm-wiki.md](llm-wiki.md).
## File naming
**Foundational / entry-point files:** ALL-CAPS
- **Root level:** `AGENTS.md`, `README.md` — discovery paths for agents and humans.
- **Agent instruction** (under `.agents/`): `OIKOS.md`, `HERMES.md` — foundational docs agents read before acting.
- **Reference docs:** `GLOSSARY.md` — lookup reference (like classic repo conventions: LICENSE, CHANGELOG, GLOSSARY).
**Content / narrative pages:** lowercase-with-dashes, date-prefixed as needed
- **Container pages:** `<id>-<name>.md` (e.g. `101-jellyfin.md`, `132-rclone.md`). The `<id>` is the LXC/VM ordinal from `inventory.yaml`.
- **Infrastructure / cross-cutting pages:** `<topic>.md` (e.g. `dns.md`, `auto-deploy.md`, `mesh.md`). Describes a system, not a specific node.
- **Plans / investigations:** `YYYY-MM-DD-<slug>.md` (e.g. `2026-07-05-oikos-prometheus-lxc.md`). Date-sorted; slug is lowercase.
- **Section indices:** `README.md` (lowercase, conventional). Prefer in folders; `index.md` only if both intro prose and listing coexist.
**Skills / runbooks:** special case
- **Folder structure:** `<name>/SKILL.md` where `<name>` is lowercase-with-dashes (e.g. `client-enrollment/SKILL.md`).
- **The filename SKILL.md is always uppercase** — it acts as a signpost so tools and humans instantly recognize it as a skill.
**General rules:** All paths use lowercase letters, numbers, and hyphens (no underscores). Uppercase is reserved for foundational docs (entry points + instruction) and filenames that signify document type (SKILL.md, GLOSSARY.md, etc.).
## Voice
Concise, technical, sysadmin-to-sysadmin. No marketing prose, no exclamation marks. Full rules in
[writing-style.md](writing-style.md).
## Page templates
### Container page (`containers/<id>-<name>.md`)
```markdown
# <id> — `<name>`
One-sentence purpose.
## At a glance
- **Hostname:** `<name>`
- **IP:** `192.168.8.x`
- **Privilege:** privileged | unprivileged
- **Resources:** N cores / M GiB RAM / D GiB rootfs
- **Mounts:** `/mnt/library``/mnt/library` (if any)
- **Public hostname:** `<sub>.hubris.network` (if proxied)
## Role
What it does, what it talks to.
## Service / port map
| Service | Listen | Notes |
## Storage / config paths
## Auto-deploy
(if any) — link to [auto-deploy](../infrastructure/auto-deploy.md)
## Related
- [Caddy](121-caddy.md) (if proxied)
- [DNS](../infrastructure/dns.md) (if has subdomain)
- [Authentik](124-authentik.md) (if SSO)
- ...
## Changelog
### YYYY-MM-DD — short title
What changed, why, link to investigation if any.
```
### Cross-cutting page (`infrastructure/<topic>.md`)
```markdown
# <Topic>
One-sentence summary.
## Why
Design rationale — what it replaces, what it solves.
## Components
Where it runs, what files matter.
## How to apply / use
Recipes.
## Gotchas
## Related
Links to nodes that host or depend on this.
## Changelog
```
### Plan (`plans/YYYY-MM-DD-slug.md`)
```markdown
# YYYY-MM-DD — <title>
## Goal
What this change achieves and why.
## Current topology / state
Diagram or description of what exists now.
## Target topology / state
What it looks like after.
## Pre-flight checklist
## Step-by-step procedure
## Verification
## Post-migration
Changelog entries to write, index status to update.
```
### Investigation (`knowledge/sources/investigations/YYYY-MM-DD-slug.md`)
```markdown
# YYYY-MM-DD — <title>
## Summary
1-3 sentences.
## Timeline
## Root cause
## Mitigations applied
## Open questions
```
## Linking discipline
- Every container page links to every cross-cutting page it participates in.
- Every cross-cutting page lists the nodes that participate.
- Every investigation links to the nodes it implicates *and* gets back-linked from each node's changelog.
- Every plan links to the infrastructure pages it affects. When done, update the plan's status in `plans/index.md` and write changelog entries on affected node pages.
## Changelog hygiene
- Reverse-chronological (newest first).
- One entry per discrete change, even if you make several in one day.
- If a change spans nodes, repeat the entry on each affected page (different perspective is fine).
- Don't rewrite history — entries are append-only. Mistakes get a follow-up entry that supersedes them.
## Same-session update rule
When you make a change to a node — migrate an LXC, update an IP, change a
mount, deploy a new service — **update every relevant doc page in the same
session.** A change that touches a container page must also update:
- The `containers/index.md` table (IPs, host, mounts, status)
- The `README.md` table (if the change affects listed columns)
- The Caddy page site list (if the change affects `*.hubris.network` routing)
- The DNS / ingress infrastructure pages (if the change affects routing)
- The `hosts/{hubris,strong}.md` host page (if container count changes)
- The `inventory.yaml` host entry (source of truth for the `hosts/*.yaml` generation)
- The `infrastructure/topology.md` (generated from inventory, but regen if needed)
The pattern of updating only one page and leaving stale references on others
is a bug. If you're doing a multi-step migration, document the intermediate
state with a changelog entry that says "pending — will finalize after Phase
N."
This rule is why Phase 2 of the strong migration (2026-07-05) caused
widespread stale data: individual container pages were updated in the
changelog but never had their At-a-glance sections, IPs, mount paths, or
host attribution updated. Don't repeat that.

View File

@@ -1,75 +0,0 @@
# Writing Style
Write like a technical reference, not a marketing page. Every sentence conveys new information.
These rules govern **committed documentation** — wiki pages, READMEs, schemas, skills, `AGENTS.md`,
plans, investigations, and code comments. They are separate from [caveman.md](caveman.md), which
governs an agent's *chat responses*; the two do not conflict.
New or rewritten pages follow these patterns from day one. Existing pages get updated the next time
they are touched.
## Vocabulary — never use these
- Significance puffers: "pivotal", "crucial", "vital", "groundbreaking", "transformative", "testament", "paramount", "invaluable".
- Analytical verbs: "delve", "leverage", "utilize", "facilitate", "foster", "showcase", "underscore", "streamline", "harness".
- Poetic nouns: "tapestry", "landscape" (figurative), "realm", "paradigm", "ecosystem" (figurative), "journey" (figurative), "nexus", "cornerstone".
- Promotional adjectives: "robust", "seamless", "innovative", "cutting-edge", "meticulous", "holistic", "comprehensive".
- Opening crutches: "In today's world", "In the ever-evolving landscape of", "It's worth noting that", "It is important to note that".
Use short, common words: "use" not "utilize", "help" not "facilitate", "show" not "demonstrate".
## Voice
Describe what systems do and how they work.
- **Reference prose** (node pages, cross-cutting infrastructure descriptions, `## Role`, `## Why`,
`At a glance`) is third-person: state facts about the system, not instructions to a reader.
- **Recipes, runbooks, and skills** are the exception: second-person imperative is allowed and
preferred where it makes a procedure clearer ("Edit the Caddyfile, commit + push", "Verify with
`dig +short`"). This matches how the operator actually works. The vocabulary, structure, and
cross-reference rules below still apply.
## Page shape
Every doc-level page follows the same shape so a reader scans it in one pass.
1. **One H1 = the page title.** Node pages use `# <id> — \`<name>\``; topic pages use `# <Topic>`.
2. **Opening definition.** First paragraph, 13 sentences, says what the thing is. No motivation, no marketing, no setup.
3. **Body sections** in the natural order for the topic. Reuse the section templates in [page-templates.md](page-templates.md).
4. **`## Changelog`** at the bottom of every node/topic page — reverse-chronological, append-only. This section is machine-parsed (`get_changelog` in `mcp/server.py`); keep the `### YYYY-MM-DD — title` shape.
5. **Related links** only at the bottom, only when a reference cannot be woven inline.
## Section indexes (folder READMEs)
A folder's `README.md` opens with a 13 sentence prose intro that says what the section covers, then
a single navigation table — `| Document | What it covers |` — and nothing else. No stale counts, no
duplicated prose, no narrative between the intro and the table.
## Structure rules
- Make every sentence information-dense. Cut filler, qualifiers, and setup phrases. Lead with the concrete fact or action, not why it matters.
- No participial tack-ons (", highlighting the importance of…"). If the clause adds information, make it a separate sentence.
- **No meta-commentary about the content itself.** Do not narrate the page's own structure or linking strategy.
- Prefer **tables** for enumerable items with internal structure (service/port maps, field lists, status grids). Reserve bullets for short non-structured lists.
- Use the **bold-leading-phrase pattern** for structured points: `**Read-only by construction.** The MCP server never mutates state.` — a bold noun phrase, a period, then the explanation.
- When enumerating across services or nodes, give each its own `###` sub-section or a table row, not one run-on paragraph.
- Use backticks for code, paths, hostnames, and file names (`inventory.yaml`, `192.168.8.77`, `pct config`); italics for first-mention terminology.
- Use `>` blockquotes for caveats and gaps that interrupt the main flow: `> **Outstanding gap.** DNS-vs-inventory drift check not yet wired.` One thought per blockquote.
## Diagrams
- Mermaid is the default for topology and flow diagrams. `infrastructure/topology.md` is generated by `oikos/gen-topology.py` — do not hand-edit it.
- ASCII box diagrams are fine for small shape diagrams; keep them to one screen.
## Sourcing and cross-references
- **Factual discipline.** Every claim is grounded in a cited source, an adjacent linked page, or a directly observable fact (`pct config`, `docker inspect`, running config). Do not write sentences that sound sourced but are inference. When docs disagree with live state, fix the doc and note it in the changelog.
- **One-sided cross-references.** When two pages relate, the link lives in the page where the connection makes organizational sense. Do not add a back-pointer unless that direction also carries content the reader needs.
- **Cross-references are content, not catalog.** Inline links arise from the surrounding prose; the linked page must be needed to understand the current sentence. A bottom-of-page "Related" list is the fallback, not the default.
- Pages link with standard relative markdown links (e.g. a container page links to `../infrastructure/dns.md`), forming a navigable graph. Orphans are a bug.
## Code comments and commit/PR prose
- Comments explain intent, trade-offs, or constraints the code cannot convey. No diff narration, no type restatement, no section-divider comments.
- Commit messages and PR descriptions are problem → change → risk → verification, not a file-by-file diff restatement.
- The banned vocabulary applies the same way in comments and commit messages.

View File

@@ -1,43 +0,0 @@
---
name: client-enrollment
risk_class: config_mutation
inputs: [hostname, kind, role]
verification: "homelab doctor (on the new client)"
docs_update_checklist: [hosts_narrative_page_if_lxc_or_vm]
---
# Client enrollment
Goal: bring a new host (workstation, LXC, VM) into inventory and the
secrets model, with mesh membership only where it's actually needed.
This wraps the existing `homelab client add` flow — see
[operations/agent-enrollment.md](../../operations/agent-enrollment.md) for
the full walkthrough; this runbook is the risk/lifecycle framing.
1. On any enrolled client: `homelab client add <hostname>` — appends a
`hosts.<name>:` block to `inventory.yaml` (lifecycle `state: planned`
`provisioning`, per [oikos/ontology.yaml](../../../oikos/ontology.yaml)),
commits + pushes.
2. Netbird join is **optional, not a required step** — only needed for
hosts that must be reachable off-LAN (workstations that roam, e.g.
`republic-laptop`, `mac-mini`). A node reachable on the household LAN
(192.168.8.0/24 — most LXCs/VMs) doesn't need it: it's already
reachable directly, and off-LAN clients reach it too via hubris's
routed `192.168.8.0/24` Netbird network resource. Skip this step for
LAN-only nodes; do it (out-of-band, console or setup key) only for
hosts that need independent off-LAN reachability.
3. On the new host: run `bootstrap.sh` (add `--with-hermes` to also
enroll the Hermes agent). This provisions `/etc/age/key.txt`, the
sync timer, and prints an age pubkey.
4. Back on an enrolled client: `homelab client add <hostname>
--finalize-pubkey <age1...>` — sets `age_pubkey`, grants shared
secrets, re-keys SOPS, commits + pushes. This is the
`provisioning → active` transition.
5. Verify: `homelab doctor` on the new client should show all checks
green (clone, sync timer, age key, CLI symlink, MCP reachable).
Docs-update checklist: if the new host is an LXC/VM, add its narrative
page under `containers/` or `vms/` and set `doc_page` in its inventory
entry (host-level cards don't have a `doc_page` field yet — services do;
narrative pages are still found via the generated `see_also` in
`hosts/<name>.yaml`).

View File

@@ -1,34 +0,0 @@
---
name: config-change-deploy
risk_class: config_mutation
inputs: [service_name, change_description]
verification: "curl -sf <service_url> (or homelab service <name> health)"
docs_update_checklist: [doc_page, changelog]
---
# Config change + deploy
Goal: change a tracked config repo (Caddy, Gitea customizations, an app's
own repo) and get it live, safely.
1. `homelab change preflight <service>` — current health, the service's
`config_repo`, its risk class, and the verification command to run
after. If risk class requires approval (`config_mutation` or
`destructive`), stop and get operator sign-off before editing — see
`oikos/policy.yaml`.
2. Clone/pull the `config_repo` (never edit the backend's working tree
directly — tracked configs change by commit + push, per
[OIKOS.md](../../OIKOS.md) conventions).
3. Make the change, commit, push to `main`.
4. The Gitea webhook fires the deploy pipeline for that repo (see
[infrastructure/auto-deploy.md](../../../knowledge/wiki/infrastructure/auto-deploy.md) for
the exact receiver/reload for this service).
5. Run the preflight's verification command. If it fails, check
`homelab service <name> log` for the reload/restart error.
6. Record the change: once `oikos/ledger.py` is wired into deploy tooling
(Week 3), this is automatic; until then, note the change and outcome
in the relevant investigation/plan doc.
Docs-update checklist: update the service's `doc_page` if the change
alters its behavior, ingress route, or ownership; add a changelog entry
if the page has one.

View File

@@ -1,25 +0,0 @@
---
name: docs-lint
risk_class: read_only
inputs: [paths]
verification: "python3 .agents/skills/docs-lint/lint.py"
docs_update_checklist: []
---
# Docs lint
Check committed documentation against the mechanical rules in
[writing-style.md](../../shared/writing-style.md): banned vocabulary and broken relative markdown
links. Prose-voice rules are not machine-checkable — those stay a review responsibility.
Run from the repo root:
python3 .agents/skills/docs-lint/lint.py # default: knowledge/ .agents/ operations/ investigations/ plans/
python3 .agents/skills/docs-lint/lint.py knowledge/wiki/containers/104-gitea.md
Exit code is non-zero when any violation is found, so it can gate a commit. The banned-vocabulary
list mirrors `writing-style.md`; update both together if the standard changes.
> **Known baseline.** `knowledge/wiki/containers/101-jellyfin.md` links into a sibling repo
> (`devops/homelab-authentik-admin`) that this checkout does not contain — expected, not a bug.
> Any other broken link is a real regression; investigate before dismissing it as baseline noise.

View File

@@ -1,69 +0,0 @@
#!/usr/bin/env python3
"""Lint committed docs against .agents/shared/writing-style.md.
Checks two mechanical rules:
1. Banned vocabulary (significance puffers, analytical verbs, poetic nouns,
promotional adjectives, opening crutches).
2. Broken relative markdown links.
Prose-voice rules are not machine-checkable; this covers the parts that are.
Run from the repo root: python3 .agents/skills/docs-lint/lint.py [paths...]
Exit 1 if any violation is found.
"""
import os, re, sys
BANNED = [
"pivotal", "crucial", "vital", "groundbreaking", "transformative", "testament",
"paramount", "invaluable", "delve", "leverage", "utilize", "facilitate", "foster",
"showcase", "underscore", "streamline", "harness", "tapestry", "realm", "paradigm",
"nexus", "cornerstone", "robust", "seamless", "innovative", "cutting-edge",
"meticulous", "holistic", "comprehensive", "in today's world",
"it's worth noting", "it is important to note",
]
BAN_RE = re.compile(r'(?<![\w-])(' + "|".join(re.escape(w) for w in BANNED) + r')(?![\w-])', re.I)
LINK = re.compile(r'\]\(([^)]+)\)')
def iter_md(paths):
for p in paths:
if os.path.isfile(p) and p.endswith(".md"):
yield p
for root, dirs, files in os.walk(p):
dirs[:] = [d for d in dirs if d not in (".git", "node_modules")]
for f in files:
if f.endswith(".md"):
yield os.path.join(root, f)
def main(argv):
paths = argv or ["knowledge", ".agents", "operations", "investigations", "plans"]
violations = 0
# The style guide and this skill enumerate the banned words by definition.
ban_exempt = ("shared/writing-style.md", "skills/docs-lint/")
for f in sorted(set(iter_md(paths))):
check_banned = not any(x in f for x in ban_exempt)
fence = False
with open(f) as fh:
for ln, line in enumerate(fh, 1):
if line.lstrip().startswith("```"):
fence = not fence; continue
if fence:
continue
if check_banned:
for m in BAN_RE.finditer(line):
print(f"{f}:{ln}: banned word '{m.group(1)}'")
violations += 1
for m in LINK.finditer(line):
link = m.group(1)
if re.match(r'^(https?:|mailto:|#|/)', link):
continue
path = re.split(r'[#?]', link)[0]
if not path:
continue
tgt = os.path.normpath(os.path.join(os.path.dirname(f), path))
if not os.path.exists(tgt):
print(f"{f}:{ln}: broken link -> {link}")
violations += 1
print(f"\n{violations} violation(s)")
return 1 if violations else 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))

View File

@@ -1,34 +0,0 @@
---
name: incident-investigation
risk_class: read_only
inputs: [symptom, affected_entity]
verification: "n/a — investigation produces a written record, not a state change"
docs_update_checklist: [investigations_entry]
---
# Incident investigation
Goal: understand what broke and why, before touching anything.
1. `homelab service <name> explain` (or `homelab node <name> relations`
if the affected entity is a host) — get the blast radius and doc
pointer first. Don't start pulling logs blind.
2. `homelab service <name> health` + `homelab service <name> log` (or
MCP `get_service_status` / `tail_log`) for the affected service.
3. Walk the blast radius: is a shared dependency down (`caddy`, `dns`,
`authentik`, or the backend host itself)? `homelab node <name>
relations` shows "affected by" — check those first.
4. `homelab apt-audit` if the symptom looks like a dpkg/upgrade
interaction.
5. Check the change ledger for recent mutations to the affected entity
or anything upstream of it: `homelab service <name> history` (once
populated) or grep `ledger/*.jsonl`.
6. Write findings to a new `knowledge/sources/investigations/<date>-<slug>.md` — symptom,
timeline, root cause, fix applied, prevention. This is the durable
record; don't rely on chat history.
Docs-update checklist: always create the investigation entry. If the
root cause was stale/wrong inventory data (a `doc_page`, `config_repo`,
or `backend` that didn't match reality — this happened during Week 1
kernel work, see the `authentik` backend fix), correct `inventory.yaml`
in the same session.

View File

@@ -1,36 +0,0 @@
---
name: lifecycle-activate-node
risk_class: config_mutation
inputs: [node_name]
verification: "homelab service <name> health (if it hosts a service); homelab doctor (if it's a client)"
docs_update_checklist: [doc_page_complete]
transition: "provisioning -> active"
---
# Lifecycle: activate a node
Per [oikos/ontology.yaml](../../../oikos/ontology.yaml). Requires: age key
enrolled if it needs secrets, mesh joined if it needs off-LAN reach,
ingress live if public, health check answering, doc page complete,
ledger entry.
1. If the node is a `homelab` client: finish enrollment per
[client-enrollment.md](../client-enrollment/SKILL.md) (`--finalize-pubkey`,
mesh join, `homelab doctor` green).
2. If it hosts a public service: add the `services:` entry in
`inventory.yaml` (backend, url, doc_page, config_repo, risk_notes —
see the Week-1 service contract fields) and wire the Caddy route in
`dtoro/caddy-conf`.
3. Confirm the health check answers: `homelab service <name> health` or
a direct `curl`.
4. Flip `state: provisioning``state: active` (or delete the `state:`
field — `active` is the default) in `inventory.yaml`.
5. Complete the doc page (stub → full narrative: role, specs, how it's
configured, dependencies).
6. Record the activation: `oikos/ledger.py append host:<name> activate
config_mutation --result ok` (or let the CLI wrapper do this once
Week 3's runbook automation lands).
Regenerate derived data: `python3 mcp/build_host_files.py && python3
oikos/gen-topology.py` so `hosts/<name>.yaml`, the topology diagram, and
the context card all reflect the new state.

View File

@@ -1,35 +0,0 @@
---
name: lifecycle-deprecate-node
risk_class: config_mutation
inputs: [node_name, replacement_node_or_reason]
verification: "homelab node <name> relations — 'affected by' must be empty before completing"
docs_update_checklist: [doc_page_deprecation_note]
transition: "active -> deprecated"
---
# Lifecycle: deprecate a node
Per [oikos/ontology.yaml](../../../oikos/ontology.yaml): a node keeps running
but takes no new dependents. **Completion condition: zero remaining
inbound `depends-on`/`routes-to` edges** — this is a hard gate, not a
suggestion; `oikos/policy.yaml` `lifecycle_overrides.deprecated.refuse`
lists `new-inbound-edges` as refused going forward.
1. Set `state: deprecated` on the node.
2. `homelab node <name> relations` — read `affected_by`. Every entry
there is something still relying on this node.
3. Migrate or retire each dependent one at a time (point its `backend`/
`config_repo`/ingress route elsewhere, or deprecate it too if it's
being retired alongside).
4. Re-run `homelab node <name> relations` after each dependent is moved.
The transition to `destroyed` is only safe once `affected_by` is
empty — check this every time, don't assume from memory.
5. Note the deprecation on the doc page: reason, replacement (if any),
date.
If step 2 shows dependents you didn't expect, stop and investigate
before proceeding — that's exactly the kind of drift the Week-3 detector
will catch automatically, but until then this manual check is the gate.
Next (once `affected_by` is empty):
[lifecycle-destroy-node.md](../lifecycle-destroy-node/SKILL.md).

View File

@@ -1,42 +0,0 @@
---
name: lifecycle-destroy-node
risk_class: destructive
inputs: [node_name]
verification: "homelab node <name> relations returns unknown-entity; pct list on the backend no longer shows it"
docs_update_checklist: [archaeology_entry, containers_index_update]
transition: "deprecated -> destroyed"
---
# Lifecycle: destroy a node
**Destructive.** Requires operator approval + typed confirmation phrase
per `oikos/policy.yaml`. Requires (ontology): backups verified, secrets
recipients removed + re-keyed, ingress/DNS removed, archaeology entry,
ledger entry.
1. Confirm the node is `deprecated` with zero `affected_by` edges
(`homelab node <name> relations`) — do not skip this even if the
deprecation runbook was followed recently; state can drift.
2. If it's an enrolled client: `homelab client remove <name>` — revokes
the age key, re-keys SOPS, removes the inventory entry. This is
already destructive-class and confirmed in the CLI.
3. Remove any ingress route (Caddy config repo) and DNS record still
pointing at it.
4. Verify backups of anything on it are retained per policy before the
disk goes away (see `backs-up-to`).
5. Destroy the LXC/VM (`pct destroy` / `qm destroy`).
6. Move the `hosts.<name>:` block (if any inventory remnant survives
`client remove`, e.g. infra-only LXCs with no age key) into
inventory.yaml's `archaeology:` section: `pve_id`, `destroyed` date,
`reason`. Add a row to `containers/index.md` "Recently destroyed"
table (kept for human-readable browsing alongside the structured
data).
7. `oikos/ledger.py append host:<name> destroy destructive --result ok`.
8. Regenerate: `python3 mcp/build_host_files.py && python3
oikos/gen-topology.py` — the node drops out of `hosts/*.yaml` and
appears in the topology doc's archaeology table.
If the destroy fails partway (e.g. secrets revoked but pct destroy
errors), do not re-run step 2 — `client remove` is not idempotent
against a second revocation attempt on the issuance server. Finish the
remaining steps manually and note the partial state in an investigation.

View File

@@ -1,39 +0,0 @@
---
name: lifecycle-migrate-node
risk_class: config_mutation
inputs: [node_name, source_host, target_host]
verification: "homelab node <name> relations (re-check blast radius); homelab service <svc> health for every hosted service"
docs_update_checklist: [doc_page_migration_note, inventory_host_and_lan_ip]
transition: "active -> migrating -> active"
---
# Lifecycle: migrate a node
Modeled on the strong Phase 1+2 migration
([plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)).
Requires (ontology): preflight + backup-verified before migrating;
post-verify + Caddy backends checked + mounts checked + docs updated
before returning to `active`.
1. `homelab change preflight <every service the node hosts>` — capture
current health as a baseline.
2. Verify backups are current for anything with data at rest on the
node (see `backs-up-to` edges once populated).
3. Set `state: migrating` in `inventory.yaml`.
4. Perform the migration (pct/qm move, or create-on-target +
data-copy + destroy-source, per the specific case).
5. Update `inventory.yaml`: new `host:`, `lan_ip`, `mesh` addresses for
the node; update every `services:` entry whose `backend` pointed at
it if the backend name itself changes (usually it doesn't — only the
`host:`/`lan_ip` on the guest entry moves).
6. Post-verify: re-run the Week-1 drift check by hand — confirm Caddy's
backend IP for each affected service matches the new `lan_ip`
(automatic in Week 3's drift detector), confirm mounts still resolve.
7. `homelab service <name> health` for every service the node hosts.
8. Set `state: active`. Add a migration note to the node's doc page
(old host/IP → new, date, phase reference) — this repo's convention
for every past migration (see `containers/101-jellyfin.md`,
`containers/129-house.md`).
Regenerate: `python3 mcp/build_host_files.py && python3
oikos/gen-topology.py`.

View File

@@ -1,33 +0,0 @@
---
name: lifecycle-provision-node
risk_class: config_mutation
inputs: [node_name, kind, storage_pool]
verification: "grep 'state: provisioning' hosts/<name>.yaml"
docs_update_checklist: [doc_page_stub]
transition: "planned -> provisioning"
---
# Lifecycle: provision a node
Per [oikos/ontology.yaml](../../../oikos/ontology.yaml) `lifecycle.transitions`.
Policy note: `provisioning` nodes get a lifecycle override —
`config_mutation` actions downgrade to `reversible_low` because nothing
depends on the node yet (see `oikos/policy.yaml` `lifecycle_overrides`).
Requires (from ontology): inventory entry, IP reserved, storage pool
chosen, doc page stub.
1. Create the LXC/VM on its target Proxmox host (`pct create` /
`qm create`), choosing the storage pool deliberately — record it as
the `storage:` field once populated (Week 1 schema; not yet backfilled
for existing nodes).
2. Add the inventory entry: `homelab client add <name>` for anything that
will run the `homelab` CLI, or a direct `hosts.<name>:` block with
`state: provisioning`, `kind`, `host`, `pve_id`, `lan_ip` for
infra-only LXCs that won't self-enroll.
3. Stub the doc page (`containers/<pve_id>-<name>.md` or
`vms/<pve_id>-<name>.md`) — even a one-line "provisioning, see plan X"
is enough to satisfy the transition requirement.
4. Reserve the IP in DNS/DHCP notes if it's a fixed LAN address.
Next: [lifecycle-activate-node.md](../lifecycle-activate-node/SKILL.md).

View File

@@ -1,209 +0,0 @@
---
name: budget-import-from-csv
risk_class: config_mutation
inputs: [csv_file]
references: [containers/129-house.md]
---
# Runbook: Budget import from N26 CSV → Yuvomi
Distil a bank-export CSV into Yuvomi's Budget and Subscriptions modules using
the `yuvomi-mcp` tools. Run this whenever a new CSV period needs to be
summarised into targets and fixed costs.
---
## Prerequisites
- `yuvomi-mcp` is running on LXC 129 and connected as an MCP server in Claude.
- The CSV is an N26 export (columns: Booking Date, Value Date, Partner Name,
Partner Iban, Type, Payment Reference, Account Name, Amount (EUR), …).
- API token: `homelab secret yuvomi-api-token` (decrypts on any enrolled client).
- Direct API base: `https://house.hubris.network/api/v1`
---
## API quirks (Yuvomi ≤ 0.77.x)
- **Subscriptions live under `/budget/subscriptions`**, NOT `/subscriptions/`.
A top-level `/subscriptions` route returns 404.
- `GET /budget/subscriptions``{ data: { subscriptions: [...], summary: {...} } }`
- `GET /budget/subscriptions/meta``{ data: { categories: [...], payment_methods: [...] } }`
- `POST /budget/subscriptions` → create a subscription (name, amount, billing_cycle,
cycle_interval, next_payment_date, currency, category_id, payment_method_id required)
- `GET /budget/` (no month) → returns only **non-recurring** base entries.
Use `GET /budget/?month=YYYY-MM` to get all entries (recurring + one-time) for a month.
- `GET /budget/categories` → expense category keys + income category names (German keys
like `"Erwerbseinkommen"`, `"Sozialleistungen"`, `"Geschenke & Transfers"`).
- Budget entries: `amount` positive = income, negative = expense.
- Recurring entries: set `is_recurring: 1` + `recurrence_interval: "monthly"`.
The `date` field sets the start month.
- `recurrence_virtual: 1` smooths non-monthly amounts across all months in the summary
(e.g. 55.08 € quarterly → shows as ~18.36 €/month).
- Custom RRULE strings (`recurrence_rule`) are **not accepted** by the API — use
`cycle_interval` on the subscription instead, or `recurrence_interval` on budget entries.
---
## Subscription category IDs (as of 2026-06-26)
| id | name | budget_subcategory_key |
|---|---|---|
| 1 | Entertainment | subscription_entertainment |
| 2 | Productivity | subscription_productivity |
| 3 | Utilities | subscription_utilities |
| 4 | Health | subscription_health |
| 5 | Education | subscription_education |
| 6 | Other | subscription_other |
## Payment method IDs
| id | name |
|---|---|
| 1 | Credit Card |
| 2 | Debit Card |
| 3 | PayPal |
| 6 | Bank Transfer / SEPA |
| 7 | Other |
---
## Step 1 — Categorise the transactions
Skip these as internal/already-covered:
- Fixed costs you'll enter as **subscriptions** (Miete, SWM, SYNVIA, Hundefutter,
Netflix, Grover, Rundfunk ARD, KuKita, Lillydoo)
- Internal transfers (The Joy Pot ↔ Cookie, Hauptkonto, Tagesgeldkonto splits)
- Identified income (Cookie Share, Kindergeld, Pocket Money credits, Distributor)
- Fun Money pass-throughs (in and out same month → net zero)
**Variable expense taxonomy:**
| Category key | Subcategory key | Examples |
|---|---|---|
| `food` | `groceries` | E-Center, Knuspr, EDEKA, Tegut, VollCorner, Lidl, Netto, REWE, KoRo, Roast Market |
| `food` | `restaurants_bars` | Restaurants, Lieferando, Cafes, Zeit für Brot, Baobab, Höflinger |
| `personal_health` | `beauty_cosmetics` | DM, Rossmann |
| `personal_health` | `pharmacy` | Apotheke, MVZ Dermatologie |
| `transport` | `apps_taxi` | Uber, RYD GMBH, MVG, Handyparken |
| `shopping_clothing` | `gifts` | Children products: Schlummersack, Catchy Kids, SP EVERY., Dukal, Berger-Lernwelt |
| `shopping_clothing` | `clothes_shoes` | Zalando, Ernsting's, Schuhmair, Thalia, Vinted, Airbnb, Hotel at Booking.com |
| `shopping_clothing` | `electronics` | Amazon, AMZN Mktp DE |
| `housing` | `renovation_maintenance` | IKEA, Markus Festl, Granit, Sostrene Grene, Mol* tischdecken, Gaertnerei, Dehner |
| `education` | `courses_college` | Kathrin Orlob (PEKiP), Nerina Aupperle |
| `leisure` | `streaming` | WOW wowtv.de |
| `financial_other` | `bank_fees` | Unidentified PayPal, Ratepay, N26 fees |
| `Geschenke & Transfers` | *(income)* | One-off incoming transfers |
---
## Step 2 — Create subscriptions
```
get_subscriptions_meta() ← get category_id and payment_method_id
```
**Standard Cookie household subscriptions (as of 2026-07):**
| Name | Amount | billing_cycle | cycle_interval | category_id | payment_method_id |
|---|---|---|---|---|---|
| Miete | 1080.00 | monthly | 1 | 6 (Other) | 6 (Bank Transfer) |
| Strom (SWM) | 79.00 | monthly | 1 | 3 (Utilities) | 6 |
| Internet / TV / Telefon | 29.99 | monthly | 1 | 3 (Utilities) | 6 |
| Hundefutter | 75.00 | monthly | 1 | 6 (Other) | 6 |
| Netflix | 8.00 | monthly | 1 | 1 (Entertainment) | 6 |
| Grover | 16.90 | monthly | 1 | 6 (Other) | 2 (Debit Card) |
| Rundfunk ARD / ZDF | 55.08 | monthly | 3 | 1 (Entertainment) | 6 |
| KuKita Daycare (Leon) | 503.00 | monthly | 1 | 5 (Education) | 6 |
| Lillydoo diapers | 56.70 | monthly | 2 | 4 (Health) | 3 (PayPal) |
Monthly equivalent total: **1,838.60 €** (Yuvomi applies cycle_interval to prorate).
---
## Step 3 — Add recurring income entries
```
stage_add_budget_entry(
title="Kindergeld",
amount=55.00,
category="Sozialleistungen",
date="YYYY-MM-01",
is_recurring=True,
recurrence_interval="monthly",
)
commit_pending(pending_id)
```
**Standard recurring income:**
| Title | Amount | category |
|---|---|---|
| Kindergeld | +55.00 | Sozialleistungen |
| Cookie Share | +2650.00 | Erwerbseinkommen *(see recommended amount below)* |
---
## Step 4 — Post variable transactions
For each non-skipped CSV row, call `stage_add_budget_entry` with the mapped
category/subcategory and the actual transaction amount and date. Use the Partner
Name + Payment Reference as the title (truncate to 100 chars).
---
## Step 5 — Verify
```
get_budget_summary("YYYY-MM")
list_subscriptions()
```
Expected for a full month with KuKita:
- Fixed expenses ≥ 1,838 € (subscriptions)
- Variable expenses ≥ 500 € (groceries alone)
---
## Cookie Share: how much to transfer monthly
Calculated from JanJun 2026 data (Cookie account, one-offs stripped):
| | €/month |
|---|---|
| **Fixed costs (subscriptions)** | **1,839** |
| Miete | 1,080 |
| KuKita *(permanent from Jul 2026)* | 503 |
| Strom + SYNVIA + Rundfunk + Netflix + Grover + Hundefutter + Lillydoo | 256 |
| **Variable (6-month averages)** | **1,032** |
| Groceries | 595 |
| Children products | 142 |
| Dining & cafes | 100 |
| Transport | 66 |
| Drugstore | 52 |
| Clothing, Amazon, Pharmacy | 77 |
| **Total monthly spend** | **≈ 2,871** |
| Minus Kindergeld (fixed income) | 55 |
| Minus Pocket Money (conservative ~600 €) | 600 |
| **→ Recommended Cookie Share** | **≈ 2,650 €** |
| With 200 € buffer | **≈ 2,850 €** |
**Current Cookie Share (Jun 2026): 1,995 € — shortfall ~655 €.**
The gap was covered by irregular Pocket Money top-ups (avg 962 €/mo over 6 months, but
highly variable: 121 €3,000 €). KuKita starting in June is the biggest step-up; raising
Cookie Share to **2,650 €** makes the budget self-sufficient without relying on top-ups.
---
## Changelog
### 2026-06-29 — Corrections from first real import
- Subscriptions endpoint is `/budget/subscriptions`, NOT `/subscriptions/` (404).
- `recurrence_rule` RRULE strings are rejected by the API; use `cycle_interval` instead.
- `GET /budget/` (no filter) returns only non-recurring entries; use `?month=` for full view.
- Added Cookie Share recommendation (2,650 €/month) based on 6-month expense analysis.
- Added full category taxonomy table.
### 2026-06-29 — Initial runbook
Created from JanJun 2026 N26 Cookie account analysis.

View File

@@ -1,30 +0,0 @@
---
name: service-health-check
risk_class: read_only
inputs: [service_name]
verification: "homelab service <name> health"
docs_update_checklist: []
---
# Service health check
Goal: determine whether a service is actually healthy, without ad-hoc SSH.
1. `homelab service <name> explain` — read the context card: backend,
blast radius, doc pointer, risk notes.
2. `homelab service <name> health` — live health probe (HTTP code against
the service's `url`/`endpoint`). Once the Week-3 scheduler ships, this
reads a cached snapshot by default; pass `--live` to force a fresh probe.
3. If unhealthy, `homelab service <name> log` (or MCP `tail_log`) for the
last 200 lines.
4. Cross-check blast radius: `homelab node <name> relations` — is this
entity's own backend host healthy? A downstream failure (e.g. `strong`
down) will show up here before the service's own logs explain anything.
5. If the fix is a restart: classify first (`oikos/policy.yaml`
`service-restart` is `reversible_low` unless the service has a
`service_overrides` entry, e.g. `caddy`/`dns` are `config_mutation`).
Unattended agents may act on `reversible_low` without approval.
Docs-update checklist: none for a pure health check. If the investigation
reveals stale `risk_notes` or a wrong `doc_page`, fix `inventory.yaml` in
the same session.

View File

@@ -1,72 +0,0 @@
# Oikos CI (Gitea Actions). Gates the deploy webhook on a green run (plan M1).
# Mirrors `make lint`, `make test`, and the generated-code drift guard.
name: ci
on:
push:
branches: [main]
pull_request:
jobs:
build-test:
runs-on: ubuntu-latest
services:
postgres:
image: timescale/timescaledb:2.17.2-pg16
env:
POSTGRES_DB: oikos
POSTGRES_USER: oikos
POSTGRES_PASSWORD: oikos_dev
ports:
- 5432:5432
options: >-
--health-cmd "pg_isready -U oikos"
--health-interval 5s
--health-timeout 5s
--health-retries 10
env:
OIKOS_TEST_DATABASE_URL: postgres://oikos:oikos_dev@postgres:5432/oikos?sslmode=disable
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version: "1.26"
cache: true
- name: go vet
run: go vet ./...
- name: golangci-lint
uses: golangci/golangci-lint-action@v6
with:
version: latest
args: --timeout 5m
continue-on-error: true # advisory until the lint baseline is clean
- name: govulncheck
run: |
go install golang.org/x/vuln/cmd/govulncheck@latest
govulncheck ./... || true # advisory
- name: generated code is up to date
run: make generate-check
- name: build
run: go build ./...
- name: test (race + coverage)
run: go test -race -covermode=atomic -coverprofile=coverage.out -timeout 300s ./...
- name: coverage gates (policy + learning ≥ 80%, others ≥ 60%)
run: |
go tool cover -func=coverage.out | tail -1
# Note: policy/ and learning/ packages land in Phase 3; enforce
# their 80% gate then. For now, report total coverage.
docker-build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: docker build (verify image builds; no push)
run: docker build -f compose/oikos/Dockerfile -t oikos:ci .

10
.gitignore vendored
View File

@@ -1,10 +0,0 @@
.DS_Store
__pycache__/
*.pyc
# Regenerated every scheduler run (every 10 min); no audit value in the
# diff. Signals (signals/*.jsonl) ARE tracked — this is just the ephemeral
# health-probe cache. See oikos/scheduler.py.
oikos/state.json
.worktrees/

View File

@@ -1,277 +0,0 @@
# Plan: Migrate library SSD to ludo-mini + Proxmox gaming/media server
## Goal
Split the homelab into two Proxmox hosts:
| Host | Role | Storage |
|------|------|---------|
| **hubris** | Core services (reverse-proxy, SSO, Matrix, git, documents, HA) | SSD 1 — boot + LXC rootfs (unchanged) |
| **ludo-mini** | Gaming server + media/library services | SSD 2 — Samsung 990 EVO Plus 4 TB (moved from hubris) |
The library SSD physically moves from hubris to ludo-mini. hubris LXCs that still need `/mnt/library` access it over NFS from ludo-mini.
## Current state
### hubris hardware
- GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS, 12 vCPU, ~28 GiB RAM
- 2× Samsung 990 EVO Plus NVMe:
- nvme0: `local` (95G) + `local-lvm` (856G) — boot, ISOs, LXC rootfs
- nvme1: `library` LVM (3.7T) — `/mnt/library` ext4 via `/dev/mapper/library-library`
### LXCs binding `/mnt/library` (host-level bind-mount)
| ID | Name | Role | I/O profile |
|----|------|------|-------------|
| 101 | jellyfin | Media streaming | Read-heavy, sequential |
| 103 | paperless | Document archive | Mixed, OCR writes |
| 104 | gitea | Git server | Mixed, lots of small files |
| 105 | apps | Docker (booklore, audiobookshelf, artifacto, MCP) | Mixed, depends on container |
| 114 | nextcloud | File sync | Mixed, WebDAV |
| 119 | sophia | Workshop | Low I/O |
| 120 | mule-images | Photo management | Write-heavy (processing), iGPU |
| 122 | arriman | *arr stack + downloads | Write-heavy (downloads) |
| 126 | plato | App (sub-mount: `/mnt/library/documents/plato`) | Light |
### NFS export chain (for VM 100 zimaos)
```
/mnt/library (ext4, host) → bind-mount → LXC 102 (nfs-export) → NFSv4 → VM 100 (zimaos)
```
### ludo-mini current
- Linux workstation, wired Ethernet 2.5 Gbps, `192.168.178.181` (household LAN)
- Runs Sunshine for game streaming
- No Proxmox, no LVM config
- Connected to SODOLA switch (same switch as hubris eno1)
### Network topology
```
Fritz!Box 7590 (192.168.178.1)
└── SODOLA 2.5G switch
├── hubris eno1 → vmbr1 (192.168.178.10)
│ └── routes to vmbr0 (192.168.8.0/24) — all LXCs
└── ludo-mini (192.168.178.181)
```
hubris routes between `192.168.8.0/24` (vmbr0) and `192.168.178.0/24` (vmbr1). So LXCs can reach ludo-mini via hubris as a router.
## Key decisions
### 1. Service split — what moves, what stays
**Move to ludo-mini** (high I/O, benefits from data locality + GPU):
- 101 jellyfin — media streaming, GPU transcoding
- 120 mule-images — photo processing, iGPU passthrough
- 122 arriman — *arr stack, downloads write to library
**Stay on hubris, NFS-mount library from ludo-mini:**
- 103 paperless — documents, moderate I/O
- 104 gitea — git repos (small files, some I/O sensitivity but acceptable over NFS)
- 105 apps — Docker apps, mixed workloads
- 114 nextcloud — file sync
- 119 sophia — workshop, light use
- 126 plato — app, light use
- 100 zimaos — NAS frontend, already NFS-mounted
### 2. NFS architecture
Instead of changing every LXC's mount config, keep the bind-mount pattern on hubris:
```
ludo-mini: /mnt/library (ext4, local NVMe)
└── NFSv4 export to 192.168.8.0/24
└── hubris host: NFS-mount at /mnt/library
└── LXCs: bind-mount /mnt/library (unchanged!)
```
This is transparent to all hubris LXCs — no container config changes needed. Only the hubris host changes from ext4 local mount to NFS mount. The LXC bind-mounts "just work" because `/mnt/library` is still at the same path on the host.
### 3. Network — ludo-mini reachability from hubris LXCs
LXCs on `192.168.8.0/24` reach ludo-mini (`192.168.178.181`) through hubris routing:
- `vmbr0` → hubris kernel routing → `vmbr1` → SODOLA → ludo-mini
- Already works (IP forwarding enabled on hubris)
**Alternative (cleaner):** Add a secondary IP `192.168.8.x` on ludo-mini's physical interface so it's directly on the homelab subnet. This avoids the router hop and keeps NFS traffic off kernel forwarding path. Worth considering but not required.
### 4. Gaming on ludo-mini with Proxmox
ludo-mini runs Sunshine (game streaming). Under Proxmox:
- **Option A:** Gaming VM with GPU passthrough — Sunshine + games in a VM, full GPU access
- **Option B:** LXC with GPU device passthrough (`/dev/dri`) — lighter, shares kernel
- **Option C:** Keep Sunshine on the Proxmox host itself (not recommended, but simplest)
Option A is the cleanest for isolation. Games need a full desktop environment and GPU drivers; a VM with GPU passthrough gives them that.
### 5. What about nfs-export (LXC 102)?
Currently exports `/mnt/library` to zimaos. After migration:
- If zimaos stays on hubris and accesses library via host NFS → bind-mount → LXC 102, that's triple-hop (ludo-mini → NFS → hubris → bind-mount → LXC 102 → NFS → zimaos). Terrible.
- Better: zimaos NFS-mounts directly from ludo-mini.
- So LXC 102 gets decommissioned (or repurposed).
- zimaos gets a new NFS mount pointing directly at ludo-mini.
## Migration phases
### Phase 1 — Preparation (no downtime)
1. **Document current state** on hubris:
- `pct list` — full container inventory
- `pct config <id>` for every library-mounting LXC
- `cat /etc/fstab` — capture the library mount line
- `df -h /mnt/library` — confirm space usage
- `lsblk -f` — UUID, filesystem
- Identify the exact NVMe device (`nvme1n1`)
2. **Pre-flight on ludo-mini:**
- Confirm hardware: CPU, RAM, available M.2 slots, GPU model
- Confirm it can take the Samsung 990 EVO Plus (M.2 NVMe, PCIe 4.0 x4)
- Verify BIOS supports virtualization (VT-d/AMD-Vi for PCIe passthrough)
- Check: does ludo-mini have a second drive for Proxmox OS? If not, we need to partition the library SSD for Proxmox boot + library LVM, which complicates things significantly
3. **Install Proxmox on ludo-mini:**
- Download Proxmox VE 9.x ISO
- Install to ludo-mini's system drive (NOT the library SSD)
- Configure networking: bridge for Proxmox, IP on 192.168.178.x
- Test: web UI accessible
4. **Prepare NFS server on ludo-mini Proxmox:**
- Create NFS-export LXC (or serve from host — simpler for now)
- Prepare `/etc/exports`: `192.168.8.0/24(rw,all_squash,anonuid=33,anongid=10000,no_subtree_check,sec=sys)`
- Same squash params as current nfs-export LXC 102
### Phase 2 — Physical SSD move (planned downtime)
1. **Graceful shutdown on hubris:**
- Stop all library-mounting LXCs (101, 103, 104, 105, 114, 119, 120, 122, 126)
- Unmount `/mnt/library` on hubris host
- Edit `/etc/fstab` to comment out the library mount line
- Power off hubris
2. **Physical drive swap:**
- Remove Samsung 990 EVO Plus (library SSD) from hubris
- Install into ludo-mini M.2 slot
- Power on ludo-mini
3. **Bring library online on ludo-mini:**
- Detect the new NVMe device
- If it's the whole device with LVM, activate the VG:
```
vgscan && vgchange -ay library
mount /dev/mapper/library-library /mnt/library
```
- Add to `/etc/fstab` for auto-mount
- Verify content: `ls /mnt/library` — same tree as before
4. **Start NFS export on ludo-mini:**
- `exportfs -ra`
- Verify: `showmount -e <ludo-mini-ip>`
### Phase 3 — Reconnect hubris LXCs
1. **Power on hubris** (without library SSD — it'll boot fine, just won't mount library)
2. **Mount NFS on hubris host:**
- Install `nfs-common` if not present
- Add to `/etc/fstab`:
```
192.168.178.181:/mnt/library /mnt/library nfs rw,vers=4,soft,timeo=30,retrans=3 0 0
```
Use `soft` to prevent hangs if ludo-mini is down; `hard` with `intr` is safer for data integrity but can block processes.
- `mount /mnt/library`
- **Verify permissions:** `ls -la /mnt/library` — should show `www-data:media` ownership for shared subtrees (same uid 33, gid 10000). The NFS all_squash guarantees this.
3. **Start LXCs:**
- Start the LXCs that stayed on hubris (103, 104, 105, 114, 119, 126)
- Their bind-mounts should work — `/mnt/library` is populated via NFS
- Verify each service: web UIs, git clone, document access
4. **Update DNS/Caddy** for services that moved:
- If jellyfin, arr services moved to ludo-mini, update Caddyfile to point to ludo-mini IPs
- Update DNS entries if needed
### Phase 4 — Migrate services to ludo-mini
1. **Create LXCs/VM on ludo-mini Proxmox:**
- 101 jellyfin — privileged LXC, mount `/mnt/library`, add `media` group
- 120 mule-images — privileged LXC, mount `/mnt/library` + `/dev/dri` passthrough
- 122 arriman — privileged LXC, mount `/mnt/library`
2. **Migrate configs:**
- Copy LXC configs from hubris (`/etc/pve/lxc/<id>.conf`) as templates
- Adjust network (IPs on household subnet or Proxmox bridge)
- Restore app data from backups or copy over NFS
3. **Gaming VM:**
- Create VM with GPU passthrough
- Pass through the dGPU for gaming performance
- Install Sunshine + game libraries
- Storage: VM disk on Proxmox storage, games on library SSD
4. **Update reverse proxy:**
- Caddy on hubris (121): update backend IPs for jellyfin, jellyseerr, qbit, sab, mule-images → point to ludo-mini
- Test: `media.hubris.network` serves from ludo-mini jellyfin
### Phase 5 — Cleanup
1. **Decommission nfs-export LXC 102** on hubris (no longer needed)
2. **Update zimaos (100)** — change NFS mount from `192.168.8.200` → `192.168.178.181`
3. **Remove old LXCs** from hubris (101, 120, 122) after confirming migration works
4. **Update inventory.yaml:**
- ludo-mini: `kind: proxmox-host`, add mounts/storage, add LXCs
- Move services from hubris to ludo-mini
- Remove nfs-export
5. **Update DNS:** `nfs-export.hubris.network` → ludo-mini IP (or remove)
6. **Run `homelab sync`** to propagate changes
## Open questions / unknowns
1. **Does ludo-mini have a second drive for Proxmox OS?** If not, we'd need to repartition the library SSD — carve out ~100 GB for Proxmox, then the rest for library LVM. This is risky (data loss if partitioning goes wrong) and requires a full backup first. **Alternative:** Buy a small SSD for ludo-mini's OS.
2. **What GPU does ludo-mini have?** Proxmox GPU passthrough requires IOMMU support and a GPU that doesn't have the reset bug. Need to check the exact GPU model.
3. **NFS performance for git (gitea)?** Git operations over NFS can be problematic (locking, stat() storms). Gitea bare repos at `/mnt/library/repos/*.git` might need testing. Worst case: move gitea's repo storage to local disk and keep `/mnt/library` for large file/LFS storage only.
4. **Media permission drift.** NFS `all_squash,anonuid=33,anongid=10000` ensures all writes from hubris LXCs (over NFS) and ludo-mini LXCs (local) land as `www-data:media`. This is the same squash currently used by nfs-export (102). Should be fine.
5. **ludo-mini network — add 192.168.8.x address?** Adding a secondary IP on ludo-mini's interface directly on the homelab subnet avoids routing through hubris for NFS traffic. Cleaner, but requires Proxmox bridge setup. Worth doing during Proxmox install.
6. **Sunshine migration.** Currently runs on ludo-mini bare metal. After Proxmox install, it needs to run in a VM. What happens to existing Sunshine configs, game libraries, save files? Need to preserve these during the Proxmox install.
7. **Backup before moving.** The library SSD holds 3.7 TB of irreplaceable data (documents, photos, repos). Restic backups are currently disabled. **Before physically moving the drive, verify the data is readable and consider doing one backup** — or at minimum, `rsync` critical directories.
## Files affected
| File | Change |
|------|--------|
| `/opt/homelab-context/inventory.yaml` | ludo-mini: workstation → proxmox-host; add LXCs, mounts; remove nfs-export; move service backends |
| `/opt/homelab-context/hosts/hubris.md` | Remove library storage, add NFS mount note |
| `/opt/homelab-context/hosts/ludo-mini.yaml` | Complete rewrite — Proxmox host, storage, tenants |
| `/opt/homelab-context/containers/102-nfs-export.md` | Mark decommissioned |
| `/opt/homelab-context/containers/index.md` | Move 101, 120, 122 to ludo-mini; remove 102 |
| `/opt/homelab-context/infrastructure/dns.md` | Update nfs-export entry |
| `/opt/homelab-context/infrastructure/media-permissions.md` | Note NFS squash from ludo-mini, not hubris |
| hubris `/etc/fstab` | Replace ext4 mount with NFS mount |
| ludo-mini `/etc/fstab` | Add library ext4 mount |
| ludo-mini `/etc/exports` | Add NFS export config |
| caddy (LXC 121) Caddyfile | Backend IPs for moved services |
| DNS (LXC 107 Technitium) | Update entries for moved services |
## Validation checklist
- [ ] ludo-mini Proxmox web UI accessible
- [ ] Library SSD detected and mountable on ludo-mini
- [ ] NFS export from ludo-mini: `showmount -e <ip>` shows `/mnt/library`
- [ ] hubris host NFS mount: `df -h /mnt/library` shows NFS, not ext4
- [ ] hubris LXCs start and bind-mount /mnt/library (content visible)
- [ ] gitea: `git clone` over SSH works, repos readable
- [ ] paperless: document ingestion works, OCR processing
- [ ] nextcloud: file sync, WebDAV
- [ ] jellyfin: media plays from ludo-mini, transcoding works
- [ ] arriman: downloads write to library, jellyfin picks up new media
- [ ] mule-images: photo import and processing
- [ ] zimaos: NFS mount from ludo-mini works, Files UI shows library
- [ ] Sunshine: game streaming from ludo-mini VM works
- [ ] All `*.hubris.network` services resolve and load through Caddy

View File

@@ -1,271 +0,0 @@
# Homelab structure revision & improvement plan
## Goal
Identify structural issues in the current hubris homelab topology and propose an
actionable improvement roadmap — DNS consolidation, monitoring gaps, backup
recovery, mesh completion, resource rightsizing, and operational hygiene.
---
## Current state summary
| Dimension | Status |
|-----------|--------|
| Hypervisor | hubris (single-node Proxmox) — also does subnet routing (192.168.8.0/24) |
| LXCs | 14 active, 1 retired (124), 1 new (107 dns) |
| VMs | HAOS (108), ZimaOS (100) |
| Workstations | mac-mini (macOS), republic-laptop, ludo-mini |
| VPS | 1 IONOS box — netbird mgmt+signal+relay, traefik, authentik, coturn |
| Switch | SODOLA 5-Port 2.5Gbit (L2, flat) |
| Networking | 192.168.8.0/24 internal, fr!tz box main LAN via vmbr0→vmbr1 routing |
| Mesh | Netbird + Tailscale (migrating), ~mixed state |
| DNS | 3 sources: Technitium (CT 107), NetBird managed zone, public IONOS |
| Backups | DISABLED since 2026-04-22 |
| Monitoring | claudio-monitor (host health only, no apt/docker checks yet) |
| Agent enrollment | 3/19 hosts enrolled (hubris, apps, republic-laptop) |
---
## Issues identified
### 1. Three overlapping DNS sources (highest risk)
**Problem:** Technitium on CT 107, NetBird managed DNS zone, and the public
IONOS wildcard all answer `*.hubris.network` queries. The dns.md doc still
references the old dnsmasq on LXC 124 (though the change log says it moved).
NetBird's managed DNS bypasses Technitium entirely for some app names — there
is no single source of truth for DNS.
**Risk:** Mismatched answers → services unreachable → "works on some clients
but not others" debugging sessions. Already cost time when `auth.hubris.network`
re-pointed to the VPS.
**Proposal:**
- Phase out NetBird managed DNS zone for `hubris.network` — Technitium is the
authoritative answerer for mesh & LAN clients
- Set Technitium as the sole DNS for all LXCs (remove router DNS / Tailscale
MagicDNS fallbacks)
- Document the full authoritative chain: Technitium → upstream forwarders → public
- Track Technitium config in git (dtoro/technitium-config or equivalent)
### 2. Backups disabled with no alternative (data loss risk)
**Problem:** The only backup was restic to an external USB that caused host
crashes. It was disabled 2026-04-22 as an A/B test — host stability was
confirmed by the SODOLA migration (no more Slate AX double-NAT hangs). The
drive is still removed.
**Risk:** `/mnt/library` (~429 GB of irreplaceable data: photos, documents,
git repos, gitea data) has no off-host copy. Single-disk failure = total loss.
**Proposal:**
- Re-evaluate the USB drive stability with the new SODOLA switch topology
(direct rear USB 3.0 port, no hub chain)
- OR adopt a cloud-backed strategy: `rclone`-to-Hetzner Storage Box or
Backblaze B2 for the irreplaceable subset (docs, photos, gitea data)
- OR use hubris's own `zfs send` to a second host/disk if ZFS is feasible
- Minimum viable: at minimum restore gitea backups + sops-encrypted secrets
via an off-site cron (cheap B2 bucket)
### 3. Mesh migration still incomplete
**Problem:** Most LXCs still use Tailscale. This forces per-LXC DNS workarounds
(/etc/hosts overrides, local dnsmasq) and runs two VPN stacks in parallel.
Mesh migration doc (mesh.md) is comprehensive but execution stalled.
**Proposal:**
- Batch-migrate all LXCs from Tailscale to Netbird in one maintenance window
- Remove Tailscale from the PVE host
- Ensure all LXCs resolve `*.hubris.network` via Technitium (no more /etc/hosts
overrides)
- Document Netbird client on each LXC (netbird version, setup key rotation)
### 4. LXC resource imbalance & disk pressure
**Problem:**
| LXC | Cores | RAM | Rootfs | Disk usage |
|-----|-------|-----|--------|------------|
| mule-images (120) | 6 | 12 GiB | 60 GiB | photo AI, reasonable |
| arriman (122) | 4 | 8 GiB | 24 GiB | media stack, OK |
| nextcloud (114) | 4 | 6 GiB | 25 GiB | file sync, OK |
| apps (105) | 2 | 4 GiB | 30 GiB | 6+ services, tight |
| paperless (103) | 2 | 3 GiB | 8 GiB | 86.9% disk |
| elementsynapse (118) | 1 | 2 GiB | 8 GiB | 86.8% disk |
| gitea (104) | 1 | 1 GiB | 8 GiB | git server, adequate |
| caddy (121) | 1 | 512 MiB | 6 GiB | fine for reverse proxy |
Paperless (103) and elementsynapse (118) are at critical disk levels. apps (105)
is undersized for 6+ services.
**Proposal:**
- Resize rootfs on paperless (8→16 GiB) and elementsynapse (8→16 GiB)
- Bump apps (105) to 4 cores / 6 GiB RAM / 40 GiB rootfs
- Enable claudio-monitor's disk check to alert before next crisis
### 5. VPS is a single point of failure
**Problem:** One IONOS VM runs netbird management (control plane), traefik
(public ingress), authentik (identity), and coturn (TURN relay). If it goes
down: no remote mesh, no public services, no auth.
**Proposal:**
- Document a VPS recovery runbook (how to restore from a known-working backup)
- Consider splitting authentik into a separate host or at minimum having a
standby configuration
- Not a high priority (the VPS has been stable) but worth documenting the
blast radius and recovery path
### 6. No centralized logging
**Problem:** Each LXC has independent journald. Cross-service debugging
involves hopping between `pct exec <id> -- journalctl -u <service>`. There is
no aggregation or retention.
**Proposal:**
- Deploy a lightweight log shipper (Loki + promtail, or vector.dev) on each LXC
- Ship logs to a central Loki instance on apps (105) or a new small LXC
- Grafana dashboard optional — even a simple `logcli` query saves time
### 7. Agent enrollment incomplete
**Problem:** Only hubris, apps, and republic-laptop are enrolled in the
homelab-context system (age keys, sync timers, MCP access). mac-mini,
ludo-mini, claudio-bot, and all other LXCs are not.
**Proposal:**
- Batch-enroll remaining LXCs (jellyfin, paperless, gitea, nextcloud, etc.)
- Enroll mac-mini (macOS — exercises the launchd timer path)
- Enroll ludo-mini (needs SSH user config in inventory first)
- Wire claudio-bot into inventory-aware queries
### 8. Configuration drift on untracked configs
**Problem:** Technitium config, dnsmasq (legacy), and several service-specific
configs are not git-tracked.
**Proposal:**
- Track Technitium zone backup + compose config in a git repo
- Apply the same pattern as caddy-conf: dtoro/technitium-conf with auto-deploy
### 9. No capacity planning / resource monitoring
**Problem:** No trend data on CPU, RAM, or disk growth. The 86% disk alerts
were discovered reactively. Rootfs resize is painful (requires Proxmox stop +
resize + growfs inside).
**Proposal:**
- Enable the missing claudio-monitor checks (disk growth trend, apt upgradable
counts, docker image drift)
- Set up a simple Prometheus + node_exporter on hubris or use the PVE API
directly
- At minimum, surface disk usage in the existing homelab-mcp management tools
### 10. No standard deploy / orchestration for bare-metal LXCs
**Problem:** Some services are bare-metal CLI apps (sophia, claudio-bot),
some are Docker on apps (105), some are Portainer-managed. No consistent
deploy pattern means every new service reinvents the deployment.
**Proposal:**
- Don't over-engineer this — the current pragmatism works
- Just document the decision tree:
- Needs `/mnt/library` mount + heavy I/O → dedicated LXC
- Small stateless web service → Docker on apps (105)
- Media stack → dedicated LXC (arriman, jellyfin)
- Everything else → judge by complexity
---
## Phased implementation plan
### Phase 1 — Critical fixes (this week)
1. Resize rootfs on paperless (103) 8→16 GiB, elementsynapse (118) 8→16 GiB
2. Enable claudio-monitor disk check + disk-growth alerting
3. Pick one backup strategy and implement minimum viable (e.g. nightly
gitea dump + sops-encrypted secrets to B2 via rclone)
4. Verify Technitium is the sole DNS for all LXCs (remove NetBird managed zone
for hubris.network)
### Phase 2 — Mesh consolidation (next week)
5. Batch-migrate remaining LXCs from Tailscale to Netbird
6. Remove Tailscale from PVE host
7. Remove all per-LXC /etc/hosts DNS overrides
8. Update DNS documentation to reflect Technitium as single source
### Phase 3 — Agent enrollment & logging (next 2 weeks)
9. Enroll all LXCs in homelab-context (age keys, sync timers)
10. Enroll mac-mini (macOS launchd path — exercises untested code path)
11. Enroll ludo-mini
12. Deploy log shipper (Loki + promtail) on apps (105) + all LXCs
### Phase 4 — Resource & monitoring hardening (next month)
13. Resize apps (105) rootfs, bump RAM
14. Deploy Prometheus + node_exporter or equivalent for trend data
15. Track Technitium config in git with auto-deploy
16. Write VPS recovery runbook
### Phase 5 — Drive re-evaluation (optional, behind host-stability gate)
17. Re-attach USB backup drive with the new SODOLA topology (direct port)
18. If stable for 7 days, re-enable restic backup schedule (chunked)
19. If not stable, finalize cloud backup as permanent strategy
---
## Files likely to change
| Path | Change |
|------|--------|
| `/opt/homelab-context/inventory.yaml` | LXCs enrolled, resource updates |
| `/opt/homelab-context/secrets/*.yaml` | New recipients for enrolled LXCs |
| `/opt/homelab-context/infrastructure/dns.md` | Reflect Technitium as sole source |
| `/opt/homelab-context/infrastructure/mesh.md` | Remove Tailscale references post-migration |
| `/opt/homelab-context/infrastructure/backups.md` | New strategy |
| `/opt/homelab-context/infrastructure/monitoring.md` | Enable missing checks |
| `/opt/homelab-context/containers/103-paperless.md` | Rootfs resize |
| `/opt/homelab-context/containers/118-elementsynapse.md` | Rootfs resize |
| `/opt/homelab-context/containers/105-apps.md` | Resource bump, logging addition |
| `/opt/homelab-context/containers/index.md` | Updated resource table |
| `.sops.yaml` | New age pubkeys for enrolled LXCs |
## Verification
Each phase ends with a verification milestone:
- Phase 1: `claudio-monitor` triggers on paperless disk → confirmed alert. Backup
of gitea data lands in B2 (or equivalent). DNS query from any LXC returns
Technitium answer.
- Phase 2: `netbird status` shows all LXCs connected. Tailscale not running on
PVE host. `curl auth.hubris.network` from any LXC resolves correctly without
/etc/hosts.
- Phase 3: Every LXC has `/opt/homelab-context/` + `/etc/age/key.txt`. MCP
tools return valid host info for all enrolled LXC names. `journalctl` shows
promtail shipping to Loki.
- Phase 4: `claudio-monitor` shows disk growth trend. apps (105) can run all
6+ services without OOM.
## Risks & tradeoffs
- **Netbird migration window:** All LXCs will briefly lose mesh connectivity
during the Tailscale→Netbird cutover. Schedule in off-hours.
- **Backup cost:** B2/e2 costs ~$5/month for ~500 GB. The USB drive was free
but unstable — trade money for reliability.
- **DNS consolidation:** Removing the NetBird managed DNS zone means any
NetBird-specific names stop resolving for hubris.network — verify nothing
depends on that path.
- **Loki on apps (105):** Adds another container to an already-loaded host.
May need to bump resources before deploying.
- **Agent enrollment on every LXC:** Each enrollment creates an age keypair
and commits a pubkey to inventory. Process is scriptable via `homelab client
add` but still takes ~2 min per host for verification.
## Open questions
1. Is the USB backup drive still physically attached to hubris? If not, the
simplest "re-enable" path requires physically re-attaching it.
2. Authentik is now on the VPS — is LXC 124 (old Authentik) still running or
was it fully decommissioned? The dns.md changelog says "shut down" but
index.md lists it as "running".
3. What's the actual disk layout on hubris? nvme0n1, ZFS pool, mount structure
— needed to plan rootfs resizes safely.
4. Does the user want to keep Tailscale on any host for a specific reason, or
is full Netbird migration the clear goal?

View File

@@ -1,246 +0,0 @@
# Plan: Narrow Technitium DHCP Pool to Avoid Static-IP Conflicts
> **For Hermes:** Use subagent-driven-development skill to implement this plan task-by-task.
**Goal:** Eliminate the IP conflict risk created by the Technitium DHCP pool (`.100.240`) overlapping with all static LXC/VM IPs (`.101.239`).
**Architecture:** Shrink the DHCP pool range on Technitium so it only covers IPs that no static host uses. No LXC/VM IPs change. Single server-side change (Technitium API), plus documentation updates.
**Tech Stack:** Technitium DNS API (`/api/dhcp/scopes/set`), bash/curl, homelab-context repo for docs.
---
## Problem statement
The Technitium DHCP server on [CT 107](../../knowledge/wiki/containers/107-dns.md) serves `192.168.8.100192.168.8.240`. **Every static homelab IP except hubris (`.77`) sits inside that range:**
| Host | IP | Inside pool? |
|---|---|---|
| hubris (Proxmox) | .77 | No — below `.100` |
| haos (VM 108) | .101 | YES |
| gitea (104) | .121 | YES |
| paperless (103) | .130 | YES |
| arriman (122) | .132 | YES |
| mule-images (120) | .136 | YES |
| sophia (119) | .157 | YES |
| mac-mini | .174 | YES |
| caddy (121) | .175 | YES |
| authentik (124) | .180 | YES |
| plato (126) | .190 | YES |
| zimaos (VM 100) | .195 | YES |
| nfs-export (102) | .200 | YES |
| apps (105) | .205 | YES |
| jellyfin (101) | .206 | YES |
| nextcloud (114) | .224 | YES |
| claudio-bot (123) | .230 | YES |
| elementsynapse (118) | .239 | YES |
The docs claim "Static-IP LXCs (below `.100`) are unaffected" — this is **false**. Static IPs span `.101.239`, the DHCP pool spans `.100.240`. They overlap almost entirely.
If the DHCP server hands out `.121/.136/.224` (or any of the above) to a new dynamic client before the static LXC claims it on boot, the static service will fail to bind and the service goes dark.
---
## Proposed approach: Shrink the pool
**Move the DHCP pool start from `.100` to `.241`**, resulting in:
- **New pool:** `192.168.8.241 192.168.8.254` (14 dynamic IPs)
- **Reserved:** `.100.240` stays for static hosts, `.2` for Technitium, `.1` for gateway
- **Zero changes to any LXC, VM, Caddy, or Proxmox config.**
Why `.241.254`:
- Highest static IP is `.239` (elementsynapse) — `.241` gives a 1-IP gap
- `.255` is the broadcast address (unusable)
- 14 IPs is plenty for truly dynamic clients (new transient containers, test VMs)
- If more are ever needed, the pool can easily be widened back down
---
## Tasks
### Task 1: Verify current Technitium DHCP scope from the API
**Objective:** Confirm the active pool range matches what's documented.
**Step 1: Log in to Technitium API and get a token**
```bash
TOKEN=$(curl -sk -X POST http://192.168.8.2:5380/api/user/login \
-H "Content-Type: application/json" \
-d '{"user":"admin","pass":"'$(cat /opt/technitium/admin_password.txt)'","includeInfo":false}' \
| jq -r '.token')
echo "Token: ${TOKEN:0:10}..."
```
**Step 2: Fetch current DHCP scopes**
```bash
curl -sk "http://192.168.8.2:5380/api/dhcp/scopes/list?token=$TOKEN" | jq .
```
**Expected:** One scope named `homelab` with `startingAddress: "192.168.8.100"` and `endingAddress: "192.168.8.240"`.
**Verification:** If the scope is NOT `.100.240`, note the actual range and adjust the plan.
---
### Task 2: Update the DHCP scope to `.241.254`
**Objective:** Shrink the pool so it no longer overlaps static IPs.
**Step 1: Update the scope via API**
```bash
curl -sk -X POST "http://192.168.8.2:5380/api/dhcp/scopes/set?token=$TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "homelab",
"startingAddress": "192.168.8.241",
"endingAddress": "192.168.8.254",
"subnetMask": "255.255.255.0",
"gatewayAddress": "192.168.8.1",
"dnsServerAddresses": ["192.168.8.2"],
"leaseTime": 86400
}'
```
**Step 2: Verify the change took effect**
```bash
curl -sk "http://192.168.8.2:5380/api/dhcp/scopes/list?token=$TOKEN" | jq '.response.scopes[0] | {startingAddress, endingAddress}'
```
**Expected:**
```json
{
"startingAddress": "192.168.8.241",
"endingAddress": "192.168.8.254"
}
```
**Pitfall:** If the API returns `{"status":"error"}`, the scope name or parameter format may differ. Inspect the response body. Technitium's API might use `rangeStart`/`rangeEnd` instead of `startingAddress`/`endingAddress`. Adjust if needed (check the full scope object from Task 1 step 2 for exact key names).
---
### Task 3: Check for active DHCP leases in the old pool that would be stranded
**Objective:** Ensure no DHCP client is currently holding an IP in `.100.240` that it will lose when its lease expires.
**Step 1: List active DHCP leases**
```bash
curl -sk "http://192.168.8.2:5380/api/dhcp/leases/list?token=$TOKEN" | jq '.response.leases[] | {ip: .ipAddress, client: .clientHostname, mac: .hardwareAddress, expires: .leaseExpires}'
```
**Step 2: Interpret results**
- If the only leases are from static LXCs that configured themselves before the DHCP move (e.g., old leases from before the 2026-06-02 static-IP migration), these leases are stale and harmless.
- If a *dynamic* client (e.g., a test laptop, transient VM) holds `.195` or similar, note it — it will lose its IP on next renew and should be moved to a static assignment or into the `.241+` pool.
- **ZimaOS (VM 100) at `.195` is a DHCP lease, not static** — this is the one host that needs attention. Either:
- Set a static IP inside ZimaOS (preferred), or
- Add a DHCP reservation for MAC in Technitium to pin `.195`
**Verification:** No "surprise" dynamic clients that would break on lease expiry.
---
### Task 4: Fix ZimaOS IP stability (if needed)
**Objective:** Ensure ZimaOS at `.195` won't float or break when the pool shrinks.
**If ZimaOS already has a static IP configured inside the VM:** Nothing to do.
**If ZimaOS is DHCP-only (likely — doc says "DHCP lease, not a reservation"):**
Option A (preferred): Set a static IP inside ZimaOS via its web UI at `http://192.168.8.195` → Settings → Network → Static IP → `192.168.8.195/24`, gateway `192.168.8.1`, DNS `192.168.8.2`.
Option B: Add a DHCP reservation in Technitium for ZimaOS's MAC address:
```bash
ZIMAMAC=$(ssh root@hubris "qm config 100 | grep net0 | grep -oE '([0-9A-Fa-f]{2}:){5}[0-9A-Fa-f]{2}'")
curl -sk -X POST "http://192.168.8.2:5380/api/dhcp/reservations/add?token=$TOKEN" \
-H "Content-Type: application/json" \
-d "{\"hardwareAddress\":\"$ZIMAMAC\",\"ipAddress\":\"192.168.8.195\"}"
```
**Pitfall:** The `/api/dhcp/reservations/add` endpoint signature is unverified — confirm the exact endpoint name from Technitium's API docs or the web UI before running it. The web console at `http://192.168.8.2:5380` → DHCP → Reservations can be used as a manual fallback.
---
### Task 5: Update documentation in homelab-context
**Objective:** Fix the now-wrong claims about static IPs being "below .100".
**Files to edit:**
1. **`infrastructure/network.md`** — Line 53
- Old: `Most homelab LXCs use static IPs below \`.100\`. DHCP only covers new/transient containers.`
- New: `Static IPs span \`.101.239\` (all LXCs + VMs + workstations). DHCP pool narrowed to \`.241.254\` to avoid overlap.`
2. **`containers/107-dns.md`** — Lines 37, 42, 55
- Line 37: Update pool range: `192.168.8.241 192.168.8.254`
- Line 42: `Static-IP LXCs (below \`.100\`)` → `Static-IP LXCs (\`.101.239\`) are excluded from the pool.`
- Line 55: Add changelog entry for the pool shrink
3. **`containers/107-dns.md`** — Add changelog entry:
```markdown
### 2026-06-03 — DHCP pool narrowed to `.241.254` to exclude static IPs
Previous pool `.100.240` overlapped with all static LXCs/VMs (\`.101.239\`), creating IP conflict risk. Shrunk pool to `.241.254`. No services re-IP'd. See [plan](../plans/2026-06-03-dhcp-pool-exclude-static-ips.md).
```
4. **`infrastructure/network.md`** — Line 51: Update pool range in the DHCP table row.
5. **`plans/2026-06-01-slate-ax-to-sodola-migration.md`** — Line 60: Optionally update the pool range in the config table (or add a post-migration note). This is the historical migration plan, so a footnote rather than an edit may be better.
**Commit:**
```bash
cd /opt/homelab-context
git add infrastructure/network.md containers/107-dns.md plans/
git commit -m "docs: DHCP pool narrowed to .241-.254 to exclude static IPs"
git push
```
---
### Task 6: Verify no regressions
**Objective:** Smoke-test that DNS and key services still work after the scope change.
```bash
# 1. DNS resolution via Technitium
dig @192.168.8.2 +short git.hubris.network
# Expected: 192.168.8.175
# 2. Caddy reverse-proxy chain
curl -sI https://git.hubris.network | head -1
# Expected: HTTP/2 200
# 3. All app names resolve
for name in git cloud media paperless photos matrix auth plato artifacto; do
result=$(dig @192.168.8.2 +short ${name}.hubris.network)
printf "%-20s → %s\n" "${name}.hubris.network" "$result"
done
# 4. Technitium DHCP scope is correct
curl -sk "http://192.168.8.2:5380/api/dhcp/scopes/list?token=$TOKEN" | jq '.response.scopes[0] | {startingAddress, endingAddress}'
```
---
## Risk assessment
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| API call fails (wrong field names) | Medium | Low | Inspect live scope object first (Task 1); adjust payload |
| ZimaOS loses IP on next boot | Low | Medium | Task 4 makes ZimaOS static or reserved |
| Active DHCP client in `.100.240` gets stranded | Low | Low | Task 3 surfaces this; client just requests a new IP from `.241+` |
| Technitium admin password file missing | Low | Medium | `/opt/technitium/admin_password.txt` was created during setup; verify existence |
## Open questions
1. Is zimaos (VM 100) currently DHCP or static? The doc says DHCP lease, but it's listed as `lan_ip: 192.168.8.195` in inventory. If it's actually DHCP, it's the one host that needs a static assignment before the pool shrinks.
2. Are there any transient DHCP clients (test laptops, phones) on the homelab subnet that hold `.100.240` addresses? Check leases before cutting over.
3. Should we widen the pool slightly (e.g., `.230.254`) for more headroom? Currently 14 IPs. If 3+ transient devices are expected, `.230.254` = 25 IPs — still safe since the highest static is `.239` and `.230.239` could be excluded.
## Execution preference
All changes are on the Technitium API + homelab-context repo. No LXC/VM restarts needed. The pool shrink takes effect immediately for NEW DHCP requests; existing leases in the old range continue until expiry (24h max).

View File

@@ -1,129 +0,0 @@
# Plan: Prevent DHCP IP drift from breaking Caddy backends
**Date:** 2026-06-05
**Slug:** prevent-dhcp-ip-drift
---
## Goal
Eliminate the root cause of services becoming unreachable when DHCP lease renewals change backend IPs that Caddy's `reverse_proxy` directives hardcode.
**Triggering incident:** Paperless (LXC 103) and HAOS (VM 108) had DHCP-assigned IPs change from `.130→.243` and `.101→.241` respectively. Caddyfile still pointed at the old IPs → services unreachable from iPhone on Netbird.
## Current context
### DHCP vs static IP inventory
| Machine | Type | PVE ID | Current IP | Allocation method | Status |
|---------|------|--------|-----------|-------------------|--------|
| **jellyfin** | LXC | 101 | 192.168.8.206 | Static (`ip=.../24`) | ✅ |
| **paperless** | LXC | 103 | 192.168.8.243 | **DHCP** (`ip=dhcp`) | ❌ broken, hotfixed to .243 |
| **gitea** | LXC | 104 | 192.168.8.121 | Static | ✅ |
| **apps** | LXC | 105 | 192.168.8.205 | Static | ✅ |
| **nextcloud** | LXC | 114 | 192.168.8.224 | Static | ✅ |
| **elementsynapse** | LXC | 118 | 192.168.8.239 | Static | ✅ |
| **mule-images** | LXC | 120 | 192.168.8.136 | Static | ✅ |
| **caddy** | LXC | 121 | 192.168.8.175 | Static | ✅ |
| **arriman** | LXC | 122 | 192.168.8.132 | Static | ✅ |
| **sophia** | LXC | 119 | 192.168.8.157 | Static? | ? (not in 2026-06-02 list) |
| **nfs-export** | LXC | 102 | 192.168.8.200 | Static? | ? |
| **plato** | LXC | 126 | 192.168.8.190 | Static? | ? |
| **HAOS** | VM | 108 | 192.168.8.241 | **DHCP** (VM — OS-managed) | ❌ broken, hotfixed to .241 |
| **zimaos** | VM | 100 | 192.168.8.195 | DHCP (known stale lease, see note) | ⚠️ open issue |
| **authentik** | — | — | — | migrated to VPS (external) | N/A |
### Infrastructure facts
- **DHCP server:** Technitium on CT 107 (192.168.8.2), pool `.241.254`
- **Static IP range:** `.101.239`
- **DNS:** Split-horizon on Technitium — `*.hubris.network → 192.168.8.175` (Caddy itself)
- **Caddyfile:** Has 29 `reverse_proxy` directives, all using **hardcoded IP:port**
- **Caddy reload:** Auto-deployed via webhook on git push to `dtoro/caddy-conf`
- **Documentation:** `inventory.yaml` updated, `hosts/*.yaml` stale-before-regenerate
### Why it happened
1. Paperless LXC 103 was **missed** during the 2026-06-02 static-IP migration (Proxmox config still shows `ip=dhcp`)
2. HAOS VM 108 uses DHCP internally (HAOS manages its own network — can't just `pct set`)
3. Both got new IPs from the Technitium `.241.254` pool after a lease renewal
4. Caddyfile still pointed at the old static-range IPs → connection refused
---
## Proposed approach
Three-layer solution:
### Layer 1: Fix the immediate offenders (static assignment)
**Paperless LXC 103:**
- `pct set 103 --net0 name=eth0,bridge=vmbr0,gw=192.168.8.1,hwaddr=BC:24:11:0A:8D:C2,ip=192.168.8.130/24,ip6=auto,type=veth`
- Inside the LXC, update `/etc/network/interfaces` to match
- Reboot/restart networking
**HAOS VM 108:**
- Set a **DHCP reservation** in Technitium for the VM's MAC address, pinning it to `192.168.8.101`
- This avoids needing to reconfigure HAOS internally (which is tedious)
- Alternatively: use `ha network update` from the HAOS console to set a static IP
### Layer 2: Audit and fix all remaining DHCP hosts
Check every LXC/VM hosted on hubris:
```bash
for ct in $(pct list | awk 'NR>1{print $1}'); do
echo "=== CT $ct ==="
pct config $ct | grep "^net0"
done
```
Any with `ip=dhcp` that Caddy reverse-proxies to → convert to static.
**Known candidates to check:**
- CT 102 (nfs-export) — `.200` but not in Caddy. May not need static.
- CT 119 (sophia) — `.157` — is this static or DHCP? Not sure.
- CT 126 (plato) — `.190` — same question.
- VM 100 (zimaos) — `.195` but known to have a stale lease (see 2026-06-03 changelog)
### Layer 3: Add validation
Create a script that runs periodically (cron or homelab cronjob):
**`/opt/homelab-context/scripts/check-caddy-backends.sh`:**
1. Parse `/etc/caddy/Caddyfile` on CT 121 to extract all `reverse_proxy IP:port` targets
2. For each `IP:port`, attempt a TCP connect (timeout 3s)
3. Report any that fail
Could also run as a homelab cron job that notifies dtoro on Matrix if a backend is unreachable.
This catches any future drift proactively (before a user reports it).
### Files likely to change
| File | Change |
|------|--------|
| `inventory.yaml` | May update paperless/HAOS IPs if we choose different static IPs |
| LXC 103 Proxmux config (via `pct set`) | Set static IP |
| Technitium DHCP reservations | Add HAOS reservation |
| `/etc/caddy/Caddyfile` on CT 121 | Already fixed — only changes again if we re-assign paperless IP to `.130` |
| `scripts/check-caddy-backends.sh` | New validation script (new file in homelab-context) |
### Risks / Tradeoffs
- **Pinning paperless to `.130`** — if the LXC was reinstalled since then, `.130` may already be in use. Verify first with `arp-scan` or `nmap`.
- **HAOS static IP via Technitium reservation** vs **inside HAOS**: Technitium reservation is simpler (no HA config changes), but if HAOS's DHCP lease expires and the Technitium server is down, the reservation won't help. A static IP inside HAOS is more robust but requires poking the HA console.
- **Validation script false positives** — a service might be legitimately down for maintenance. The script should be a warning, not an alert.
- **Caddy reload** — each Caddyfile edit triggers an auto-reload via webhook. If the backend is down during reload, Caddy itself stays up (it's just a reverse_proxy target).
### Verification
1. After setting paperless static: `ssh root@192.168.8.175 "curl -s -o /dev/null -w '%{http_code}' http://192.168.8.130:8000"` → 302
2. After Technitium HAOS reservation: `curl -s -o /dev/null -w '%{http_code}' http://192.168.8.101:8123` → 200
3. Run validation script → all targets reachable
4. Confirm from iPhone: both `paperless.hubris.network` and `home.hubris.network` load
### Open questions
1. Should paperless go back to `.130` (its original), or stay at `.243` (current)? Going back to `.130` means updating the Caddyfile again, but keeps the static range allocation consistent.
2. HAOS: Technitium reservation or HAOS-internal static config? Reservation is easier; HAOS-internal is more robust.
3. Should the Caddyfile validation script run as a homelab cron job, or as a cron on the caddy LXC itself?
4. ZimaOS (VM 100) — should we also pin its IP while we're at it?

View File

@@ -1,130 +0,0 @@
# Plan: Fix Frequent Authentik Login Prompts
## Goal
Stop requiring repeated login to Authentik (several times per day) by fixing session and cookie expiry settings so the user stays logged in for longer periods (e.g., 730 days, or until explicit logout).
## Current Context
Authentik runs on the VPS (`82.165.190.79`) in Docker Compose. Traffic flows:
```
Browser → Caddy (LXC 121) → VPS Traefik → Authentik
```
Caddy's `forward_auth` uses the `(authentik)` snippet which proxies to `auth.hubris.network/outpost.goauthentik.io/auth/caddy`. The Authentik server version is **2026.5.2**.
## Root Cause Found
### Primary: `SESSION_EXPIRE_AT_BROWSER_CLOSE = True`
The Authentik Django session (`authentik_session` cookie) is configured to **expire on browser close**. Every time the user closes and reopens their browser, the session cookie is cleared. The next visit to a service that requires OAuth2 authorization (Gitea, Jellyfin, etc.) will redirect to the Authentik login page.
### Secondary: `SESSION_COOKIE_AGE = 86400` (24 hours)
Even with the browser left open continuously, the session expires after 24 hours. Combined with `SESSION_SAVE_EVERY_REQUEST = False`, activity does NOT extend the session.
### Session configuration (from Docker Python environment):
| Setting | Current Value | Default in Django |
|---------|---------------|-------------------|
| `SESSION_EXPIRE_AT_BROWSER_CLOSE` | `True` | `False` |
| `SESSION_COOKIE_AGE` | `86400` (24h) | `1209600` (14d) |
| `SESSION_SAVE_EVERY_REQUEST` | `False` | `False` |
| `SESSION_COOKIE_SAMESITE` | `Lax` | `Lax` |
### What ISN'T the problem:
- **Proxy cookie validity** — `hubris-forward-auth` has `access_token_validity = hours=24`, which is reasonable for the forward-auth token.
- **Server-side session duration** — The `user_login` stage has `session_duration = seconds=0` (indefinite).
- **Refresh tokens** — All OAuth2 providers have `refresh_token_validity = days=30`, which is fine.
- **Caddy configuration** — The forward-auth chain is correctly set up.
- **Outpost health** — All containers healthy, up for 6 days.
## Proposed Approach
Change two Django session settings via Authentik environment variables:
1. **`AUTHENTIK_SESSION_COOKIE_AGE` = 604800** (7 days) — extends session cookie lifetime from 24h to 7 days
2. **`AUTHENTIK_SESSION_EXPIRE_AT_BROWSER_CLOSE` = false** — prevents session cookie from being cleared on browser close
This keeps users logged in for up to 7 days with normal browser use (close/reopen, daily usage). The session still expires after 7 days of inactivity (`SESSION_SAVE_EVERY_REQUEST` stays False).
## Step-by-step Plan
### Step 1: Add environment variables to Docker compose
Edit `/opt/docker-compose.yml` on the VPS to add these env vars to the `authentik-server` service:
```yaml
authentik-server:
environment:
# ... existing vars ...
AUTHENTIK_SESSION_COOKIE_AGE: "604800" # 7 days (was 86400 / 24h)
AUTHENTIK_SESSION_EXPIRE_AT_BROWSER_CLOSE: "false" # was true
```
Note: the Authentik config system uses `__` (double underscore) for nesting. The env vars map to the Django settings via the config YAML path. The correct Authentik env var for `SESSION_COOKIE_AGE` would be `AUTHENTIK_SESSION__COOKIE_AGE` if it goes through the config system, or just `SESSION_COOKIE_AGE` if it's passed directly. Need to verify the exact variable name Authentik expects.
### Step 2: Verify variable naming
Check the Authentik config YAML (`/authentik/lib/default.yml` inside the container) to confirm the exact env var name mapping. Authentik uses a custom config layer that maps env vars to settings.
**Alternative if env vars don't work:** Some Authentik settings need to be set via the admin UI (under System Settings or Tenant settings). The Django session settings might need to be configured differently in this version.
### Step 3: Restart Authentik server
```bash
ssh root@82.165.190.79
docker compose -f /opt/docker-compose.yml restart authentik-server
```
### Step 4: Verify the fix
```bash
# Check session settings took effect
ssh root@82.165.190.79 'docker exec -i authentik-server python3 << "PYEOF"
import os
os.environ.setdefault("DJANGO_SETTINGS_MODULE", "authentik.root.settings")
import django
django.setup()
from django.conf import settings
print("SESSION_EXPIRE_AT_BROWSER_CLOSE:", settings.SESSION_EXPIRE_AT_BROWSER_CLOSE)
print("SESSION_COOKIE_AGE:", settings.SESSION_COOKIE_AGE)
PYEOF'
```
### Step 5: Functional test
1. Login to Authentik at `auth.hubris.network`
2. Close the browser completely
3. Re-open browser, navigate to a forward-auth-gated service (e.g., paperless.hubris.network)
4. Verify you're NOT redirected to login
5. Verify OAuth2 services (Gitea) also maintain the session
## Files Likely to Change
| File | Change |
|------|--------|
| `/opt/docker-compose.yml` | Add `AUTHENTIK_SESSION_COOKIE_AGE` and `AUTHENTIK_SESSION_EXPIRE_AT_BROWSER_CLOSE` env vars |
## Tests / Validation
1. **Config verification** — Run Python snippet inside container to confirm Django settings changed
2. **Browser test** — Close/reopen browser, verify session persists (Step 5 above)
3. **24-hour test** — Check session is still alive after 24h of normal use
## Risks, Tradeoffs, and Open Questions
| Risk | Mitigation |
|------|------------|
| Env var names don't match Authentik's config schema | First verify in the container's `default.yml` config file |
| 7-day persistent cookie is a security concern (stolen cookie = 7 days of access) | This is the same risk as any "Remember Me" feature on any web app. The tradeoff is convenience vs. security. |
| The proxy cookie (`authentik_proxy_*`) may still have its own 24h limit | That's managed separately via the OAuth2 provider's `access_token_validity` setting. If we also want to extend that, we can update `hubris-forward-auth` provider's `access_token_validity` from `hours=24` to `days=7`. |
| `SESSION_COOKIE_SECURE = False` | Should be `True` since Authentik is served behind HTTPS. However, the forward-auth subrequest from Caddy to the outpost is HTTP internally (`http://127.0.0.1:8099`), so `False` may be intentional for the outpost check. |
## Open Questions
1. **What environment variable name does Authentik use for Django session settings?** Need to check `default.yml`. The config layer may use `AUTHENTIK_SESSION__COOKIE_AGE` (double underscore) or the raw Django setting name.
2. **Should we also extend the proxy token validity?** The `hubris-forward-auth` provider has `access_token_validity = hours=24`. If we want users to not need re-login for more than 24h, we should also bump this to match the session cookie age.
3. **Which specific service triggers the most login prompts?** The forward-auth (Caddy-gated) services use proxy cookies. OAuth2 services (Gitea, Jellyfin) use the Django session. Understanding which one the user is hitting most could narrow the fix scope.

View File

@@ -1,149 +0,0 @@
# Plan: Fix Caddyfile truncation + prevent recurring outages
**Date:** 2026-06-06
**Slug:** caddyfile-truncation-permanent-fix
---
## Goal
Restore all `*.hubris.network` services that went offline when the Caddyfile on LXC 121 was truncated to only 3 photo-related site blocks, and implement automated safeguards to prevent this class of outage from recurring.
## Root cause
The Caddyfile at `/etc/caddy/Caddyfile` on LXC 121 was manually edited locally (not via the `dtoro/caddy-conf` git repo), overwriting ~260 lines (30+ site blocks + forward-auth infrastructure) with only 43 lines covering `photos.hubris.network`, `prism.hubris.network`, and a manually-added `photos2.hubris.network`.
**Evidence:**
- `git diff HEAD -- Caddyfile` shows `+3 / -159` lines diff — all other blocks deleted
- Git reflog shows HEAD at `32575ce` (`fix: sab... port 8081→8082`), but working tree diverges
- Deploy webhook log: Jun 06 12:39 — `deploy failed: git pull` (dirty tree blocks merge)
- Backup file `Caddyfile.bak.1780263919` (225 lines) confirms the full original was intact before truncation
- `origin/master` at `1b977aa` is the authoritative source — 260 lines, all blocks present
**Why "third time this week":**
| Incident | Date | Cause |
|---|---|---|
| 1 | Jun 02 | DHCP IP drift — paperless (130→243), HAOS (101→241) |
| 2 | Jun 05 | More DHCP drift — apps (205), mule-images (136 overridden by dhclient) |
| 3 | Jun 06 | **Caddyfile truncated** — unrelated to IPs, much worse |
The Caddyfile truncation is the most severe: it took down **all LAN services** except `photos.hubris.network` and `auth.hubris.network` (VPS-hosted).
## Immediate fix
### Step 1: Restore Caddyfile from origin/master and reload
On LXC 121:
```bash
cd /etc/caddy
# Stash any local changes
git stash
# Reset to origin/master
git checkout --force origin/master -- Caddyfile
# Caddyfile now has all 30+ sites
caddy validate --config /etc/caddy/Caddyfile
systemctl reload caddy
```
This restores all service blocks including: media, git, paperless, books, home, cloud, matrix, proxmox, docker, jellyseerr, qbit, sab, blog, auth, artifacto, plato, zimaos, mcp, secrets, sso + authentik forward-auth infrastructure.
### Step 2: Add `photos2.hubris.network` via git (if still needed)
The `photos2.hubris.network` block was manually added locally and is NOT in origin/master. If the user wants to keep it, submit a PR/commit to the `dtoro/caddy-conf` repo.
### Step 3: Verify
- From any LAN/mesh client: `curl -sk https://media.hubris.network/` → 200
- Run `bash /opt/homelab-context/scripts/check-caddy-backends.sh` from hubris → all targets reachable
- Flush mac-mini DNS: `sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder`
## Permanent safeguards
### Layer 1: Caddyfile integrity check (deploy hook)
Add a site-count validation to the deploy script (`/etc/caddy/scripts/deploy.sh`):
```bash
# Count site blocks (lines matching *.hubris.network {)
SITE_COUNT=$(grep -c '^[a-z].*hubris.network {' Caddyfile)
if [ "$SITE_COUNT" -lt 20 ]; then
echo "[deploy] ERROR: Only $SITE_COUNT sites found (expected 20+). Refusing to reload."
exit 1
fi
```
This catches any future truncation before `caddy reload` runs.
### Layer 2: Caddyfile backup on deploy
Add to deploy script before git pull:
```bash
cp Caddyfile "Caddyfile.bak.$(date +%s)"
```
Keep last 3 backups, auto-rotate.
### Layer 3: Dirty-tree handling in deploy webhook
The deploy webhook currently hard-fails when the working tree is dirty. Change the receiver script to handle this gracefully:
```bash
cd /etc/caddy
# If dirty, stash local changes
if ! git diff --quiet; then
echo "[deploy] Working tree dirty — stashing"
git stash push -m "auto-stash by deploy webhook $(date)"
fi
git pull --ff-only
```
This prevents the webhook from blocking on future local edits.
### Layer 4: Scheduled Caddyfile health check
Add a homelab cron job that runs `check-caddy-backends.sh` every 10 minutes and notifies if any Caddy backend is unreachable.
```yaml
# In homelab context: cronjob
schedule: "*/10 * * * *"
script: /opt/homelab-context/scripts/check-caddy-backends.sh
```
### Layer 5: DNS sync cron (fix already-deployed sync)
The `dns-sync.py` on LXC 107 at `/opt/dns-sync/sync.py` is installed but has **no crontab** — the sync never runs automatically. The NetBird managed DNS zone has drifted from Technitium. Add a systemd timer or crontab:
```bash
echo "*/10 * * * * root python3 /opt/dns-sync/sync.py >> /var/log/dns-sync.log 2>&1" > /etc/cron.d/dns-sync
```
## Files likely to change
| File | Change |
|------|--------|
| `/etc/caddy/Caddyfile` on LXC 121 | Restore from origin/master |
| `/etc/caddy/scripts/deploy.sh` on LXC 121 | Add site-count validation + backup + dirty-tree handling |
| `caddy-conf` git repo | PR with deploy.sh improvements + photos2 (if wanted) |
| `cronjob` in Hermes | Schedule `check-caddy-backends.sh` |
| `/etc/cron.d/dns-sync` on LXC 107 | New — add dns-sync cron |
## Verification
1. All `*.hubris.network` URLs load from mac-mini: `media`, `git`, `paperless`, `cloud`, `home`, `proxmox`, etc.
2. `check-caddy-backends.sh` exits 0 on hubris
3. `systemctl status caddy` shows active on LXC 121
4. `dns-sync` runs and writes to `/var/log/dns-sync.log`
## Risks / Tradeoffs
- **Restoring from origin/master overwrites photos2.hubris.network** — recreate it via proper git commit
- **Caddy staging ACME certs for prism/photos2**: The `tls dns ionos` directive uses staging env (`acme-staging-v02.api.letsencrypt.org`), which fails DNS propagation check (VPS port 53 unreachable from LXC). Once restored, these two subdomains will have the same issue. Move them to production IONOS DNS-01 by removing the staging CA directive or setting the correct `acme_issuer` in Caddyfile.
- **Dirty-tree stash could lose edits** — mitigated by `git stash push --message` + backup file creation before stash
## Open questions
1. Keep `photos2.hubris.network`? If yes, add via proper git push.
2. `prism.hubris.network` and `photos2` certs fail on staging ACME — set production `acme_issuer` in Caddyfile?
3. Should `check-caddy-backends.sh` run as a homelab cron job or as a regular cron on LXC 121?

View File

@@ -1,358 +0,0 @@
# Assessment: Which nodes can move to `strong`
## Executive summary
hubris is **memory-starved**: 28 GiB RAM, 71.9 GiB allocated across 18 LXC + 2 VM
(2.5× overcommit), 10 GiB swap in active use. strong sits **completely empty**
28 GiB RAM, 25 GiB free, 0 guests, 2.7 TiB unused storage. The single most
effective decongestion move is to shift guests off hubris onto strong.
This document assesses every guest for move-readiness, grouped by constraints
(library dependency, GPU, core-infra status), and proposes a phased migration
that does **not** require the physical library-SSD move (the blocker of the
original plan) — library access from strong is provided via NFS from hubris.
---
## Current resource state (live, 2026-07-05)
### hubris — overloaded
| Resource | Capacity | Allocated (all guests) | Actual use | Status |
|----------|----------|----------------------|------------|--------|
| RAM | 28 GiB | 70.2 GiB (2.5× overcommit) | 18 GiB used + 10 GiB swap | ⚠️ heavy swap pressure |
| CPU | 12 vCPU (6c/12t) | 43 vCPU (3.6× overcommit) | ~43% scaling MHz | OK (shares) |
| local-lvm | 856 GiB | 582 GiB allocated (29% thin) | — | OK |
| library (lvmthin) | 3.7 TiB | — | 1.2 TiB used (34%) | OK, 2.3 TiB free |
### strong — empty, ready
| Resource | Capacity | Used | Status |
|----------|----------|------|--------|
| RAM | 28 GiB | 2.3 GiB (host only) | 25 GiB free |
| CPU | 16 vCPU (8c/16t) | idle (0.08 load) | 100% free |
| local-lvm | 856 GiB | 0 | empty |
| ludo-lvm | 1.8 TiB | 0 | empty |
| Guests | — | 0 LXC, 0 VM | nothing running |
### Network topology constraint
```
Fritz!Box (192.168.178.1)
└── SODOLA 2.5G switch
├── hubris eno1 → vmbr1 (192.168.178.10) → vmbr0 (192.168.8.0/24)
│ └── all 20 guests on 192.168.8.x
└── strong vmbr0 (192.168.178.181)
└── no internal bridge yet, guests would be on 192.168.178.x
```
strong reaches `192.168.8.0/24` via the Fritz static route through hubris.
Guests on strong get `192.168.178.x` IPs unless we add an internal bridge
on strong (Phase 0 prerequisite — see below).
---
## Per-guest assessment
### Tier 1 — Move immediately (no library dependency, no core-infra)
These guests mount **no** `/mnt/library` and are not part of the core
infrastructure spine (caddy/dns/auth/mcp). They are the easiest wins.
| ID | Name | Cores | RAM | Library? | GPU? | Notes |
|----|------|-------|-----|----------|------|-------|
| 118 | elementsynapse | 2 | 4 GiB | ❌ | ❌ | Matrix homeserver. Public via Caddy (`matrix.hubris.network`). Only change: Caddy backend IP. **Easiest move in the fleet.** |
| 129 | house | 2 | 3 GiB | ❌ | ❌ | Yuvomi family planner (Docker). No public Caddy route yet (uses VPS traefik directly). Self-contained. |
**Combined RAM freed from hubris: 7 GiB.** No NFS, no library, no GPU.
### Tier 2 — Move with library NFS (high resource consumers)
These are the heaviest guests and the original migration plan's primary
targets. They mount `/mnt/library` and two use the iGPU. Moving them
requires an NFS export from hubris → strong (reverse of the original
plan's direction, since the physical SSD hasn't moved).
| ID | Name | Cores | RAM | Library? | GPU? | I/O profile | Notes |
|----|------|-------|-----|----------|------|-------------|-------|
| 120 | mule-images | 6 | 12 GiB | ✅ mp0 | ✅ iGPU | Write-heavy (photo processing) | **#1 RAM consumer.** Strong has Radeon 680M iGPU (VAAPI works). |
| 122 | arriman | 4 | 8 GiB | ✅ mp0 | ❌ | Write-heavy (downloads) | *arr stack + qbit + sab. Mounts library for download writes. |
| 101 | jellyfin | 4 | 8 GiB | ✅ mp0 | ✅ iGPU | Read-heavy sequential | Media streaming + transcode. Strong 680M handles VAAPI. |
**Combined RAM freed: 28 GiB.** This alone would eliminate hubris's swap
pressure entirely.
### Tier 3 — Could move, low urgency
| ID | Name | Cores | RAM | Library? | Notes |
|----|------|-------|-----|----------|-------|
| 130 | grimmory | 1 | 2 GiB | ✅ mp0 | Book library (Docker). Migrated from apps LXC recently. |
| 131 | teddycloud | 1 | 1 GiB | ✅ mp0 | New (not in inventory.yaml yet). |
| 132 | rclone | 1 | 2 GiB | ✅ mp0 (ro) | Backup container. Read-only library mount. |
| 128 | trmnl | 1 | 768 MiB | ❌ | TRMNL middleware. No library. Could move but tiny. |
| 119 | sophia | 2 | 1 GiB | ✅ mp0 | Workshop. Light use. |
### Stay on hubris (core infrastructure)
| ID | Name | Cores | RAM | Why it stays |
|----|------|-------|-----|--------------|
| 121 | caddy | 1 | 512 MiB | Reverse proxy — terminates all `*.hubris.network`. Must stay on hubris for LAN-side reachability. **Needs backend IP updates** when guests move. |
| 107 | dns | 1 | 1 GiB | Technitium DNS, split-horizon. Core. |
| 106 | auth-outpost | 1 | 512 MiB | Authentik SSO enforcement. Core. |
| 105 | apps | 2 | 4 GiB | homelab MCP + secrets-issuance + artifacto. Core infra. Mounts library. |
| 104 | gitea | 1 | 1 GiB | Git server. NFS would hurt git lock/stat perf. Mounts library (bare repos). |
| 103 | paperless | 2 | 3 GiB | Document archive. Moderate I/O, OCR writes. Mounts library. |
| 102 | nfs-export | 1 | 512 MiB | Exports library to zimaos via NFS. Must stay with the physical library. |
| 100 | zimaos | 4 | 8 GiB (VM) | NAS frontend eval. Already NFS-mounts library from 102. |
| 108 | haos | 2 | 4 GiB (VM) | Home Assistant OS. Hardware access, low latency. |
---
## Constraints & prerequisites
### 1. Network — strong needs an internal bridge (Phase 0)
strong currently has only `vmbr0` on `192.168.178.0/24`. Guests created there
get household-LAN IPs, not homelab-subnet IPs. Two options:
- **Option A (recommended):** Add `vmbr1` on strong as a portless internal
bridge with a `192.168.8.x/24` address (e.g. `192.168.8.3`). Route between
strong's `vmbr0` and `vmbr1` the same way hubris does. Guests go on `vmbr1`
and get `192.168.8.x` IPs — transparent to Caddy, DNS, and inter-LXC refs.
Requires adding a static route on Fritz (or relying on hubris's existing
route — strong would need IP forwarding + a route to 192.168.8.0/24 via vmbr1).
- **Option B (simpler, messier):** Put guests on `192.168.178.x` directly.
Caddy can still reach them (hubris routes to 192.168.178.0/24). But DNS
records, inter-LXC references, and firewall rules all assume `192.168.8.x`.
More config churn per guest.
### 2. Storage — rootfs migration (no shared storage)
`local-lvm` is per-node (not shared). Moving an LXC requires either:
- `vzdump` → restore on strong (clean, but needs temp disk space + downtime)
- `rsync` the rootfs to a new LXC on strong (faster for large rootfs like 120's 100G)
- `pct migrate` only works with shared storage — **not applicable here**
For VMs (100, 108): `qm migrate` also needs shared storage. Not moving VMs.
### 3. Library access — NFS from hubris to strong
Since the physical library SSD is still on hubris, strong's guests that need
`/mnt/library` must NFS-mount it from hubris. Options:
- **Export from hubris host directly** (simplest): add `/mnt/library` to
`/etc/exports` on hubris with the same squash params as LXC 102
(`rw,all_squash,anonuid=33,anongid=10000,no_subtree_check`). Mount on strong
at `/mnt/library`. Strong's guests bind-mount it just like hubris's guests do.
- **Use existing nfs-export LXC 102**: strong NFS-mounts from `192.168.8.200`
(LXC 102). This already has the right squash config. Less host-level change.
**This is the path of least resistance.**
### 4. GPU — iGPU passthrough on strong
strong has a Ryzen 7 PRO 6850U with Radeon 680M iGPU. For jellyfin (VAAPI
transcoding) and mule-images (photo processing), we need:
- `/dev/dri/renderD128` passed to the LXC (`lxc.cgroup2.devices.allow` +
`lxc.mount.entry` or Proxmox's `dev0:` passthrough)
- `video` / `render` group membership inside the container
- Confirm `amdgpu` driver loads on strong's host kernel (it should — same APU family)
### 5. Quorum — 2-node cluster, no QDevice
Moving guests to strong does NOT fix the quorum issue but **reduces blast
radius**: if hubris reboots (its known thermal instability), the guests on
strong keep running independently. Consider adding a QDevice as a separate
follow-up — it's orthogonal to this migration.
---
## Revised migration phases
The original plan's NFS-over-LAN approach has been superseded. Instead,
**media library data moves to ludo-lvm** on strong so migrated guests access
it as a local ext4 mount. Data is split by origin:
```
hubris (stays): library SSD (3.7T, 1.2T used)
└── /mnt/library/{documents,images,cloud,homecloud,notes,repos,sophia}
↑ user-generated content (docs, photos, cloud sync, notes, repos, workshop)
strong (moves): ludo-lvm (1.8T, 0 used at start)
└── /mnt/media_local ← 1.5T thin volume
└── {downloads,movies,music,tv,anime,books}
↑ non-user-generated content (media arr stack, book library)
```
| Category | Stays on hubris | Moves to strong |
|----------|----------------|-----------------|
| Media | — | downloads (25G), movies (51G), music (29G), tv (30G), anime (206G) |
| Books | — | books (2.6G) |
| Docs/Photos | documents (249M), images (4K) | — |
| Cloud sync | cloud (287G), homecloud (367G) | — |
| Personal | notes (6.7M), repos (84M), sophia (151G) | — |
| **Total** | **~805G** | **~344G** |
ludo-lvm (1.8T) fits all media + books with ~1.15T headroom for growth.
hubris library SSD (3.7T, 1.2T used) retains the user-generated content.
Both sides keep their data local — no cross-node NFS needed for daily I/O.
---
### Phase 2a — Prepare ludo-lvm on strong
1. Create a ext4 filesystem on ludo-lvm for media:
```bash
lvcreate -n media -L 1.5T ludo-lvm
mkfs.ext4 /dev/ludo-lvm/media
```
2. Mount at `/mnt/media_local` on strong, add to `/etc/fstab`
3. rsync media directories from hubris → strong:
```bash
rsync -av --progress /mnt/library/{movies,tv,anime,downloads,music,books} strong:/mnt/media_local/
```
### Phase 2b — Migrate arriman (122) to strong
1. Stop arriman on hubris, dump rootfs (24G)
2. Restore on strong with IP `192.168.8.245/28` on vmbr1
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (downloads land locally)
4. Update Caddy: jellyseerr/qbit/sab backends → new IP
5. Update inventory.yaml
### Phase 2c — Migrate jellyfin (101) to strong
1. Stop jellyfin on hubris, dump rootfs (16G)
2. Restore on strong with IP `192.168.8.246/28` on vmbr1
3. Pass `/dev/dri/renderD128` + `/dev/dri/card0` (Radeon 680M + RX 7600)
4. Mount `/mnt/media_local` → `/mnt/library` via mp0 (media reads locally)
5. Update Caddy: `media.hubris.network` → new IP
6. Reinstall `sso-inject.js` in web dir (lost on every apt upgrade)
7. Test VAAPI transcoding, SSO login, media playback
### Phase 2d — Migrate grimmory (130) to strong
1. Stop grimmory on hubris, dump rootfs (16G)
2. Restore on strong with IP `192.168.8.247/28` on vmbr1
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (books read locally)
4. Update Caddy: `books.hubris.network` → new IP
5. Update inventory.yaml
6. Test: book browsing, calibre-web access
### No NFS export needed
With the data split by origin, hubris guests that only need user-generated
content (documents, images, cloud, repos, sophia) still access them from the
original library SSD — no cross-node NFS required. The two sides are
independent.
**Result after Phase 2: hubris frees 26 GiB RAM (4 migrated guests) + 344G of
library I/O burden. Strong becomes the media/books powerhouse.**
---
### Phase 3 — Migrate mule-images (120) to strong
Move photo management (12 GiB RAM, 6 cores, iGPU) last because it needs:
- `/mnt/library` access (now NFS from strong — already set up in Phase 2d)
- `/dev/dri/renderD128` (Radeon 680M — confirm VAAPI compatibility first)
Steps:
1. Stop mule-images on hubris, rsync the 100G rootfs to strong (faster than vzdump)
2. Restore on strong with IP on vmbr1
3. Pass Radeon 680M iGPU
4. Reconfigure library paths → `/mnt/media_local` (or keep NFS mount)
5. Update Caddy: `photos.hubris.network` → new IP
6. Test photo import + processing pipeline
---
### Phase 4 — Tier 3 moves (optional)
Migrate grimmory (130), teddycloud (131), rclone (132), trmnl (128), sophia (119)
as needed — each frees 12 GiB. Not urgent; do when convenient.
---
### Phase 5 — Follow-up
- **QDevice**: add a tiebreaker for 2-node quorum
- **Gaming VM**: strong's 6850U has enough cores alongside migrated LXCs
- **Hubris library cleanup**: after all guests are confirmed working, decide
whether to keep the original library SSD as backup or repurpose it
---
## Resource math after Phase 3 (all Tier 1 + 2 moved)
| | hubris | strong |
|---|--------|--------|
| Guests | 11 LXC + 2 VM | 5 LXC |
| RAM allocated | ~25 GiB | ~45 GiB |
| RAM capacity | 28 GiB | 28 GiB |
| Overcommit | 0.9× (under-committed) | 1.6× (manageable) |
| Library disk | Local ext4 (3.7T) → NFS client | Local ext4 on ludo-lvm (1.8T) |
| GPU | Radeon 760M (idle) | Radeon 680M (jellyfin + mule-images) |
strong becomes the media/library powerhouse. hubris becomes a lean core-infra
node (DNS, auth, git, docs, caddy, HA).
---
---
## Risk register
| Risk | Impact | Mitigation |
|------|--------|------------|
| NFS latency for library reads (jellyfin, arriman) | Media playback stutter, slow downloads | Test iperf between strong↔hubris first. If 2.5G link, NFS throughput is fine (~1 Gbit/s). |
| GPU passthrough on strong (680M vs 760M) | Transcode quality/compat differences | Both are AMD VAAPI — same driver stack. Test `vainfo` inside LXC before going live. |
| Caddy backend IP churn | Service outage if IP wrong | Update Caddyfile in git repo (caddy-conf), test each route before destroying old LXC. |
| vzdump/restore downtime | Service unavailable during migration | Schedule off-hours. Use rsync for large rootfs (120's 100G) to minimize freeze window. |
| 2-node quorum still fragile | If hubris goes down, strong /etc/pve goes read-only | Guests keep running. Add QDevice as follow-up. |
| Library data integrity during NFS transition | Permission drift | NFS `all_squash,anonuid=33,anongid=10000` matches existing LXC 102 config. Verify with `ls -la /mnt/library` after mount. |
---
## Open questions for operator
1. **Internal bridge on strong**: proceed with `vmbr1` on `192.168.8.3/24`
(Option A), or use `192.168.178.x` guest IPs (Option B)?
2. **Migration method**: `vzdump`/restore (clean, downtime) vs `rsync` rootfs
(faster for large disks, needs manual config copy)?
3. **Phase 1 priority**: move elementsynapse + house first (quick wins), or
go straight to Phase 2 (mule-images/jellyfin/arriman) for maximum relief?
4. **Should we add a QDevice now** before moving anything, to protect
management plane during the migration?
---
## Changelog
### 2026-07-05 — Phase 2d complete (grimmory migrated; media NFS to zimaos)
grimmory (130) → 192.168.8.247 on strong. Rsync'd /books (2.6G) to ludo-lvm.
LXC 102 (nfs-export) now mounts strong's NFS at /mnt/media and exports it as
a second share alongside /mnt/library. Zimaos mounts both: /media/library
(hubris user-generated) and /media/media (strong media+books).
See hosts/strong.md changelog.
### 2026-07-05 — Phase 2 complete (arriman + jellyfin migrated; library on ludo-lvm)
arriman (122) → 192.168.8.245, jellyfin (101) → 192.168.8.246. Created 1.5T
thin volume on ludo-lvm, rsync'd 363G of media data. Both containers use local
ext4 mount — no NFS. Jellyfin has 680M + RX 7600 GPU passthrough.
Caddy backends updated. See hosts/strong.md changelog.
### 2026-07-05 — Phase 1 complete (elementsynapse + house migrated to strong)
Both Tier 1 guests moved: elementsynapse (118) → 192.168.8.242, house (129) → 192.168.8.244.
Strong now has vmbr1 at 192.168.8.241/28. Hubris has proxy ARP + /32 routes for strong
guest range. DHCP scope narrowed to 192.168.8.100-239 to avoid conflicts.
Teddycloud (LXC 131) given static IP 192.168.8.150 due to IP conflict with
previous DHCP allocation at 192.168.8.243.
See hosts/strong.md changelog for full steps.
### 2026-07-05 — assessment created
Built from live `pct config` + `pvesm status` + `free -h` data pulled from
both nodes. Supersedes the storage-migration framing of the original
library-SSD plan — this assessment treats the SSD move as optional and
focuses on guest relocation via NFS.

View File

@@ -23,11 +23,7 @@ creation_rules:
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4,
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6
- path_regex: ^secrets/gitea-pat\.yaml$
# Write-scoped Gitea PAT (dtoro user). Same recipient list as hello.yaml
@@ -36,21 +32,17 @@ creation_rules:
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4,
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6
- path_regex: ^secrets/gitea-tokens\.yaml$
# Workstations only.
age: >-
# placeholder — fill with age_pubkey of: republic-laptop, mac-mini, strong, hubris
# placeholder — fill with age_pubkey of: republic-laptop, mac-mini, ludo-mini, hubris
- path_regex: ^secrets/webhook-hmacs\.yaml$
# LXCs that run a webhook receiver.
age: >-
# placeholder — fill with age_pubkey of: apps, caddy
# placeholder — fill with age_pubkey of: apps, caddy, claudio-bot, claudio-monitor host
- path_regex: ^secrets/turn-shared-secret\.yaml$
# coturn TURN long-term-credential password. Consumed by hubris (which
@@ -60,10 +52,7 @@ creation_rules:
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6
- path_regex: ^secrets/netbird-authentik-oidc\.yaml$
# Authentik OIDC client secret for the netbird-dashboard provider.
@@ -72,20 +61,7 @@ creation_rules:
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
- path_regex: ^secrets/netbird-pat\.yaml$
# NetBird API Personal Access Token. Consumed by the dns-sync job on the
# `dns` LXC (107) to reconcile Technitium -> NetBird managed DNS zone.
# (When 107 is enrolled, add its age_pubkey here and updatekeys.)
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6
- path_regex: ^secrets/openrouter-api-key\.yaml$
# OpenRouter API key consumed by the `hermes` wrapper (bin/hermes) when
@@ -95,42 +71,5 @@ creation_rules:
# See operations/hermes-agent.md.
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
- path_regex: ^secrets/yuvomi-api-token\.yaml$
# Named Bearer token for the Yuvomi REST API, consumed by yuvomi-mcp on
# LXC 129 (house).
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
- path_regex: ^secrets/hermes-house-users\.yaml$
# Signal number → Yuvomi user_id mapping (PII). Consumed by hermesd on LXC 129.
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
- path_regex: ^secrets/oikos-approval-hmac\.yaml$
# HMAC signing key for Oikos approval-grant tokens (oikos/approve.py).
# Recipients: apps (105, runs the approval engine alongside homelab-mcp)
# and hubris (admin/debug decrypt). See OIKOS.md "Approval engine".
age: >-
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0
- path_regex: ^secrets/oikos-console-deploy-secret\.yaml$
# Shared HMAC secret for the Gitea deploy webhook (id 14) ->
# oikos-console-deploy.service on apps (105). Generated + registered
# with Gitea before the apps-side install ran (see
# oikos/console/deploy/README.md "Status") — write this exact value
# into /etc/oikos-console-deploy/secret rather than letting
# webhook/install.sh generate a fresh one.
age: >-
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6
# webhook noop 2026-05-20T18:16:57+02:00

File diff suppressed because one or more lines are too long

View File

@@ -4,20 +4,6 @@ You are running on a machine that is part of the **hubris** homelab. The full
context is in this checkout at `/opt/homelab-context/`. This file is the entry
point. Read it once at start, then keep working.
The operating model — OODA loop, risk classes, approval rules, the ontology,
and node lifecycle — is defined in [OIKOS.md](.agents/OIKOS.md). Before any mutation,
classify the action against `oikos/policy.yaml`; when the class requires
approval, stop and ask the operator.
Agent-facing instruction is separated from human content under `.agents/`:
`.agents/shared/` holds the conventions every agent applies
([writing-style](.agents/shared/writing-style.md), [caveman](.agents/shared/caveman.md),
[page-templates](.agents/shared/page-templates.md), [llm-wiki](.agents/shared/llm-wiki.md)), and
`.agents/domains/` holds the per-domain schemas
([knowledge](.agents/domains/knowledge/schema.md), [operations](.agents/domains/operations/schema.md)).
The narrative wiki lives under `knowledge/wiki/`; the machine-readable substrate
(`inventory.yaml`, `hosts/*.yaml`, `oikos/`) stays at the repo root.
## 1. Who you are
Run `hostname` (Linux) or `scutil --get LocalHostName` (macOS), then read:
@@ -33,14 +19,13 @@ the operator to run `homelab client add <hostname>` from an existing client.
- `/opt/homelab-context/inventory.yaml` — every host, LXC, VM, and workstation
with their mesh addresses, roles, and service mappings. Treat this file as
authoritative; anything you read in narrative pages should agree with it.
- `/opt/homelab-context/knowledge/wiki/infrastructure/mesh.md` — Tailscale → Netbird state.
- `/opt/homelab-context/infrastructure/mesh.md` — Tailscale → Netbird state.
Both meshes are accepted today; Netbird is preferred for new traffic.
- `/opt/homelab-context/knowledge/wiki/infrastructure/dns.md` — split-horizon DNS via
Technitium on [dns (107)](knowledge/wiki/containers/107-dns.md). `*.hubris.network`
resolves to 192.168.x.x on the LAN and to mesh addresses off-LAN.
- `/opt/homelab-context/.agents/operations/commands.md` — the operator's cheatsheet
for pct, caddy, DNS, and the Oikos command surface. Use these verbs when
you take actions.
- `/opt/homelab-context/infrastructure/dns.md` — split-horizon DNS via
dnsmasq on LXC 124. `*.hubris.network` resolves to 192.168.x.x on the LAN
and to mesh addresses off-LAN.
- `/opt/homelab-context/operations/commands.md` — the operator's cheatsheet
for pct, caddy, dnsmasq. Use these verbs when you take actions.
## 3. The MCP server
@@ -59,15 +44,8 @@ Available tools:
get_service_status(service), tail_log(service, lines=200),
list_lxcs(), get_lxc_state(lxc), ping_service(service)
Oikos (read-only; see OIKOS.md):
explain(service) — compact context card, cheaper than search_docs+get_page
preflight(service) — risk class, approval requirement, verification command
get_relations(entity) — ontology blast-radius query (host: or service: id)
get_change_history(entity, limit=20) — change-ledger entries
get_state_snapshot() — last scheduler Observe-pass (health, disk, drift count)
Mutations are **not** exposed via MCP. Use the `homelab` CLI for those, with
operator confirmation — see OIKOS.md's risk classes and approval flow.
operator confirmation.
**When to prefer MCP over grepping the clone:** any time you need to resolve a
name to an address, look up service status, or search the wiki by content.
@@ -75,25 +53,17 @@ Grep is fine for browsing or when MCP is unreachable.
## 4. Wiki conventions
See [page-templates.md](.agents/shared/page-templates.md) for file naming, page
structure, and the tone standard. Quick reference:
- **File naming:** Foundational docs are ALL-CAPS (AGENTS.md, OIKOS.md, GLOSSARY.md);
containers use `<id>-<name>.md`; infrastructure uses lowercase-with-dashes; plans
and investigations use `YYYY-MM-DD-slug.md`; skills are `<name>/SKILL.md`.
- **Where pages live:** Narrative under `knowledge/wiki/{containers,hosts,vms,infrastructure}/`;
incident records under `knowledge/sources/investigations/`; runbook procedures under
`.agents/skills/<name>/SKILL.md`; operator reference under `.agents/operations/`;
design docs under `plans/`. Cross-link liberally; orphans are bugs.
- **Changelog format:** Every page ends with a `## Changelog` section, entries in
reverse-chronological order:
- Pages live under `containers/`, `hosts/`, `vms/`, `infrastructure/`,
`investigations/`, `operations/`. Cross-link liberally; orphans are bugs.
- Every page ends with a `## Changelog` section, entries in reverse-chrono
order:
### YYYY-MM-DD — short title
one or two lines describing what changed and why.
- **Live state precedence.** If you observe a discrepancy between the docs and
running state, update the docs *in the same session* (per the same-session update
rule in [page-templates.md](.agents/shared/page-templates.md#same-session-update-rule)).
- Investigation files are dated and slugged: `YYYY-MM-DD-slug.md`.
- Live state takes precedence over docs. If you observe a discrepancy, update
the docs *in the same session* (per the same-session update rule).
## 5. Acting on the homelab
@@ -106,44 +76,14 @@ structure, and the tone standard. Quick reference:
demand using the per-client age key at `/etc/age/key.txt`. Secrets ARE
available in this system — `list_my_secrets()` (MCP) shows what you can
decrypt.
- **Mutations** (restart, edit configs, etc.): classify against
`oikos/policy.yaml` first (`homelab decide <action> <entity>`).
`reversible_low` actions just need the interactive confirmation prompt;
`config_mutation`/`destructive` actions are mechanically refused without
a valid `--approval-id` from `homelab approval request` — see OIKOS.md.
For ad-hoc work, SSH and edit directly — but commit changes that touch
tracked configs (caddy, gitea custom, artifacto, mule-image, etc.; see
`knowledge/wiki/infrastructure/auto-deploy.md`).
- **Mutations** (restart, edit configs, etc.): the `homelab` CLI's mutating
subcommands ask for confirmation. For ad-hoc work, SSH and edit directly —
but commit changes that touch tracked configs (caddy, gitea custom,
artifacto, mule-image, etc.; see `infrastructure/auto-deploy.md`).
- **Wiki updates**: same-session rule applies to any meaningful state change
this client makes.
## 6. Communication mode
Read and apply `/opt/homelab-context/.agents/shared/caveman.md` (if present). It defines the lab's
terse-communication standard — drop filler, keep substance, use fragments.
## 7. Auto-setup mechanism
The homelab-context repo ships tooling that gets automatically installed
on every client after `git pull`. This is handled by `tools/post-pull.sh`
(replaces the raw git pull in the sync timer) which runs any script matching
`tools/*.setup.sh` after pull.
Currently auto-setup:
- **Caveman + templates** (`tools/setup-caveman.sh`): Installs Caveman npm
package, wrapper scripts, and compact output templates for token-efficient
CLI output. Wrapper at `~/bin/caveman_wrapper.sh`.
- **Hermes agent persona** (`tools/setup-hermes-soul.sh`): Provisions
`~/.hermes/SOUL.md` from `HERMES.md` on Hermes agents. This ensures every
Hermes agent follows the canonical homelab persona (token efficiency, source
of truth hierarchy). No-op on non-Hermes agents.
To add a new auto-setup, create `tools/<name>.setup.sh` in the repo,
commit and push. All enrolled clients pick it up within 5 minutes.
To trigger sync manually: `sudo homelab sync` or wait for the 5-min timer.
## 8. When in doubt
## 6. When in doubt
Run `homelab mcp search_docs <query>` or `homelab mcp get_host <name>`.
The clone is the fallback; MCP is the index.

98
CONTRIBUTING.md Normal file
View File

@@ -0,0 +1,98 @@
# Contributing to the Homelab Wiki
## Voice
Concise, technical, sysadmin-to-sysadmin. No marketing prose, no exclamation marks.
## Page templates
### Container page (`containers/<id>-<name>.md`)
```markdown
# <id> — `<name>`
One-sentence purpose.
## At a glance
- **Hostname:** `<name>`
- **IP:** `192.168.8.x`
- **Privilege:** privileged | unprivileged
- **Resources:** N cores / M GiB RAM / D GiB rootfs
- **Mounts:** `/mnt/library``/mnt/library` (if any)
- **Public hostname:** `<sub>.hubris.network` (if proxied)
## Role
What it does, what it talks to.
## Service / port map
| Service | Listen | Notes |
## Storage / config paths
## Auto-deploy
(if any) — link to [auto-deploy](../infrastructure/auto-deploy.md)
## Related
- [Caddy](121-caddy.md) (if proxied)
- [DNS](../infrastructure/dns.md) (if has subdomain)
- [Authentik](124-authentik.md) (if SSO)
- ...
## Changelog
### YYYY-MM-DD — short title
What changed, why, link to investigation if any.
```
### Cross-cutting page (`infrastructure/<topic>.md`)
```markdown
# <Topic>
One-sentence summary.
## Why
Design rationale — what it replaces, what it solves.
## Components
Where it runs, what files matter.
## How to apply / use
Recipes.
## Gotchas
## Related
Links to nodes that host or depend on this.
## Changelog
```
### Investigation (`investigations/YYYY-MM-DD-slug.md`)
```markdown
# YYYY-MM-DD — <title>
## Summary
1-3 sentences.
## Timeline
## Root cause
## Mitigations applied
## Open questions
```
## Linking discipline
- Every container page links to every cross-cutting page it participates in.
- Every cross-cutting page lists the nodes that participate.
- Every investigation links to the nodes it implicates *and* gets back-linked from each node's changelog.
## Changelog hygiene
- Reverse-chronological (newest first).
- One entry per discrete change, even if you make several in one day.
- If a change spans nodes, repeat the entry on each affected page (different perspective is fine).
- Don't rewrite history — entries are append-only. Mistakes get a follow-up entry that supersedes them.

View File

@@ -1,50 +0,0 @@
.PHONY: build test test-db lint generate generate-check dev migrate seed export clean tidy
BINARY := oikos
GO ?= go
build:
$(GO) build -o $(BINARY) -tags timetzdata ./cmd/oikos
test:
$(GO) test -race -cover ./...
# Integration tests against the compose Postgres (starts it if needed)
test-db:
docker compose up -d postgres
@sleep 3
OIKOS_TEST_DATABASE_URL="postgres://oikos:$${OIKOS_DB_PASSWORD:-oikos_dev}@localhost:5432/oikos?sslmode=disable" \
$(GO) test -race -count=1 ./internal/db/ ./internal/httpapi/ ./internal/mcp/
lint:
$(GO) vet ./...
@command -v golangci-lint >/dev/null 2>&1 && golangci-lint run || echo "golangci-lint not installed, skipping"
generate:
$(GO) run github.com/oapi-codegen/oapi-codegen/v2/cmd/oapi-codegen@v2.4.1 \
-config api/codegen.yaml api/openapi.yaml
$(GO) run github.com/sqlc-dev/sqlc/cmd/sqlc@v1.29.0 generate
# CI drift guard: regenerate and fail if the committed output changed.
generate-check: generate
@git diff --exit-code -- internal/httpapi/gen internal/db/sqlcgen \
|| (echo "generated code is stale — run 'make generate' and commit" && exit 1)
migrate:
$(GO) run ./cmd/oikos migrate
seed:
$(GO) run ./cmd/oikos seed
export:
$(GO) run ./cmd/oikos export
dev:
docker compose --profile dev up -d
clean:
rm -f $(BINARY)
$(GO) clean -testcache
tidy:
$(GO) mod tidy

233
README.md
View File

@@ -1,190 +1,77 @@
# Homelab OS
# Homelab Wiki — `hubris`
Living documentation for the **hubris** Proxmox homelab + Oikos operating system.
Living documentation for the **hubris** Proxmox homelab. Every node, every cross-cutting system, and every meaningful incident is its own page; pages are linked so you can start anywhere and walk the graph.
**For agents running on enrolled clients:** start with [AGENTS.md](AGENTS.md), then [OIKOS.md](.agents/OIKOS.md).
> Last refreshed against live state: **2026-04-28**.
---
## Map
## For Agents — Navigation & Entry Points
### Hosts
- [`hubris`](hosts/hubris.md) — single Proxmox VE node, GMKtec NucBox M6 Ultra, `192.168.8.77`
### You are running on a client enrolled in the hubris homelab
### VMs
- [100 — `zimaos`](vms/100-zimaos.md) — ZimaOS 1.6.1, NAS frontend (evaluation)
- [108 — `haos-16.3`](vms/108-haos.md) — Home Assistant OS
1. **First:** Read [AGENTS.md](AGENTS.md) once. It explains who you are, the topology, available tools, conventions, and how to act.
2. **Before any mutation:** Read [OIKOS.md](.agents/OIKOS.md). It defines the operating model, risk classes, approval flow, and the ontology you'll consult.
3. **For specific workflows:** Load the matching skill from `.agents/skills/<name>/SKILL.md` (e.g., [service-health-check](.agents/skills/service-health-check/SKILL.md)).
4. **When in doubt:** Use MCP tools (`search_docs`, `get_page`, `explain`, `get_changelog`) — they're cheaper and more reliable than grepping.
### LXC containers
See the full table in [`containers/index.md`](containers/index.md). Quick links:
### Key References for Agents
| ID | Name | IP | Role |
| --- | ---------------- | --------------- | --------------------------------------------- |
| 101 | [jellyfin](containers/101-jellyfin.md) | 192.168.8.206 | Media server |
| 102 | [nfs-export](containers/102-nfs-export.md) | 192.168.8.200 | NFSv4 re-export of /mnt/library for ZimaOS |
| 103 | [paperless](containers/103-paperless.md) | 192.168.8.130 | Document mgmt |
| 104 | [gitea](containers/104-gitea.md) | 192.168.8.121 | Git server |
| 105 | [apps](containers/105-apps.md) | 192.168.8.205 | Docker host (Artifacto / Booklore / PlantUML / Portainer / WriteFreely) |
| 114 | [nextcloud](containers/114-nextcloud.md) | 192.168.8.224 | Personal cloud |
| 118 | [elementsynapse](containers/118-elementsynapse.md) | 192.168.8.239 | Matrix Synapse |
| 119 | [sophia](containers/119-sophia.md) | 192.168.8.157 | Sophia |
| 120 | [mule-images](containers/120-mule-images.md) | 192.168.8.136 | Mule-image / mulita photos |
| 121 | [caddy](containers/121-caddy.md) | 192.168.8.175 | Reverse proxy |
| 122 | [arriman](containers/122-arriman.md) | 192.168.8.132 | Docker host (\*arr stack) |
| 123 | [claudio-bot](containers/123-claudio-bot.md) | 192.168.8.230 | Matrix control plane |
| 124 | [authentik](containers/124-authentik.md) | 192.168.8.180 | SSO + split-horizon DNS |
| 126 | [plato](containers/126-plato.md) | 192.168.8.190 | Plato (notes/discovery workspace) |
- **What am I?** → `/opt/homelab-context/hosts/<hostname>.yaml` (read on first run)
- **Live topology** → `inventory.yaml` + `hosts/*.yaml` (canonical, always wins)
- **Risk & approval** → [oikos/policy.yaml](oikos/policy.yaml) (enforced, not advisory)
- **Runbooks & workflows** → [.agents/skills/](.agents/skills/) (risk class + verification checklist included)
- **State of Oikos** → [OIKOS.md build status](.agents/OIKOS.md#build-status-30-day-roadmap) (scheduled probes, drift detectors, signals, approval engine)
### Cross-cutting infrastructure
- [DNS — split-horizon](infrastructure/dns.md)
- [Ingress — Caddy + VPS traefik](infrastructure/ingress.md)
- [Mesh — Tailscale → Netbird migration](infrastructure/mesh.md)
- [Monitoring — claudio-monitor](infrastructure/monitoring.md)
- [Media permissions — `media` GID 10000](infrastructure/media-permissions.md)
- [SSH access](infrastructure/ssh-access.md)
- [Backups — restic on external drive (disabled)](infrastructure/backups.md)
- [Auto-deploy — gitea-webhook pipelines](infrastructure/auto-deploy.md)
- [VPS hardening — IONOS / netbird control plane](infrastructure/vps-hardening.md)
- [Homelab context distribution](infrastructure/homelab-context.md) — cross-client `/opt/homelab-context` + MCP + secrets-issuance
### When to Use MCP vs Files vs Shell
### Investigations
Time-stamped incident notes / experiments in [`investigations/`](investigations/index.md).
| Task | Use | Tool |
|------|-----|------|
| Resolve hostname → address | MCP | `get_host(name)` or `list_services()` |
| Search wiki by content | MCP | `search_docs(query)` |
| Read a wiki page | MCP or file | `get_page(path)` or `cat knowledge/wiki/.../...md` |
| Get changelog entries | MCP | `get_changelog(page, since?)` |
| Understand a service | MCP | `explain(service)` — compact context card, cheaper than search+read |
| Blast-radius query | MCP | `get_relations(entity)` (ontology walk) |
| List available secrets | MCP | `list_my_secrets()` (scoped to your age key) |
| Browse or grep | File | Raw `grep` when MCP unreachable, or exploratory browsing |
**When MCP is unreachable:** fall back to grepping the clone at `/opt/homelab-context/`. The local files are the same; MCP is just an index.
---
## Understanding the Operating Model
Before you act, **classify your action against [oikos/policy.yaml](oikos/policy.yaml)**.
### The Oikos OODA Loop + Decision Tree
```mermaid
flowchart TD
Observe["**Observe**<br/>probes, drift detectors, agent signals"]
Orient["**Orient**<br/>ontology, context, state, entity relations"]
Decide{"**Decide**<br/>classify against oikos/policy.yaml"}
Auto["Auto-act<br/>(unattended)"]
Escalate["Escalate<br/>homelab approval request"]
Act["**Act**<br/>homelab CLI, runbooks, skills"]
Verify["**Verify**<br/>checklist from SKILL.md"]
Ledger["**Ledger**<br/>mutation record: who/what/risk"]
Document["**Document**<br/>wiki update, same-session rule"]
Observe --> Orient --> Decide
Decide -->|read_only, reversible_low| Auto
Decide -->|config_mutation, destructive| Escalate
Auto --> Act
Escalate -->|approval granted| Act
Act --> Verify --> Ledger --> Document
Document -.loop.-> Observe
```
### Risk Classes (enforced, not advisory)
From [oikos/policy.yaml](oikos/policy.yaml):
- **read_only** — status, logs, docs, inventory queries. Unattended. MCP tools are all read_only.
- **reversible_low** — restart, cache clear, sync pull. Unattended + ledger entry.
- **config_mutation** — tracked-config edits (commit+push, never local), deploys, upgrades, DNS/ingress changes. **Operator approval required.**
- **destructive** — destroy, format, wipe, rotate, revoke. **Approval + typed confirmation phrase.**
### Decision Flow
1. **Decide:** Use `homelab decide <action> <entity>` to classify (risk class × blast radius × confidence).
2. **Escalate if needed:** `homelab approval request` (Matrix-delivered to operator; see [operations/commands.md](.agents/operations/commands.md)).
3. **Execute:** Use `homelab` CLI (not ad-hoc SSH) — it enforces policy, logs mutations, and verifies outcomes.
4. **Document:** Update wiki in the same session (per [AGENTS.md §5](AGENTS.md#5-acting-on-the-homelab) and the [same-session rule](.agents/shared/page-templates.md#same-session-update-rule)).
### The Ontology Graph
Everything that can break, be changed, or hold data has an entity in `inventory.yaml` + `oikos/ontology.yaml`. Blast-radius questions ("what breaks if strong goes down?") are graph walks via `homelab node <name> relations`, not doc archaeology.
**See:** [OIKOS.md](.agents/OIKOS.md) (full operating model, OODA loop, primitives, lifecycle gates, build status).
---
## Finding & Understanding Information
The narrative documentation is organized in **layers**:
| Layer | What it is | Where | Immutable? | How agents use it |
|-------|-----------|-------|-----------|-------------------|
| **Sources** | Raw evidence: incidents, external refs, live state | `knowledge/sources/investigations/` | Yes | Read to understand root causes; do not rewrite |
| **Wiki** | Synthesized current-state: one page per node & per system | `knowledge/wiki/{containers,hosts,vms,infrastructure}/` | No | This is the reference layer — if wiki disagrees with live state, update it *in the same session* |
| **Index** | Pure listings — every page in scope with one-line summary | `index.md` / folder `README.md` | No | Navigation aid; keep it current when wiki restructures |
| **Log** | Append-only doc-maintenance record (restructures, ingests, lints) | `knowledge/log.md` | Yes (append-only) | Read to understand past doc changes; never edit directly |
**Changelog ≠ Log:** Each wiki page ends with a `## Changelog` (infrastructure changes to that node, machine-parsed). That's not the Log; the Log records *doc operations* only.
**See:** [llm-wiki.md](.agents/shared/llm-wiki.md) (full rules, page structure, immutability contract).
---
## Map & Quick Navigation
### Agent Entry Points (Start Here)
- **You are an agent** → [AGENTS.md](AGENTS.md) (on deployed clients: `/opt/homelab-context/AGENTS.md`)
- **Operating model & risk policy** → [OIKOS.md](.agents/OIKOS.md)
- **Specific workflows** → [.agents/skills/](.agents/skills/) (load the matching SKILL.md before acting)
- **Operations cheatsheet** → [.agents/operations/commands.md](.agents/operations/commands.md)
- **Tools & MCP reference** → [AGENTS.md §3 — The MCP server](AGENTS.md#3-the-mcp-server)
### Topology & Infrastructure
Node counts, IPs, and service lists change often — treat `inventory.yaml` and the index pages below as the source of truth, not this README.
- **Proxmox hosts** → [knowledge/wiki/hosts/index.md](knowledge/wiki/hosts/index.md)
- **VMs** → [knowledge/wiki/vms/index.md](knowledge/wiki/vms/index.md)
- **LXC containers** → [knowledge/wiki/containers/index.md](knowledge/wiki/containers/index.md)
- **Cross-cutting infrastructure** (DNS, ingress, mesh, backups, monitoring, auto-deploy, VPS) → [knowledge/wiki/infrastructure/index.md](knowledge/wiki/infrastructure/index.md)
### Knowledge & References
- **Glossary** — [GLOSSARY.md](knowledge/GLOSSARY.md)
- **Incidents & investigations** — [knowledge/sources/investigations/index.md](knowledge/sources/investigations/index.md) (active + [archive](knowledge/sources/investigations/archive/))
- **Plans & design docs** — [plans/index.md](plans/index.md)
- **Hermes agent** (for Hermes-enrolled clients) — [HERMES.md](.agents/HERMES.md)
---
### Operations
- [Command cheatsheet](operations/commands.md)
- [Agent enrollment](operations/agent-enrollment.md) — bootstrap a new client (workstation, LXC, VM) into the homelab context system
## Conventions
All pages follow:
- **Each node page** ends with a `## Changelog` section. Reverse-chronological. Entry format:
```
### YYYY-MM-DD — short title
one or two lines on what changed and why.
```
- **Cross-linking is mandatory.** If a page references another node or system, link to it. Treat orphans as a bug.
- **Live state wins.** When something here disagrees with `pct config` / `docker inspect` / running config, fix the wiki *and* note the change in the relevant changelog.
- **Tracked configs.** A node whose config lives in a Gitea repo (Caddy, Gitea customizations, Artifacto, mule-image, claudio-bot) is auto-deployed via webhook — see [auto-deploy](infrastructure/auto-deploy.md). Edits there must be pushed, not left local.
- **No secrets.** This is a private repo on `git.hubris.network`, but still: paths to secret files are fine, secret values are not.
- **File naming.** Foundational docs (entry-points, agent instruction, references) are ALL-CAPS (`AGENTS.md`, `OIKOS.md`, `GLOSSARY.md`); containers use `<id>-<name>.md`; infrastructure pages use lowercase-with-dashes; plans and incidents use `YYYY-MM-DD-slug.md`; skills are `<name>/SKILL.md`. See [page-templates.md](.agents/shared/page-templates.md#file-naming) for the full rules.
- **Voice & vocabulary.** Concise, technical, sysadmin-to-sysadmin. No marketing prose, no puffers (seamless, robust, leverage, etc.). Full rules in [writing-style.md](.agents/shared/writing-style.md).
- **Cross-linking is mandatory.** If a page references a node or system, link to it. Treat orphans as a bug.
- **Live state wins.** When something here disagrees with `pct config` / `docker inspect` / running state, fix the wiki *and* add a changelog entry *in the same session*.
- **Tracked configs.** Pages for configs living in git repos (Caddy, Gitea, Artifacto, mule-image) must note the repo. Edits go through commit+push, never local changes. See [auto-deploy](knowledge/wiki/infrastructure/auto-deploy.md).
- **No secrets.** This is a private repo, but still: reference secret *paths*, never secret *values*.
## Maintaining this wiki
**For agents:** Read [caveman.md](.agents/shared/caveman.md) (terse communication standard). Use templates at [page-templates.md](.agents/shared/page-templates.md) when creating pages.
When you change a node:
1. Update the relevant page (config snapshot, ports, mounts).
2. Add a changelog entry at the bottom of that page.
3. If the change touches a cross-cutting system (DNS, Caddy, Authentik, mesh), update *that* page too and link it from the changelog entry.
4. If it's an incident, add an entry to [`investigations/`](investigations/index.md).
---
## See also
## Updating the Wiki
### When You Change Infrastructure
1. Update the relevant page (config snapshot, ports, mounts, IP address).
2. Add a `### YYYY-MM-DD — title` entry to the page's `## Changelog` section (reverse chronological order).
3. If the change touches a cross-cutting system (DNS, Caddy, Authentik, mesh), update *that* page too and link from the changelog.
4. If it's an incident, add a record to [`knowledge/sources/investigations/`](knowledge/sources/investigations/index.md).
### When You Restructure the Wiki
1. Update the relevant `index.md` / `README.md` in that section.
2. Add a single-line entry to [`knowledge/log.md`](knowledge/log.md): `## [YYYY-MM-DD] <operation> | <summary>` (e.g., `## [2026-07-06] restructure | split infrastructure/dns into dns.md + dns-advanced.md`).
### The Same-Session Update Rule
**Any meaningful state change made in this session requires a wiki update before the session closes.** A change that touches a container page must also update:
- The `containers/index.md` table (IPs, host, mounts, status)
- The root `README.md` table (if affected)
- The Caddy page site list (if affects `*.hubris.network` routing)
- The DNS / ingress infrastructure pages (if affects routing)
- The `hosts/hubris.md` or `hosts/strong.md` page (if container count changes)
- The `inventory.yaml` host entry (source of truth for `hosts/*.yaml` generation)
- The `knowledge/wiki/infrastructure/topology.md` (regenerate if needed)
Not updating all linked places is a bug. See [page-templates.md — same-session update rule](.agents/shared/page-templates.md#same-session-update-rule).
---
## More Information
- **For Hermes agents** → [HERMES.md](.agents/HERMES.md) (persona, source-of-truth hierarchy, token efficiency)
- **For manual workflows** → [.agents/operations/](.agents/operations/) (commands cheatsheet, agent enrollment, Hermes guide)
- **For skills/runbooks** → [.agents/skills/](.agents/skills/) (load the matching SKILL.md before acting; includes risk class + verification)
- **MCP tools** → [AGENTS.md §3](AGENTS.md#3-the-mcp-server) (available tools, when to use MCP vs files)
- **Page templates & voice** → [.agents/shared/](.agents/shared/) (page-templates.md, writing-style.md, caveman.md, llm-wiki.md)
- **Machine-readable substrate** → `inventory.yaml`, `oikos/policy.yaml`, `oikos/ontology.yaml` (not part of the wiki; see [llm-wiki.md](.agents/shared/llm-wiki.md#rules))
- [`CONTRIBUTING.md`](CONTRIBUTING.md) — page templates and tone

View File

@@ -1,8 +0,0 @@
# oapi-codegen config — `make generate` regenerates internal/httpapi/gen.
package: gen
output: internal/httpapi/gen/api.gen.go
generate:
models: true
chi-server: true
strict-server: true
embedded-spec: true

File diff suppressed because it is too large Load Diff

View File

@@ -1,7 +0,0 @@
# Redocly lint config for api/openapi.yaml (CI runs: redocly lint api/openapi.yaml)
extends:
- recommended
rules:
# Every operation declares `default` → RFC 9457 problem+json instead of
# enumerating each 4XX (plan R3-3); oapi-codegen handles `default` fine.
operation-4xx-response: off

View File

@@ -37,21 +37,6 @@ INVENTORY = CONTEXT / "inventory.yaml"
HOSTS_DIR = CONTEXT / "hosts"
AGE_KEY = Path(os.environ.get("SOPS_AGE_KEY_FILE", "/etc/age/key.txt"))
# Oikos kernel modules (policy classification, ontology relations, change
# ledger). Optional at import time so a stale/partial checkout degrades to
# "feature unavailable" instead of crashing every subcommand.
sys.path.insert(0, str(CONTEXT))
try:
from oikos import approve as oikos_approve
from oikos import decide as oikos_decide
from oikos import ledger as oikos_ledger
from oikos import policy as oikos_policy
from oikos import relations as oikos_relations
from oikos import signal as oikos_signal
except ImportError:
oikos_approve = oikos_decide = oikos_ledger = oikos_policy = None
oikos_relations = oikos_signal = None
# ---------- helpers ----------
@@ -187,25 +172,6 @@ def service_backend_host(name: str) -> str:
return service(name)["backend"]
def _record_change(entity: str, action: str, risk: str, *,
result: str | None = None, verification: str | None = None,
approval_ref: str | None = None) -> None:
"""Append a ledger entry and push it standalone (used by mutations that
don't already go through push_inventory, e.g. restart)."""
if oikos_ledger is None:
return
oikos_ledger.append(entity, action, risk, result=result, verification=verification,
approval_ref=approval_ref)
try:
subprocess.run(["git", "add", "ledger/"], check=True, cwd=CONTEXT)
if subprocess.run(["git", "diff", "--cached", "--quiet"], cwd=CONTEXT).returncode != 0:
subprocess.run(["git", "commit", "-m", f"ledger: {entity} {action} ({risk})"],
check=True, cwd=CONTEXT)
subprocess.run(["git", "push"], check=True, cwd=CONTEXT)
except subprocess.CalledProcessError as e:
print(f"warning: could not commit/push ledger entry: {e}", file=sys.stderr)
def push_inventory(message: str, extra_paths: list[str] | None = None) -> None:
"""Stage + commit + push inventory + regenerated hosts/ (+ any extras)."""
subprocess.run(["python3", str(CONTEXT / "mcp" / "build_host_files.py")],
@@ -551,19 +517,6 @@ def cmd_ssh(args: argparse.Namespace) -> int:
os.execvp(cmd[0], cmd)
def cmd_ssh_config(args: argparse.Namespace) -> int:
"""Generate SSH config from inventory.yaml."""
script = CONTEXT / "ssh" / "gen-config.py"
if not script.exists():
die(f"ssh-config generator not found: {script}")
cmd = [sys.executable or "python3", str(script)]
if args.install:
cmd.append("--install")
return subprocess.call(cmd)
def cmd_pct(args: argparse.Namespace) -> int:
lxc = args.lxc
inv = inventory()
@@ -604,31 +557,11 @@ def cmd_restart(args: argparse.Namespace) -> int:
svc = args.service
host_name = service_backend_host(svc)
unit = service(svc).get("systemd_unit", svc)
risk = (oikos_policy.classify_action("service-restart", svc)
if oikos_policy else "reversible_low") or "reversible_low"
approval = oikos_policy.approval_for(risk) if oikos_policy else "none"
# Mechanical gate: config_mutation/destructive risk classes require a
# live grant regardless of -y/interactivity — an agent (or a human
# bypassing the confirm() prompt with -y) cannot mutate a gated service
# without a real oikos/approve.py approval. See oikos/policy.yaml.
if approval != "none":
if not args.approval_id:
die(f"restarting '{svc}' is risk class '{risk}' (approval: {approval}) — "
f"pass --approval-id <id> from an approved 'homelab approval request'")
ok, reason = oikos_approve.check_grant(args.approval_id, f"service:{svc}", "service-restart")
if not ok:
die(f"approval {args.approval_id} not valid for this action: {reason}")
if not args.yes:
if not confirm(f"restart systemd unit '{unit}' on {host_name}?"):
return 1
base = ssh_base(host_name)
rc = subprocess.call(base + ["--", "systemctl", "restart", unit])
_record_change(f"service:{svc}", "restart", risk,
result=("ok" if rc == 0 else f"failed rc={rc}"),
approval_ref=args.approval_id)
return rc
return subprocess.call(base + ["--", "systemctl", "restart", unit])
def cmd_open(args: argparse.Namespace) -> int:
@@ -1085,19 +1018,14 @@ def cmd_refresh_creds(args: argparse.Namespace) -> int:
def cmd_sync(args: argparse.Namespace) -> int:
# Unlike every other mutating command here, this one had no os.geteuid()
# guard — always shelled out to sudo. Fails outright with "No such file
# or directory: 'sudo'" on minimal root-only Linux images (no sudo
# binary installed at all) reached via `ssh root@host`, e.g. strong.
needs_sudo = os.geteuid() != 0
if sys.platform == "darwin":
cmd = ["launchctl", "kickstart", "-k",
"system/network.hubris.homelab-context-sync"]
else:
cmd = ["systemctl", "start", "homelab-context-sync.service"]
if needs_sudo:
cmd = ["sudo"] + cmd
return subprocess.call(cmd)
return subprocess.call(
["sudo", "launchctl", "kickstart", "-k",
"system/network.hubris.homelab-context-sync"]
)
return subprocess.call(
["sudo", "systemctl", "start", "homelab-context-sync.service"]
)
def cmd_mcp(args: argparse.Namespace) -> int:
@@ -1157,11 +1085,9 @@ def cmd_client_add(args: argparse.Namespace) -> int:
print("granting hermes-only secrets...")
_grant_shared_secrets(pubkey, HERMES_SECRETS)
commit_subject = f"client-add: {name} (finalize age_pubkey + grant shared + hermes secrets)"
if oikos_ledger is not None:
oikos_ledger.append(f"host:{name}", "client-add-finalize", "config_mutation", result="ok")
push_inventory(
commit_subject,
extra_paths=[".sops.yaml", "secrets/", "ledger/"],
extra_paths=[".sops.yaml", "secrets/"],
)
print(f"finalized {name}.")
return 0
@@ -1223,11 +1149,9 @@ def cmd_client_remove(args: argparse.Namespace) -> int:
print(f" issuance revoke failed: {e}")
# 4. Commit + push (extras: .sops.yaml + secrets/ may also have changed).
if oikos_ledger is not None:
oikos_ledger.append(f"host:{name}", "client-remove", "destructive", result="ok")
push_inventory(
f"client-remove: {name}",
extra_paths=[".sops.yaml", "secrets/", "ledger/"],
extra_paths=[".sops.yaml", "secrets/"],
)
print()
@@ -1240,211 +1164,6 @@ def cmd_client_remove(args: argparse.Namespace) -> int:
return 0
def _require_oikos() -> None:
if oikos_policy is None or oikos_relations is None or oikos_ledger is None:
die("oikos/ kernel modules not importable — is this checkout up to date?")
def cmd_service(args: argparse.Namespace) -> int:
"""Service Console v0 — explain/health/docs/log/actions/history for one service."""
_require_oikos()
name = args.name
svc = service(name) # dies with a clear message if unknown
if args.action == "explain":
card = CONTEXT / "oikos" / "cards" / f"service-{name}.md"
if not card.exists():
die(f"no context card for {name} — run: python3 oikos/gen-topology.py")
print(card.read_text())
return 0
if args.action == "health":
url = svc.get("url") or svc.get("endpoint")
if not url:
die(f"service {name} has no url/endpoint in inventory")
if not args.live:
try:
from oikos import scheduler as oikos_scheduler
cached = oikos_scheduler.cached_service_health(name)
except ImportError:
cached = None
if cached is not None and cached.get("checked"):
status = "ok" if cached.get("ok") else "unhealthy"
print(f"{name}: {cached.get('checked_url', url)} -> "
f"{cached.get('http_code') or 'no response'} ({status}, "
f"as of {cached['as_of']} — pass --live to force a fresh probe)")
return 0 if cached.get("ok") else 1
proc = subprocess.run(
["curl", "-sS", "-o", "/dev/null", "-w", "%{http_code}", "--max-time", "5", url],
capture_output=True, text=True,
)
code = proc.stdout.strip() or "no response"
print(f"{name}: {url} -> {code} (live probe)")
return 0 if code.startswith(("2", "3")) else 1
if args.action == "docs":
doc = svc.get("doc_page")
if not doc:
die(f"no doc_page recorded for {name} in inventory.yaml")
path = CONTEXT / doc
if not path.exists():
die(f"doc_page {doc} does not exist")
print(path.read_text())
return 0
if args.action == "log":
return cmd_logs(argparse.Namespace(service=name, lines=args.lines, follow=False))
if args.action == "actions":
for a in oikos_policy.safe_actions_for_service(name, svc):
print(f"{a['action']:<24} {a['risk']:<16} approval={a['approval']}")
return 0
if args.action == "history":
entries = oikos_ledger.history(f"service:{name}", limit=args.limit)
if not entries:
print(f"(no ledger entries for service:{name} yet)")
for e in entries:
print(json.dumps(e))
return 0
die(f"unknown service action: {args.action}")
def cmd_change_preflight(args: argparse.Namespace) -> int:
"""Dry-run report before mutating a service: health, risk class, approval
requirement, and the verification command to run after."""
_require_oikos()
name = args.service
svc = service(name)
risk = (oikos_policy.classify_action("tracked-config-edit", name)
if svc.get("config_repo")
else oikos_policy.classify_action("service-restart", name)) or "config_mutation"
approval = oikos_policy.approval_for(risk)
print(f"Preflight: {name}")
print(f" risk class: {risk} (approval: {approval})")
url = svc.get("url") or svc.get("endpoint")
if url:
proc = subprocess.run(
["curl", "-sS", "-o", "/dev/null", "-w", "%{http_code}", "--max-time", "3", url],
capture_output=True, text=True,
)
print(f" current health: {url} -> {proc.stdout.strip() or 'no response'}")
if svc.get("config_repo"):
print(f" config repo: {svc['config_repo']} "
f"(verify the backend's working tree is clean before editing)")
if svc.get("risk_notes"):
print(f" risk notes: {svc['risk_notes']}")
print(f" verification after change: "
+ (f"curl -sf {url}" if url else f"homelab logs {name}"))
if approval != "none":
print(f" requires operator approval before mutating ({approval})")
return 0
def cmd_node_relations(args: argparse.Namespace) -> int:
"""Walk the ontology graph both directions for a host or service name."""
_require_oikos()
results = oikos_relations.relations_for_name(args.name)
if not results:
die(f"unknown entity: {args.name}")
for r in results:
print(f"entity: {r['entity']}")
print(f" impacts: {', '.join(r['impacts']) or '(none)'}")
print(f" affected by: {', '.join(r['affected_by']) or '(none)'}")
print(f" full blast radius: {', '.join(r['blast_radius']) or '(none)'}")
return 0
def cmd_decide(args: argparse.Namespace) -> int:
"""Route a proposed action: auto-act or escalate. See oikos/decide.py."""
_require_oikos()
result = oikos_decide.classify(args.action, args.entity, service_name=args.service_name,
record=not args.no_record)
print(json.dumps(result, indent=2))
return 0 if result["route"] == "auto-act" else 1
def cmd_approval_request(args: argparse.Namespace) -> int:
_require_oikos()
entry = oikos_approve.request(
args.entity, args.action, args.risk, args.evidence,
verification=args.verification, requires_phrase=args.requires_phrase,
ttl_hours=args.ttl_hours,
)
print(json.dumps({k: v for k, v in entry.items() if k != "matrix_message"}, indent=2))
print()
print("--- post this to Matrix ---")
print(entry["matrix_message"])
return 0
def cmd_approval_list(args: argparse.Namespace) -> int:
_require_oikos()
for e in oikos_approve.list_approvals(state=args.state):
print(json.dumps(e))
return 0
def cmd_approval_reply(args: argparse.Namespace) -> int:
_require_oikos()
try:
entry = oikos_approve.reply(args.id, args.decision, phrase=args.phrase,
decided_by=args.decided_by)
except (ValueError, RuntimeError) as e:
die(str(e))
print(json.dumps(entry, indent=2))
return 0
def cmd_approval_check(args: argparse.Namespace) -> int:
_require_oikos()
ok, reason = oikos_approve.check_grant(args.id, args.entity, args.action)
print(f"{'GRANTED' if ok else 'DENIED'}: {reason}")
return 0 if ok else 1
def cmd_signal_raise(args: argparse.Namespace) -> int:
_require_oikos()
action = None
if args.action_runbook or args.action_risk:
action = {"runbook": args.action_runbook, "risk": args.action_risk}
entry = oikos_signal.raise_signal(args.kind, args.severity, args.entity, args.evidence,
likely_cause=args.likely_cause,
recommended_action=action,
verification=args.verification)
print(json.dumps(entry, indent=2))
return 0
def cmd_signal_list(args: argparse.Namespace) -> int:
_require_oikos()
for e in oikos_signal.list_signals(state=args.state, entity=args.entity,
severity=args.severity, kind=args.kind):
print(json.dumps(e))
return 0
def cmd_signal_ack(args: argparse.Namespace) -> int:
_require_oikos()
print(json.dumps(oikos_signal.acknowledge(args.id, args.note), indent=2))
return 0
def cmd_signal_resolve(args: argparse.Namespace) -> int:
_require_oikos()
print(json.dumps(oikos_signal.resolve(args.id, args.note), indent=2))
return 0
def cmd_signal_mute(args: argparse.Namespace) -> int:
_require_oikos()
print(json.dumps(oikos_signal.mute(args.id, args.ttl_hours, args.note), indent=2))
return 0
def cmd_nuke(args: argparse.Namespace) -> int:
name = args.name
if not args.yes:
@@ -1740,12 +1459,6 @@ def main() -> int:
sp.add_argument("command", nargs=argparse.REMAINDER)
sp.set_defaults(func=cmd_ssh)
sp = sub.add_parser("ssh-config",
help="generate ~/.ssh/config.d/homelab from inventory.yaml")
sp.add_argument("--install", "-i", action="store_true",
help=f"write to ~/.ssh/config.d/homelab and wire Include into main config")
sp.set_defaults(func=cmd_ssh_config)
sp = sub.add_parser("pct", help="proxy pct commands via ssh to hubris")
sp.add_argument("lxc")
sp.add_argument("action")
@@ -1762,9 +1475,6 @@ def main() -> int:
sp = sub.add_parser("restart", help="restart a service")
sp.add_argument("service")
sp.add_argument("--yes", "-y", action="store_true")
sp.add_argument("--approval-id", default=None,
help="required if the service's risk class needs approval "
"(see 'homelab approval request')")
sp.set_defaults(func=cmd_restart)
sp = sub.add_parser("open", help="open a service's URL in browser")
@@ -1820,101 +1530,6 @@ def main() -> int:
help="skip the pre-flight dpkg-audit gate AND proceed past snapshot failures")
sp.set_defaults(func=cmd_apt_upgrade)
sp = sub.add_parser("service", help="Service Console v0 — explain/health/docs/log/actions/history")
sp.add_argument("name")
sp.add_argument("action", choices=["explain", "health", "docs", "log", "actions", "history"])
sp.add_argument("--lines", "-n", type=int, default=200, help="for 'log'")
sp.add_argument("--limit", type=int, default=20, help="for 'history'")
sp.add_argument("--live", action="store_true",
help="for 'health': force a fresh probe instead of the scheduler's cache")
sp.set_defaults(func=cmd_service)
change = sub.add_parser("change", help="change/mutation workflow")
chsub = change.add_subparsers(dest="action", required=True)
ch_preflight = chsub.add_parser("preflight")
ch_preflight.add_argument("service")
ch_preflight.set_defaults(func=cmd_change_preflight)
sp = sub.add_parser("node", help="ontology queries on a host/service")
sp.add_argument("name")
sp.add_argument("action", choices=["relations"])
sp.set_defaults(func=cmd_node_relations)
sp = sub.add_parser("decide", help="classify a proposed action: auto-act or escalate")
sp.add_argument("action")
sp.add_argument("entity")
sp.add_argument("--service-name", default=None)
sp.add_argument("--no-record", action="store_true",
help="skip writing this classification to the change ledger")
sp.set_defaults(func=cmd_decide)
approval = sub.add_parser("approval", help="approval-engine requests (escalate route)")
apsub = approval.add_subparsers(dest="action", required=True)
ap_req = apsub.add_parser("request")
ap_req.add_argument("entity")
ap_req.add_argument("action")
ap_req.add_argument("risk")
ap_req.add_argument("evidence")
ap_req.add_argument("--verification")
ap_req.add_argument("--requires-phrase", action="store_true")
ap_req.add_argument("--ttl-hours", type=int, default=24)
ap_req.set_defaults(func=cmd_approval_request)
ap_list = apsub.add_parser("list")
ap_list.add_argument("--state", choices=["pending", "approved", "denied", "expired", "executed"])
ap_list.set_defaults(func=cmd_approval_list)
ap_reply = apsub.add_parser("reply")
ap_reply.add_argument("id")
ap_reply.add_argument("decision", choices=["approve", "deny"])
ap_reply.add_argument("--phrase")
ap_reply.add_argument("--decided-by")
ap_reply.set_defaults(func=cmd_approval_reply)
ap_check = apsub.add_parser("check")
ap_check.add_argument("id")
ap_check.add_argument("entity")
ap_check.add_argument("action")
ap_check.set_defaults(func=cmd_approval_check)
signal = sub.add_parser("signal", help="the attention layer (oikos/signal.py)")
sigsub = signal.add_subparsers(dest="action", required=True)
sig_raise = sigsub.add_parser("raise")
sig_raise.add_argument("kind")
sig_raise.add_argument("severity", choices=["info", "warning", "critical"])
sig_raise.add_argument("entity")
sig_raise.add_argument("evidence")
sig_raise.add_argument("--likely-cause")
sig_raise.add_argument("--action-runbook")
sig_raise.add_argument("--action-risk")
sig_raise.add_argument("--verification")
sig_raise.set_defaults(func=cmd_signal_raise)
sig_list = sigsub.add_parser("list")
sig_list.add_argument("--state", choices=["raised", "acknowledged", "acting", "resolved", "muted"])
sig_list.add_argument("--entity")
sig_list.add_argument("--severity", choices=["info", "warning", "critical"])
sig_list.add_argument("--kind")
sig_list.set_defaults(func=cmd_signal_list)
sig_ack = sigsub.add_parser("ack")
sig_ack.add_argument("id")
sig_ack.add_argument("--note")
sig_ack.set_defaults(func=cmd_signal_ack)
sig_resolve = sigsub.add_parser("resolve")
sig_resolve.add_argument("id")
sig_resolve.add_argument("--note")
sig_resolve.set_defaults(func=cmd_signal_resolve)
sig_mute = sigsub.add_parser("mute")
sig_mute.add_argument("id")
sig_mute.add_argument("--ttl-hours", type=int, default=24)
sig_mute.add_argument("--note")
sig_mute.set_defaults(func=cmd_signal_mute)
sp = sub.add_parser("nuke", help="shred /etc/age/key.txt + /opt/homelab-context on a host")
sp.add_argument("name")
sp.add_argument("--yes", "-y", action="store_true")
@@ -1929,7 +1544,7 @@ def main() -> int:
csub_add.add_argument("--with-hermes", action="store_true",
help="also grant secrets/openrouter-api-key.yaml so this "
"host can run the Hermes agent (see "
".agents/operations/hermes-agent.md). Combine with --finalize-pubkey.")
"operations/hermes-agent.md). Combine with --finalize-pubkey.")
csub_add.set_defaults(func=cmd_client_add)
csub_rm = csub.add_parser("remove")
csub_rm.add_argument("name")

View File

@@ -7,11 +7,7 @@
# curl ... | sudo bash -s -- --with-mcp # also wire Claude's .mcp.json
# curl ... | sudo bash -s -- --with-hermes # also install Goose + Hermes wrapper
# curl ... | sudo bash -s -- --dry-run # show what would happen
# curl ... | sudo bash -s -- --no-secrets # skip age-key issuance entirely
# curl ... | sudo bash -s -- --no-mesh # get secrets over LAN only, skip
# # installing/connecting Netbird
# # (host must be on 192.168.8.0/24
# # or otherwise reach secrets.hubris.network)
# curl ... | sudo bash -s -- --no-secrets # skip age-key issuance
#
# Prerequisites the script verifies:
# - running as root
@@ -27,7 +23,7 @@ REPO_HTTPS="${HOMELAB_REPO_URL:-https://git.hubris.network/dtoro/Homelab-Docs.gi
CLONE_DIR="${HOMELAB_CONTEXT_DIR:-/opt/homelab-context}"
ISSUANCE_URL_NETBIRD="${HOMELAB_ISSUANCE_NETBIRD:-https://secrets.hubris.network/issue}"
ISSUANCE_URL_TAILSCALE="${HOMELAB_ISSUANCE_TAILSCALE:-https://secrets.hubris.network/issue}"
MCP_URL="${HOMELAB_MCP_URL:-https://mcp.hubris.network/mcp}"
MCP_URL="${HOMELAB_MCP_URL:-https://mcp.hubris.network/sse}"
HERMES_MCP_URI="${HOMELAB_HERMES_MCP_URI:-https://mcp.hubris.network/mcp}"
HERMES_MODEL="${HOMELAB_HERMES_MODEL:-nousresearch/hermes-4-405b}"
@@ -35,7 +31,6 @@ WITH_MCP=0
WITH_HERMES=0
DRY_RUN=0
NO_SECRETS=0
NO_MESH=0
GITEA_TOKEN="${HOMELAB_GITEA_TOKEN:-}"
GITEA_USER="${HOMELAB_GITEA_USER:-dtoro}"
@@ -46,7 +41,6 @@ while [ $# -gt 0 ]; do
--with-hermes) WITH_HERMES=1; shift ;;
--dry-run) DRY_RUN=1; shift ;;
--no-secrets) NO_SECRETS=1; shift ;;
--no-mesh) NO_MESH=1; shift ;;
--gitea-token) GITEA_TOKEN="$2"; shift 2 ;;
--gitea-user) GITEA_USER="$2"; shift 2 ;;
--help|-h)
@@ -89,21 +83,6 @@ run() {
fi
}
# Run a command as the enrolling human user when one exists (i.e. this
# script was invoked via `sudo bash bootstrap.sh` from a real login), and
# directly otherwise. Minimal Linux images (bare Proxmox/Debian installs
# reached via `ssh root@host`) often don't even have a `sudo` binary
# installed — calling `sudo -u root ...` on those unconditionally fails
# with "sudo: command not found" even though we're already root and don't
# need to switch users at all.
run_as() {
if [ -n "${SUDO_USER:-}" ] && [ "$SUDO_USER" != "root" ]; then
sudo -u "$SUDO_USER" -- "$@"
else
"$@"
fi
}
# -------- preflight --------
if [ "$(id -u)" -ne 0 ]; then
echo "bootstrap.sh must run as root (use sudo)." >&2
@@ -139,22 +118,6 @@ fi
if [ "$NO_SECRETS" -eq 0 ]; then
for cmd in age sops; do command -v "$cmd" >/dev/null || missing+=("$cmd"); done
fi
# sops isn't a real Debian/Fedora package (there is no apt/dnf "sops"), so it
# always needs the direct-binary-download path, on both distros. Only Darwin
# (brew) can install it via a package manager.
install_sops_binary() {
local sops_version=v3.9.4
local arch
arch="$(uname -m)"
case "$arch" in
x86_64|amd64) arch=amd64 ;;
aarch64|arm64) arch=arm64 ;;
*) echo "[bootstrap] unsupported arch for sops binary download: $arch" >&2; return 1 ;;
esac
curl -fsSL "https://github.com/getsops/sops/releases/download/${sops_version}/sops-${sops_version}.linux.${arch}" \
-o /usr/local/bin/sops && chmod +x /usr/local/bin/sops
}
if [ "${#missing[@]}" -gt 0 ]; then
if [ "$DRY_RUN" -eq 1 ]; then
echo "+ would install missing tools: ${missing[*]}"
@@ -175,23 +138,13 @@ if [ "${#missing[@]}" -gt 0 ]; then
for m in "${missing[@]}"; do
case "$m" in
python3-yaml) dnf_list+=("python3-pyyaml") ;;
sops) install_sops_binary ;;
*) dnf_list+=("$m") ;;
esac
done
[ "${#dnf_list[@]}" -gt 0 ] && dnf install -y "${dnf_list[@]}"
dnf install -y "${dnf_list[@]}"
elif command -v apt-get >/dev/null 2>&1; then
apt_list=()
for m in "${missing[@]}"; do
case "$m" in
sops) install_sops_binary ;;
*) apt_list+=("$m") ;;
esac
done
if [ "${#apt_list[@]}" -gt 0 ]; then
DEBIAN_FRONTEND=noninteractive apt-get update
DEBIAN_FRONTEND=noninteractive apt-get install -y "${apt_list[@]}"
fi
DEBIAN_FRONTEND=noninteractive apt-get update
DEBIAN_FRONTEND=noninteractive apt-get install -y "${missing[@]}"
else
echo "[bootstrap] no supported package manager for: ${missing[*]}" >&2
echo "[bootstrap] install with your package manager + re-run" >&2
@@ -212,14 +165,10 @@ if [ "${#missing[@]}" -gt 0 ]; then
fi
# -------- ensure netbird is installed + connected (workstation/VM hosts) --------
# Skipped on --no-secrets (LXCs that route via the LAN already), --no-mesh
# (explicit opt-out — secrets issuance still works if the mesh check below
# falls back to LAN reachability), and --dry-run. Installs netbird if
# missing, then drives `netbird up` against the homelab management server.
# The operator clicks the printed device-code URL once — this blocks
# indefinitely if nobody approves it, so don't skip --no-mesh on a host
# nobody's watching interactively.
if [ "$NO_SECRETS" -eq 0 ] && [ "$DRY_RUN" -eq 0 ] && [ "$NO_MESH" -eq 0 ]; then
# Skipped on --no-secrets (LXCs that route via the LAN already) and --dry-run.
# Installs netbird if missing, then drives `netbird up` against the homelab
# management server. The operator clicks the printed device-code URL once.
if [ "$NO_SECRETS" -eq 0 ] && [ "$DRY_RUN" -eq 0 ]; then
if ! command -v netbird >/dev/null 2>&1 && ! command -v tailscale >/dev/null 2>&1; then
echo "[bootstrap] no mesh CLI found; installing netbird..."
if [ "$OS" = "Darwin" ]; then
@@ -481,9 +430,9 @@ if [ "$WITH_HERMES" -eq 1 ]; then
echo "+ would run upstream goose installer and symlink to /usr/local/bin/goose"
else
# Upstream installer drops the binary at ~/.local/bin/goose for the
# invoking user. We run it as $H_USER (via run_as) then symlink
# system-wide.
run_as env CONFIGURE=false \
# invoking user. We run it as $H_USER then symlink system-wide.
sudo -u "$H_USER" \
env CONFIGURE=false \
bash -c 'curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash'
if [ -x "$H_HOME/.local/bin/goose" ]; then
ln -sfn "$H_HOME/.local/bin/goose" /usr/local/bin/goose
@@ -507,7 +456,7 @@ if [ "$WITH_HERMES" -eq 1 ]; then
Linux) HERMES_LINK=/root/HERMES.md ;;
Darwin) HERMES_LINK=/etc/HERMES.md ;;
esac
run "ln -sfn '$CLONE_DIR/.agents/HERMES.md' '$HERMES_LINK'"
run "ln -sfn '$CLONE_DIR/HERMES.md' '$HERMES_LINK'"
echo "[bootstrap] linked HERMES.md → $HERMES_LINK"
# 4. Drop the Goose config. Idempotent YAML merge — preserves any keys the
@@ -563,7 +512,7 @@ PYEOF
# 5. Symlink HERMES.md as the global .goosehints — Goose injects it into
# the system prompt on every session start.
run "ln -sfn '$CLONE_DIR/.agents/HERMES.md' '$GOOSEHINTS'"
run "ln -sfn '$CLONE_DIR/HERMES.md' '$GOOSEHINTS'"
if [ "$DRY_RUN" -eq 0 ]; then
chown -h "$H_USER" "$GOOSEHINTS" 2>/dev/null || true
fi
@@ -646,7 +595,7 @@ if [ "$HKIND" != "lxc" ] && [ "$HKIND" != "vm" ]; then
# Make sure pipx is available; OS-specific install.
if ! command -v pipx >/dev/null 2>&1; then
if [ "$OS" = "Darwin" ] && command -v brew >/dev/null 2>&1; then
run_as brew install pipx 2>&1 | tail -2 || true
sudo -u "${SUDO_USER:-$USER}" brew install pipx 2>&1 | tail -2 || true
elif command -v dnf >/dev/null 2>&1; then
dnf install -y pipx 2>&1 | tail -2 || true
elif command -v apt-get >/dev/null 2>&1; then
@@ -654,9 +603,9 @@ if [ "$HKIND" != "lxc" ] && [ "$HKIND" != "vm" ]; then
fi
fi
if command -v pipx >/dev/null 2>&1; then
INVOKING_USER="${SUDO_USER:-root}"
run_as bash -lc "pipx install 'mcp[cli]'" 2>&1 | tail -3 || true
run_as bash -lc "pipx ensurepath" >/dev/null 2>&1 || true
INVOKING_USER="${SUDO_USER:-$USER}"
sudo -u "$INVOKING_USER" -- bash -lc "pipx install 'mcp[cli]'" 2>&1 | tail -3 || true
sudo -u "$INVOKING_USER" -- bash -lc "pipx ensurepath" >/dev/null 2>&1 || true
echo "[bootstrap] mcp CLI installed for $INVOKING_USER via pipx"
else
echo "[bootstrap] WARNING: pipx unavailable; install manually: pipx install 'mcp[cli]'" >&2

View File

@@ -1,254 +0,0 @@
package main
import (
"context"
"fmt"
"log/slog"
"os"
"os/signal"
"syscall"
"net/http"
"github.com/dtoro/oikos/internal/config"
"github.com/dtoro/oikos/internal/db"
"github.com/dtoro/oikos/internal/httpapi"
"github.com/dtoro/oikos/internal/observability"
"github.com/jackc/pgx/v5"
)
// SchedulerRunner is set by the scheduler init() to avoid circular imports.
var SchedulerRunner func(context.Context, *db.Pool, config.Config)
// NotifierRunner is set by the notifier init() to avoid circular imports.
var NotifierRunner func(context.Context, *db.Pool, config.Config)
func main() {
if len(os.Args) < 2 {
usage()
os.Exit(1)
}
role := os.Args[1]
cfg := config.FromEnv()
// Structured logging (slog)
logger := observability.NewLogger(cfg.Debug)
slog.SetDefault(logger)
slog.Info("starting oikos", "role", role, "config", cfg)
ctx, cancel := signal.NotifyContext(context.Background(),
syscall.SIGTERM, syscall.SIGINT)
defer cancel()
switch role {
case "migrate":
if err := runMigrate(ctx, cfg); err != nil {
slog.Error("migrate failed", "error", err)
os.Exit(1)
}
case "seed":
if err := runSeed(ctx, cfg); err != nil {
slog.Error("seed failed", "error", err)
os.Exit(1)
}
case "export":
if err := runExport(ctx, cfg); err != nil {
slog.Error("export failed", "error", err)
os.Exit(1)
}
case "api":
if err := runAPI(ctx, cfg); err != nil {
slog.Error("api failed", "error", err)
os.Exit(1)
}
case "scheduler":
if SchedulerRunner != nil {
SchedulerRunner(ctx, nil, cfg)
} else {
slog.Error("scheduler not compiled in (import internal/scheduler)")
os.Exit(1)
}
case "notifier":
if NotifierRunner != nil {
NotifierRunner(ctx, nil, cfg)
} else {
slog.Error("notifier not compiled in (import internal/notifier)")
os.Exit(1)
}
case "all":
slog.Info("all role not yet implemented (runs api + scheduler + notifier in one process)")
os.Exit(1)
case "version":
fmt.Println("oikos dev (Phase 1)")
case "help", "--help", "-h":
usage()
default:
fmt.Fprintf(os.Stderr, "unknown role: %s\n", role)
usage()
os.Exit(1)
}
}
func usage() {
fmt.Println(`oikos — the homelab OS
Usage: oikos <role> [flags]
Roles:
migrate Run database migrations (forward-only, idempotent)
seed Ingest seed YAML files into the database
export Export DB state back to seed YAMLs (DR / version control)
api Run the REST + MCP API server (Phase 2)
scheduler Run the observe + act loop (Phase 3)
notifier Run the notification service (Phase 3)
all Run all roles in one process (dev mode)
version Print version info
Environment:
OIKOS_DATABASE_URL Postgres connection string
OIKOS_API_LISTEN API listen address (default :8090)
OIKOS_ENV Environment (dev, prod)
OIKOS_DEBUG Enable verbose logging (true/1)
OIKOS_SEEDS_DIR Path to seeds directory (default: seeds)
OIKOS_MCP_BEARER_TOKEN Shared secret for MCP auth`)
}
func runMigrate(ctx context.Context, cfg config.Config) error {
pool, err := db.New(ctx, cfg.DatabaseURL)
if err != nil {
return err
}
defer pool.Close()
slog.Info("running migrations")
if err := pool.Migrate(ctx); err != nil {
return err
}
slog.Info("migrations complete")
return nil
}
func runSeed(ctx context.Context, cfg config.Config) error {
pool, err := db.New(ctx, cfg.DatabaseURL)
if err != nil {
return err
}
defer pool.Close()
// Ensure migrations are applied first
if err := pool.Migrate(ctx); err != nil {
return fmt.Errorf("migrations: %w", err)
}
seedsDir := cfg.SeedsDir
if seedsDir == "" {
seedsDir = "seeds"
}
// Ingest ontology seed
ontoContent, err := os.ReadFile(seedsDir + "/ontology.yaml")
if err != nil {
return fmt.Errorf("read ontology seed: %w", err)
}
err = pool.SeedIngest(ctx, "ontology.yaml", ontoContent,
func(ctx context.Context, tx pgx.Tx, data map[string]any) error {
r, err := db.IngestOntologySeed(ctx, tx, data)
if err != nil {
return err
}
slog.Info("ontology ingested",
"lifecycles", r.Lifecycles,
"entity_types", r.EntityTypes,
"relationship_types", r.RelationshipTypes)
return nil
})
if err != nil {
return err
}
// Ingest inventory seed
invContent, err := os.ReadFile(seedsDir + "/inventory.yaml")
if err != nil {
return fmt.Errorf("read inventory seed: %w", err)
}
err = pool.SeedIngest(ctx, "inventory.yaml", invContent,
func(ctx context.Context, tx pgx.Tx, data map[string]any) error {
r, err := db.IngestInventorySeed(ctx, tx, data)
if err != nil {
return err
}
slog.Info("inventory ingested",
"entities", r.Entities,
"relationships", r.Relationships)
return nil
})
if err != nil {
return err
}
// Ingest policy seed
polContent, err := os.ReadFile(seedsDir + "/policy.yaml")
if err != nil {
return fmt.Errorf("read policy seed: %w", err)
}
err = pool.SeedIngest(ctx, "policy.yaml", polContent,
func(ctx context.Context, tx pgx.Tx, data map[string]any) error {
r, err := db.IngestPolicySeed(ctx, tx, data)
if err != nil {
return err
}
slog.Info("policy ingested",
"risk_classes", r.RiskClasses,
"approval_rules", r.ApprovalRules,
"autonomy_settings", r.AutonomySettings)
return nil
})
if err != nil {
return err
}
slog.Info("seed ingest complete")
return nil
}
func runAPI(ctx context.Context, cfg config.Config) error {
pool, err := db.New(ctx, cfg.DatabaseURL)
if err != nil {
return err
}
defer pool.Close()
if err := pool.Migrate(ctx); err != nil {
return fmt.Errorf("migrations: %w", err)
}
err = httpapi.ListenAndServe(ctx, pool, cfg)
if err == http.ErrServerClosed {
return nil
}
return err
}
func runExport(ctx context.Context, cfg config.Config) error {
pool, err := db.New(ctx, cfg.DatabaseURL)
if err != nil {
return err
}
defer pool.Close()
exports, err := db.ExportToYAML(ctx, pool)
if err != nil {
return err
}
for name, content := range exports {
path := cfg.SeedsDir + "/" + name
if err := os.WriteFile(path, content, 0644); err != nil {
return fmt.Errorf("write %s: %w", path, err)
}
slog.Info("exported", "file", path, "bytes", len(content))
}
return nil
}

View File

@@ -1,21 +0,0 @@
# Multi-stage Dockerfile for Oikos (ADR 0001: single binary)
FROM golang:1.26-alpine AS builder
RUN apk add --no-cache git ca-certificates
WORKDIR /build
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -o /oikos -tags timetzdata -ldflags="-s -w" ./cmd/oikos
# --- Runtime: distroless static ---
FROM gcr.io/distroless/static:nonroot
COPY --from=builder /oikos /oikos
COPY --from=builder /build/seeds /seeds
COPY --from=builder /build/migrations /migrations
ENTRYPOINT ["/oikos"]

View File

@@ -0,0 +1,34 @@
# 101 — `jellyfin`
Media server: serves the movies / TV / anime / music / audiobooks / podcasts libraries from `/mnt/library` to LAN clients.
## At a glance
- **Hostname:** `jellyfin`
- **IP:** `192.168.8.206`
- **Privilege:** **unprivileged** + idmap (so it can write to the `media` group on `/mnt/library`)
- **Resources:** 2 cores / 4 GiB RAM / 16 GiB rootfs
- **Mounts:** `/mnt/library``/mnt/library`
- **Public hostname:** [`media.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md) → `:8096`
## Service / port map
| Service | Listen | Notes |
| -------- | ------ | ----- |
| jellyfin | `:8096` | HTTP (caddy terminates TLS) |
## Permissions
Member of the [media GID 10000](../infrastructure/media-permissions.md) standard. Service user `jellyfin` is in the `media` group inside the container; idmap block in `/etc/pve/lxc/101.conf` maps in-container GID 10000 to host GID 10000.
## Related
- [Caddy reverse proxy](121-caddy.md)
- [Media permissions](../infrastructure/media-permissions.md)
- [arriman](122-arriman.md) — \*arr stack writes the libraries jellyfin reads
- [DNS split-horizon](../infrastructure/dns.md)
## Changelog
### 2026-04-28 — wiki entry created
Initial documentation. No config changes.
### 2026-04-20 — joined the `media` GID 10000 standard
Idmap block applied; in-container `media` group at GID 10000 mapped to host GID 10000. See [media permissions](../infrastructure/media-permissions.md). Config backup: `/root/101.conf.bak.*`.

View File

@@ -8,7 +8,7 @@ Dedicated, single-purpose LXC that re-exports `/mnt/library` over NFSv4 to clien
- **LAN DNS:** `nfs-export.hubris.network``192.168.8.200` (direct, no Caddy)
- **Privilege:** privileged (`unprivileged: 0`) + `lxc.apparmor.profile: unconfined` — required for `nfs-kernel-server`
- **Resources:** 1 core / 512 MiB RAM / 2 GiB rootfs / 256 MiB swap
- **Mounts:** host `/mnt/library` ↔ container `/mnt/library` (same path on both sides — matches the bind-mount convention used by jellyfin, paperless, arriman, nextcloud, mule-images, apps)
- **Mounts:** host `/mnt/library` ↔ container `/mnt/library` (same path on both sides — matches the bind-mount convention used by jellyfin, paperless, arriman, nextcloud, mule-images, plato, apps)
## What it does
@@ -54,7 +54,7 @@ We considered three options before building this:
| Option | Outcome |
|---|---|
| **NFS on hubris bare-metal host** | Best performance, but adds long-lived NFS/RPC daemons to a host with a recent crash episode ([hubris crash 2026-04-21/22](../../sources/investigations/index.md)). Rejected. |
| **NFS on hubris bare-metal host** | Best performance, but adds long-lived NFS/RPC daemons to a host with a recent crash episode ([hubris crash 2026-04-21/22](../investigations/index.md)). Rejected. |
| **SMB on host** | Same host-blast-radius problem, plus 3050% lower throughput than NFS on Linux↔Linux. Rejected. |
| **NFS in a dedicated LXC** ← this | Within ~2% of host performance (LXC is namespace isolation; IO path is unchanged), zero new daemons on hubris, matches the existing fleet pattern. Selected. |

View File

@@ -18,16 +18,16 @@ Paperless-ngx for document management. Ingests scans / PDFs from `/mnt/library/d
| paperless-task-queue, paperless-scheduler, paperless-consumer | — | systemd workers |
## Auth
Behind [Authentik forward-auth](106-auth-outpost.md). API path `/api/*` bypasses forward-auth (mobile clients can't follow the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates matching users on first SSO login.
Behind [Authentik forward-auth](124-authentik.md). API path `/api/*` bypasses forward-auth (mobile clients can't follow the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates matching users on first SSO login.
## Storage
- Documents at `/mnt/library/documents` (owner `www-data:www-data`, mode 750 — *not* on the `media` group, by design).
## Known issues
- ~~Disk usage was 86.9% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
- Disk usage was 86.9% at last claudio-monitor reading on 2026-04-21. Monitor or grow rootfs.
## Related
- [Authentik](106-auth-outpost.md)
- [Authentik](124-authentik.md)
- [Caddy](121-caddy.md)
- [DNS](../infrastructure/dns.md)
- [Monitoring](../infrastructure/monitoring.md)

View File

@@ -45,9 +45,6 @@ LXC has `/etc/hosts` override mapping `auth.hubris.network → 192.168.8.175` (r
## Changelog
### 2026-06-24 — terminalito deploy webhook (id 12)
Push webhook on `dtoro/terminalito``http://192.168.8.211:9797/deploy` ([trmnl (128)](128-trmnl.md)); `app.ini` `ALLOWED_HOST_LIST` extended with `192.168.8.211`. See [auto-deploy](../infrastructure/auto-deploy.md).
### 2026-04-28 — wiki entry created
Initial documentation.
@@ -55,7 +52,7 @@ Initial documentation.
Added `192.168.8.205`. See [Artifacto auto-deploy on apps (105)](105-apps.md).
### 2026-04-21 — `/etc/hosts` override for `auth.hubris.network` added
For OIDC integration with [authentik (124)](106-auth-outpost.md). Outside the PVE markers, with a hubris-hosts-override.service for idempotency.
For OIDC integration with [authentik (124)](124-authentik.md). Outside the PVE markers, with a hubris-hosts-override.service for idempotency.
### 2026-04-20 — gitea customizations + auto-deploy pipeline shipped
`dtoro/gitea-customizations` repo created; webhook receiver at loopback `:9797` validates HMAC and runs `deploy.sh`. CAD and PlantUML loaders live in `footer.tmpl`.

View File

@@ -1,6 +1,6 @@
# 105 — `apps`
Docker host for everything that doesn't justify its own LXC. Currently runs Artifacto, PlantUML server, Portainer (and historically WriteFreely / blog), plus the [homelab-context distribution services](../infrastructure/homelab-context.md) (MCP + secrets-issuance) since 2026-05-20. Booklore migrated to [grimmory (130)](130-grimmory.md) on 2026-06-29.
Docker host for everything that doesn't justify its own LXC. Currently runs Artifacto, Booklore, PlantUML server, Portainer (and historically WriteFreely / blog), plus the [homelab-context distribution services](../infrastructure/homelab-context.md) (MCP + secrets-issuance) since 2026-05-20.
## At a glance
- **Hostname:** `apps`
@@ -15,6 +15,7 @@ Docker host for everything that doesn't justify its own LXC. Currently runs Arti
| Hostname | Container | Backend port | Notes |
| --------------------------------- | ---------------- | ------------ | ----- |
| `docker.hubris.network` | Portainer | `:9443` | Native OAuth2 via Authentik. Trusted-origins requires hostname only (no scheme/port). |
| `books.hubris.network` | Booklore | `:6060` | Native OIDC. Redirect URI `/oauth2-callback`. |
| `artifacto.hubris.network` | Artifacto | `:3100` | Public `/p/*`, `/static/*`, `/healthz` exposed via [VPS traefik](../infrastructure/ingress.md). |
| `blog.hubris.network` | WriteFreely | `:8080` | Native OIDC via `[oauth.generic]`. |
| `git.hubris.network/_plantuml/*` | PlantUML server | `:8079` | Same-origin route from [gitea (104)](104-gitea.md). |
@@ -45,6 +46,11 @@ Receiver at `/opt/artifacto-deploy/` (outside the app repo): `deploy.sh` + `webh
### Portainer
Native OAuth2 (Settings → Authentication → OAuth → Custom). Manual endpoints (no OIDC discovery). Uses `portainer-uid` custom-claim scope from Authentik. Container is **not** compose-managed — safe to `docker run` recreate; data lives in named volume `portainer_data`. CLI flag: `--trusted-origins docker.hubris.network` (hostname only — `IsTrustedOrigin` rejects strings containing `://`).
### Booklore
Native OIDC via Authentik (Settings → OIDC). Redirect URI `/oauth2-callback` (NOT `/api/oidc`). Container needs `extra_hosts: auth.hubris.network:192.168.8.175`. **Edit via Portainer UI** if it's a Portainer-managed stack.
> ⚠️ **Never `docker compose up` Portainer-managed stacks from the host shell.** Portainer's compose state lives at `/var/lib/docker/volumes/portainer_data/_data/compose/<N>/`. Running `docker compose up -d <svc>` from the host triggers recreates of OTHER services in the stack and silently destroys bind-mounted data. **This wiped Booklore's mariadb data on 2026-04-22.** Use the Portainer UI editor for compose changes. See [mesh migration](../infrastructure/mesh.md#critical-never-docker-compose-up-portainer-managed-stacks) for the full warning.
### homelab-mcp (`/opt/homelab-mcp/`)
FastMCP server (Python venv at `/opt/homelab-mcp/.venv`). Reads from
`/opt/homelab-context/` (this LXC is itself an enrolled
@@ -54,7 +60,7 @@ is `dtoro/Homelab-Docs/mcp/server.py`; service unit
disabled at the FastMCP layer because mesh+LAN gating is the actual
trust boundary.
- Endpoint: `https://mcp.hubris.network/mcp` (Caddy → `:9810`). StreamableHTTP transport (POST `/mcp`).
- Endpoint: `https://mcp.hubris.network/sse` (Caddy → `:9810`).
- 14 tools registered: `get_host`, `list_services`, `find_service`,
`get_topology`, `search_docs`, `get_page`, `get_changelog`, `whoami`,
`list_my_secrets` (context); `get_service_status`, `tail_log`,
@@ -100,23 +106,20 @@ Native OIDC via `[oauth.generic]` in `config/config.ini`. `host = https://auth.h
## Related
- [Gitea (104)](104-gitea.md) — uses the PlantUML server
- [Caddy (121)](121-caddy.md)
- [Authentik (124)](106-auth-outpost.md)
- [Authentik (124)](124-authentik.md)
- [DNS](../infrastructure/dns.md)
- [Auto-deploy](../infrastructure/auto-deploy.md)
- [Public ingress (Artifacto + blog)](../infrastructure/ingress.md)
## Changelog
### 2026-06-29 — Booklore migrated to Grimmory on LXC 130
Booklore stack removed from Portainer. MariaDB dump taken first, then restored into [grimmory (130)](130-grimmory.md)'s fresh MariaDB. `books.hubris.network` Caddy backend updated to `192.168.8.213:6060`. Authentik OIDC provider updated to Public client type (PKCE) for Grimmory compatibility.
### 2026-05-20 — homelab-mcp + secrets-issuance live
Two new services from the [homelab-context distribution plan](../infrastructure/homelab-context.md):
`homelab-mcp.service` on `:9810` (MCP read+management surface) and
`secrets-issuance.service` on `:9820` (per-client age-key provisioning).
Caddy fronts both with Let's Encrypt; new vhosts on
[caddy](121-caddy.md), split-horizon DNS entries on
[authentik (124)](106-auth-outpost.md). Gitea webhook ids 10 + 11 wire
[authentik (124)](124-authentik.md). Gitea webhook ids 10 + 11 wire
auto-deploy. LXC is itself an enrolled context client
(`/opt/homelab-context/`).

View File

@@ -11,7 +11,7 @@ Personal cloud / file collaboration. Source-of-truth for the photo libraries sur
- **Public hostname:** [`cloud.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md)
## Auth
Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin` via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider <name> --mapping-uid=nc_uid` and `--scope="openid profile email <app>-uid"`. See [Authentik](106-auth-outpost.md#per-app-username-override-pattern-authentik) for the full pattern.
Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin` via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider <name> --mapping-uid=nc_uid` and `--scope="openid profile email <app>-uid"`. See [Authentik](124-authentik.md#per-app-username-override-pattern-authentik) for the full pattern.
Redirect URI: `/index.php/apps/user_oidc/code` (NOT `/apps/...` — pretty URLs aren't on).
@@ -74,7 +74,7 @@ Apply with `systemctl restart mariadb` (not reload — `innodb_log_file_size` ne
## Related
- [mulita (120)](120-mule-images.md) — reads NC user trees + writes back via WebDAV
- [Authentik (124)](106-auth-outpost.md)
- [Authentik (124)](124-authentik.md)
- [Caddy (121)](121-caddy.md)
- [DNS](../infrastructure/dns.md)
- [Mesh migration (DNS overrides explained)](../infrastructure/mesh.md)

View File

@@ -1,11 +1,10 @@
# 118 — `elementsynapse`
Matrix homeserver (Synapse). Backs `@dtoro:avispero`.
Matrix homeserver (Synapse). Backs `@dtoro:avispero` and `@claudio:avispero`.
## At a glance
- **Hostname:** `elementsynapse`
- **IP:** `192.168.8.242`
- **Host:** **strong** (migrated from hubris 2026-07-05)
- **IP:** `192.168.8.239`
- **Privilege:** **unprivileged**
- **Resources:** 1 core / 2 GiB RAM / **16 GiB rootfs** (grown from 8 GiB on 2026-05-15 after disk-full incident)
- **Mounts:** none from `/mnt/library`
@@ -34,29 +33,16 @@ All five bridges run as plain `docker compose` stacks under `/root/mautrix-<name
- `/var/lib/matrix-synapse/media_store` is the dominant grower (~2 GiB at last check). If disk pressure returns, purge remote media via the Synapse admin API before resizing further.
## Known issues
- ~~Disk usage was 86.8% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
- ~~Disk usage was 86.8% at last claudio-monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
## Related
- ~~[claudio-bot (123)](archive/123-claudio-bot.md)~~ — decommissioned 2026-06-04, replaced by Hermes Agent
- [claudio-bot (123)](123-claudio-bot.md) — connects directly to `192.168.8.239:8008` (avoids hairpin-NAT TLS issue on the public URL)
- [Caddy](121-caddy.md)
- [DNS](../infrastructure/dns.md)
- [Monitoring](../infrastructure/monitoring.md)
## Changelog
### 2026-06-06 — DHCP drift fixed: internal `/etc/network/interfaces` was `dhcp` despite Proxmox static config
**Symptom:** Matrix was down. Caddy at `192.168.8.175` couldn't reach `192.168.8.239:8008` — the LXC was actually at `192.168.8.244` because the guest-side dhclient had overridden the PVE-assigned static IP.
**Root cause:** During the 2026-06-02 static-IP migration, `pct set 118 --net0 ... ip=192.168.8.239/24` was applied to the Proxmox config, but the internal `/etc/network/interfaces` still had `iface eth0 inet dhcp`. On every DHCP lease renewal, dhclient grabbed `.244` from Technitium's pool.
**Fix:**
- Replaced `iface eth0 inet dhcp` with `iface eth0 inet static` + `address 192.168.8.239/24` + `gateway 192.168.8.1`
- `ifdown eth0 && ifup eth0` applied the static IP
- Killed lingering dhclient process
- Verified: `curl http://192.168.8.239:8008` returns 302 from Caddy's LXC
**Prevention:** The `check-caddy-backends.sh` cron on hubris now runs every 10 minutes, which would have caught this drift within 10 minutes of occurrence.
### 2026-05-15 — phantom-notification cleanup for `@admin`
After the disk-full incident, the mobile (Element X) badge showed ~125 unread but every room read clean in the UI. Root cause: stale rows in `event_push_actions` that were never reaped — Synapse's read-receipt-driven cleanup didn't catch up. Two contributors:
1. **8 of 12 affected rooms** had read receipts past the "unread" stream_ordering — pure stale state, likely from the disk-full window stalling rotation/cleanup.

View File

@@ -60,7 +60,7 @@ For pushes from inside the LXC, gitea creds at `/etc/mule-deploy/git-credentials
## Related
- [Nextcloud (114)](114-nextcloud.md) — source of truth for photo libraries
- [Authentik (124)](106-auth-outpost.md)
- [Authentik (124)](124-authentik.md)
- [Caddy (121)](121-caddy.md)
- [DNS](../infrastructure/dns.md)
- [Auto-deploy](../infrastructure/auto-deploy.md)

View File

@@ -11,16 +11,16 @@ The reverse proxy. Terminates TLS for every `*.hubris.network` hostname on the L
- **Config:** `/etc/caddy/Caddyfile` is a [git checkout of `dtoro/caddy-conf`](#auto-deploy)
- **Cert source:** Let's Encrypt **DNS-01** via IONOS API (`IONOS_AUTH_API_TOKEN`).
## Sites currently served (live as of 2026-07-06)
## Sites currently served (live as of 2026-04-28)
- `artifacto.hubris.network` → [apps (105)](105-apps.md) `:3100`
- `auth.hubris.network` → [authentik (124)](124-authentik.md) `:9000`
- `blog.hubris.network` → [apps (105)](105-apps.md) `:8080`
- `books.hubris.network` → [grimmory (130)](130-grimmory.md) `:6060`
- `books.hubris.network` → [apps (105)](105-apps.md) `:6060`
- `cloud.hubris.network` → [nextcloud (114)](114-nextcloud.md) `:443`
- `docker.hubris.network` → [apps (105)](105-apps.md) `:9443`
- `git.hubris.network` → [gitea (104)](104-gitea.md) `:3000` (+ `handle_path /_plantuml/*` → apps `:8079`)
- `home.hubris.network` → [haos VM (108)](../vms/108-haos.md) `192.168.8.101:8123`
- `house.hubris.network` → [house (129)](129-house.md) `:3000`
- `jellyseerr.hubris.network` → [arriman (122)](122-arriman.md) `:5056`
- `matrix.hubris.network` → [elementsynapse (118)](118-elementsynapse.md) `:8008`
- `media.hubris.network` → [jellyfin (101)](101-jellyfin.md) `:8096`
@@ -28,18 +28,15 @@ The reverse proxy. Terminates TLS for every `*.hubris.network` hostname on the L
- `photos.hubris.network` → [mule-images (120)](120-mule-images.md) `:3000`
- `proxmox.hubris.network` → [hubris host](../hosts/hubris.md) `:8006`
- `qbit.hubris.network` → [arriman (122)](122-arriman.md) `:8080`
- `roms.hubris.network` → [romm (134)](134-romm.md) `:80`
- `sab.hubris.network` → [arriman (122)](122-arriman.md) `:8082` (Authentik forward-auth)
- `teddy.hubris.network` → LXC 131 `192.168.8.150:8443`
- `trmnl.hubris.network` → [trmnl (128)](128-trmnl.md) `:9851`
- `sab.hubris.network` → [arriman (122)](122-arriman.md) `:8081`
> **Reminder:** Caddy alone isn't enough to make a new subdomain reachable on the LAN. Each one needs an entry in [DNS split-horizon](../infrastructure/dns.md) too.
## Snippet: `(authentik)` forward-auth
A snippet at the top of the Caddyfile (used as `import authentik` in any site block) wires forward-auth to the embedded Authentik outpost. It points at `http://192.168.8.180:9000` directly (NOT `https://auth.hubris.network`) to avoid Caddy-to-self round-tripping that strips `X-Forwarded-Host`. The forward-auth block must explicitly set `header_up X-Forwarded-Host {host}`. See [Authentik](106-auth-outpost.md#forward-auth-domain-level-setup).
A snippet at the top of the Caddyfile (used as `import authentik` in any site block) wires forward-auth to the embedded Authentik outpost. It points at `http://192.168.8.180:9000` directly (NOT `https://auth.hubris.network`) to avoid Caddy-to-self round-tripping that strips `X-Forwarded-Host`. The forward-auth block must explicitly set `header_up X-Forwarded-Host {host}`. See [Authentik](124-authentik.md#forward-auth-domain-level-setup).
For apps with mobile clients, `/api/*` (or equivalent) bypasses forward-auth — see the per-app gotchas in [Authentik](106-auth-outpost.md).
For apps with mobile clients, `/api/*` (or equivalent) bypasses forward-auth — see the per-app gotchas in [Authentik](124-authentik.md).
## Caddy environment
@@ -60,7 +57,7 @@ Gitea webhook id 2 on `dtoro/caddy-conf`. Receiver, deploy script, install scrip
## Related
- [DNS split-horizon](../infrastructure/dns.md) — must add entry for every new subdomain
- [Authentik (124)](106-auth-outpost.md) — forward-auth + IdP
- [Authentik (124)](124-authentik.md) — forward-auth + IdP
- [Auto-deploy](../infrastructure/auto-deploy.md)
- [Public ingress (VPS traefik)](../infrastructure/ingress.md) — mirrors Caddy's certs to the VPS for public exposure
- [Gitea (104)](104-gitea.md) — webhook source
@@ -68,28 +65,6 @@ Gitea webhook id 2 on `dtoro/caddy-conf`. Receiver, deploy script, install scrip
## Changelog
### 2026-06-13 — sab.hubris.network gated with Authentik forward-auth; port fixed :8081→:8082
`sab.hubris.network` now uses `import authentik` inside a `handle` block. SABnzbd own auth disabled, `local_ranges` set for transparent proxy. Port bumped from `:8081` to `:8082` (fix from 2026-06-04) now documented.
### 2026-06-06 — Caddyfile truncated to 43 lines; restored from origin/master + safeguards added
**Symptom:** All `*.hubris.network` hosts except `photos` and `auth` (VPS-hosted) returned `tlsv1 alert internal error` or timeout. Only 3 site blocks (`photos`, `prism`, `photos2`) remained in the Caddyfile.
**Root cause:** The Caddyfile was manually edited directly on LXC 121 (not via the `dtoro/caddy-conf` git repo), overwriting 260 lines / 30+ site blocks with 43 lines of photo-only config.
**Fix:**
- Restored Caddyfile from `origin/master` (`git checkout --force origin/master -- Caddyfile`)
- `systemctl reload caddy`
**Permanent safeguards added to `/etc/caddy/scripts/deploy.sh`:**
- **Site-count guard:** refuses to reload if fewer than 20 `*.hubris.network` blocks detected
- **Dirty-tree auto-stash:** stashes local changes before `git pull --ff-only` so the webhook doesn't fail on local edits
- **Auto-backup:** saves `Caddyfile.bak.<timestamp>` before any modifications, keeps last 5
Also: [elementsynapse LXC 118](118-elementsynapse.md) found to have DHCP-overridden static IP (actual `.244` vs config `.239`) during incident investigation — fixed.
### 2026-06-02 — caddy.service unit missing; recreated
After the Slate AX → SODOLA network migration, Caddy was not listening (ports 80/443 dead). Root cause: the custom hubris1 Debian package (`caddy_1:2.11.3-hubris1_amd64`) does not ship a systemd service unit file. The unit had previously existed but was lost (likely on a package reinstall). Recreated at `/lib/systemd/system/caddy.service` with standard Caddy service config + `EnvironmentFile=/etc/caddy/caddy.env` (already present in `caddy.service.d/override.conf`). **Risk:** the unit will be lost again if the package is reinstalled without the file being tracked. Fix: add the service unit to the `caddy-conf` repo or rebuild the hubris1 package to include it.
### 2026-04-28 — wiki entry created
Initial documentation. 16 active sites at this date.

View File

@@ -4,11 +4,10 @@ Docker host running the \*arr stack via [`ezarr`](https://github.com/ezarr/ezarr
## At a glance
- **Hostname:** `arriman`
- **IP:** `192.168.8.245`
- **Host:** **strong** (migrated from hubris 2026-07-05)
- **IP:** `192.168.8.132`
- **Privilege:** privileged
- **Resources:** 4 cores / 8 GiB RAM / 24 GiB rootfs
- **Mounts:** `/mnt/media_local``/mnt/library`
- **Mounts:** `/mnt/library``/mnt/library`
- **Public hostnames:** `jellyseerr` / `qbit` / `sab` (see below)
## Compose
@@ -23,27 +22,19 @@ Docker host running the \*arr stack via [`ezarr`](https://github.com/ezarr/ezarr
## Service / port map
All services route through gluetun's network namespace. Ports are exposed via
the gluetun container:
| Service | Host:Container | Public hostname |
| ------------- | -------------- | ------------------------------------ |
| gluetun (VPN) | — | — |
| sonarr | `8989:8989` | direct only (via gluetun) |
| radarr | `7878:7878` | direct only (via gluetun) |
| lidarr | `8686:8686` | direct only (via gluetun) |
| prowlarr | `9696:9696` | direct only (via gluetun) |
| bazarr | `6767:6767` | direct only (via gluetun) |
| sonarr | `8989:8989` | direct only |
| radarr | `7878:7878` | direct only |
| lidarr | `8686:8686` | direct only |
| prowlarr | `9696:9696` | direct only |
| bazarr | `6767:6767` | direct only |
| jellyseerr | `5056:5055` | [`jellyseerr.hubris.network`](../infrastructure/dns.md) |
| qbittorrent | `8080:8080` | [`qbit.hubris.network`](../infrastructure/dns.md) |
| sabnzbd | `8082:8082` HTTP, `9090:9090` HTTPS | [`sab.hubris.network`](../infrastructure/dns.md) |
| sabnzbd | `8081:8080` | [`sab.hubris.network`](../infrastructure/dns.md) |
| flaresolverr | `8191:8191` | internal only |
| homarr | `7575:7575` | internal only |
Internal *arr ↔ *arr / *arr ↔ qBit/SAB/flaresolverr comms run on `localhost:<port>`
(services share gluetun's shared network namespace). External services reach them
via `gluetun:<port>` (e.g. Sonarr → qBittorrent at `localhost:8080` or
`gluetun:8080`).
Internal *arr ↔ *arr / *arr ↔ qBit/SAB/flaresolverr comms run on `ezarr_default` using docker service names.
## Categories (qBit + SAB + *arr)
@@ -57,30 +48,18 @@ via `gluetun:<port>` (e.g. Sonarr → qBittorrent at `localhost:8080` or
Path mapping: host `/mnt/library/<cat>` ↔ container `/data/media/<cat>`. Downloads: host `/mnt/library/downloads/<torrents|usenet>/<cat>` ↔ container `/data/torrents/<cat>` and `/data/usenet/<cat>`.
## Auth (reverse-proxy + Authentik forward-auth)
## Auth (qBit reverse-proxy + Authentik forward-auth)
### qBit
Auto-login behind forward-auth via IP whitelist. `qBittorrent.conf` lines:
- `WebUI\\AuthSubnetWhitelist=172.18.0.0/16, 172.17.0.0/16, 192.168.8.175/32`
- `WebUI\\ReverseProxySupportEnabled=true`
- `WebUI\\TrustedReverseProxiesList=192.168.8.175, 172.18.0.0/16`
qBit auto-login behind forward-auth via IP whitelist. `qBittorrent.conf` lines:
- `WebUI\AuthSubnetWhitelist=172.18.0.0/16, 172.17.0.0/16, 192.168.8.175/32`
- `WebUI\ReverseProxySupportEnabled=true`
- `WebUI\TrustedReverseProxiesList=192.168.8.175, 172.18.0.0/16`
> **Stop the container before editing `qBittorrent.conf`.** qBit writes its in-memory config on graceful shutdown and clobbers any live edits. Recipe: `docker stop qbittorrent && sed -i ... && docker start qbittorrent`.
Mobile/desktop clients keep working via `/api/v2/*` path bypass on Caddy.
### SABnzbd
Gated with Authentik forward-auth (applied 2026-06-13). Caddy `sab.hubris.network` block uses `import authentik` inside a `handle` block. SABnzbd's own web auth is disabled:
- `html_login = 0` → no HTML login form
- `username` / `password` cleared → CherryPy Basic Auth not activated
- `local_ranges = 172.18.0.0/16, 192.168.8.0/24, 127.0.0.0/8` → proxied requests from Caddy (192.168.8.x) and Docker-proxy (172.18.x) pass without auth
**API key** (`67ef5a45e4e04157994e977005a33878`) still works for internal service-to-service calls (Sonarr/Radarr/Lidarr via Docker internal networking — they talk to SAB at `localhost:8082`, not through Caddy).
`host_whitelist`: `sabnzbd, localhost, 127.0.0.1, 192.168.8.132, sab.hubris.network` — extend before accessing SAB from a new host.
SABnzbd `host_whitelist`: `sabnzbd, localhost, 127.0.0.1, 192.168.8.132, sab.hubris.network` — extend before accessing SAB from a new host.
## Credentials
@@ -115,39 +94,13 @@ Member of [media GID 10000](../infrastructure/media-permissions.md). The LXC has
## Related
- [Caddy (121)](121-caddy.md)
- [Authentik (124)](106-auth-outpost.md) — forward-auth wiring + per-app `/api/*` bypass
- [Authentik (124)](124-authentik.md) — forward-auth wiring + per-app `/api/*` bypass
- [DNS](../infrastructure/dns.md)
- [Media permissions](../infrastructure/media-permissions.md)
- [Hubris host](../hosts/hubris.md)
## Changelog
### 2026-06-13 — SABnzbd gated with Authentik forward-auth
SABnzbd now uses Authentik forward-auth (same `import authentik` Caddy pattern as qBit). SABnzbd's own web auth disabled: `html_login=0`, credentials cleared, `local_ranges` extended to cover Docker bridge + homelab LAN. API key still works for internal *arr service calls. See [Auth section](#auth-reverse-proxy--authentik-forward-auth) above.
### 2026-06-04 — all arr services moved behind gluetun VPN; SAB port conflict fixed
- All services (sonarr, radarr, lidarr, bazarr, prowlarr, jellyseerr, homarr,
flaresolverr) now use `network_mode: service:gluetun` — whole stack routes
through the VPN
- Port mappings moved from individual services to gluetun container
- **Fixed SABnzbd port conflict**: was crashing in a restart loop because
qBittorrent held port 8080 inside the shared gluetun namespace. Changed
SAB internal port to 8082 (config at `/config/sabnzbd-config/sabnzbd.ini`)
- Caddy `sab.hubris.network` updated to point to `:8082`
- Jellyseerr's `extra_hosts` (auth.hubris.network) moved to gluetun since
`extra_hosts` conflicts with `network_mode`
### 2026-06-02 — ProtonVPN added (gluetun); LXC IP set static
- Added `gluetun` container to compose as a WireGuard VPN sidecar (ProtonVPN, server AL#57, located in Tirana, Albania)
- **qbittorrent** and **sabnzbd** now use `network_mode: service:gluetun` — all traffic routes through the VPN
- Ports 8080 (qBit WebUI), 6881 tcp/udp (qBit BT), 8081 (SAB WebUI) exposed through gluetun
- gluetun config at `gluetun-config/wireguard/wg0.conf` (read-only mount)
- Healthcheck on gluetun; qBit/SAB wait for `service_healthy` before starting
- LXC IP changed from DHCP to static (`192.168.8.132`) via `pct set` + `/etc/network/interfaces`
- **After first start:** Sonarr/Radarr/Lidarr download client host needs updating from `qbittorrent``gluetun` (SAB similarly `sabnzbd``gluetun`)
- **Also fixed:** 7 other DHCP LXCs (101 jellyfin, 103 paperless, 104 gitea, 105 apps, 114 nextcloud, 118 elementsynapse, 120 mule-images, 121 caddy) set to static IPs to prevent floating on reboot. See infrastructure/dns.md.
### 2026-04-28 — wiki entry created
Initial documentation.

View File

@@ -1,9 +1,4 @@
# 123 — `claudio-bot` (DEPRECATED — destroyed 2026-06-04)
> **This LXC was destroyed on 2026-06-04.** Replaced by Hermes Agent on mac-mini.
> Monitoring migrated to `homelab-hardware-health` skill + 15-min Hermes cronjob.
> Repos `dtoro/claudio-bot` and `dtoro/claudio-monitor` archived (read-only) on Gitea.
> See [deprecation plan](../../../../plans/done/2026-06-04_130000-deprecate-claudio-bot.md) for full details.
# 123 — `claudio-bot`
Matrix-resident control plane. Bot account `@claudio:avispero` joined to a private room; accepts slash commands and natural language; relays infra notifications.
@@ -17,7 +12,7 @@ Matrix-resident control plane. Bot account `@claudio:avispero` joined to a priva
## Stack
Repo `dtoro/claudio-bot`, checkout at `/opt/claudio-bot`, systemd unit `claudio-bot.service`. Connects to Matrix at `http://192.168.8.239:8008` (direct LAN to [synapse (118)](../118-elementsynapse.md), avoids hairpin-NAT TLS issue on `matrix.hubris.network`).
Repo `dtoro/claudio-bot`, checkout at `/opt/claudio-bot`, systemd unit `claudio-bot.service`. Connects to Matrix at `http://192.168.8.239:8008` (direct LAN to [synapse (118)](118-elementsynapse.md), avoids hairpin-NAT TLS issue on `matrix.hubris.network`).
Room: `!dEUVJArVKPorxHHJZK:avispero` (invite-only, `@dtoro:avispero` allowed).
@@ -43,8 +38,8 @@ Currently set to `lmstudio` → `google/gemma-4-e4b` on the Mac mini at `192.168
## IPC
`http://192.168.8.230:9090/{notify,propose,status}`. Header `X-Bot-Token` must match `/etc/claudio-bot/ipc.token`. Used by:
- [claudio-monitor on hubris](../../infrastructure/monitoring.md) for edge-triggered alerts (token in `/etc/claudio-monitor/bot.token`)
- The (currently disabled) [restic backup wrapper](../../infrastructure/backups.md) (token in `/etc/restic/bot.token`)
- [claudio-monitor on hubris](../infrastructure/monitoring.md) for edge-triggered alerts (token in `/etc/claudio-monitor/bot.token`)
- The (currently disabled) [restic backup wrapper](../infrastructure/backups.md) (token in `/etc/restic/bot.token`)
> Token files at the source side **must hold the same value as `/etc/claudio-bot/ipc.token`** — rotate together.
@@ -61,22 +56,16 @@ Active plugins:
Push to `dtoro/claudio-bot` → gitea webhook → `http://192.168.8.230:9797/deploy` → pull + `pip install` + restart. Same shape as caddy-conf.
`app.ini` `ALLOWED_HOST_LIST` on [gitea](../104-gitea.md) includes `192.168.8.230`.
`app.ini` `ALLOWED_HOST_LIST` on [gitea](104-gitea.md) includes `192.168.8.230`.
## Related
- [elementsynapse (118)](../118-elementsynapse.md)
- [Monitoring (claudio-monitor)](../../infrastructure/monitoring.md)
- [Backups (disabled)](../../infrastructure/backups.md)
- [Auto-deploy](../../infrastructure/auto-deploy.md)
- [elementsynapse (118)](118-elementsynapse.md)
- [Monitoring (claudio-monitor)](../infrastructure/monitoring.md)
- [Backups (disabled)](../infrastructure/backups.md)
- [Auto-deploy](../infrastructure/auto-deploy.md)
## Changelog
### 2026-06-04 — LXC destroyed; replaced by Hermes Agent
LXC 123 destroyed via `pct destroy 123 --purge`. Bot service stopped, systemd
units disabled. `dtoro/claudio-bot` and `dtoro/claudio-monitor` archived on
Gitea. Monitoring replaced by Hermes `homelab-health-watchdog` cron job.
`@claudio:avispero` Matrix account decommissioned.
### 2026-04-28 — wiki entry created
Initial documentation.
@@ -84,7 +73,7 @@ Initial documentation.
`backend: lmstudio``google/gemma-4-e4b` on the Mac mini. Anthropic key still present so the swap is reversible by flipping the config field.
### 2026-04-21 — `monitor` plugin added
Receives events from [claudio-monitor](../../infrastructure/monitoring.md). Slash commands + tools registered. See `dtoro/claudio-bot` commit `e56da25`.
Receives events from [claudio-monitor](../infrastructure/monitoring.md). Slash commands + tools registered. See `dtoro/claudio-bot` commit `e56da25`.
### 2026-04-20 — claudio-bot deployed
LXC 123 provisioned. Repo, systemd unit, Matrix wiring, `system` + `backup` plugins, IPC server.

186
containers/124-authentik.md Normal file
View File

@@ -0,0 +1,186 @@
# 124 — `authentik`
Central Identity Provider for the lab. Also runs the [split-horizon dnsmasq](../infrastructure/dns.md) — ergo "the SSO and DNS box".
## At a glance
- **Hostname:** `authentik`
- **IP:** `192.168.8.180` (statically configured — the only LXC with a static IP)
- **Privilege:** privileged
- **Resources:** 2 cores / 4 GiB RAM / 20 GiB rootfs
- **Mounts:** none from `/mnt/library`
- **Public hostname:** [`auth.hubris.network`](../infrastructure/dns.md) → [caddy (121)](121-caddy.md) → `:9000`
- **Container DNS (in `/etc/pve/lxc/124.conf`):** `192.168.8.1 1.1.1.1` (router DNS plus a fallback added 2026-04-21 because router DNS flakes intermittently — Authentik is the resolver itself for the *rest* of the LAN, but its own LXC uses upstream).
## Authentik stack (`/opt/authentik/`)
Upstream `docker-compose.yml` + `.env`. Services: `postgresql` (16-alpine), `server`, `worker`. Authentik 2026.x dropped the Redis dependency.
- `.env` mode 600, **untracked**, holds `AUTHENTIK_SECRET_KEY` and `PG_PASS`.
- `AUTHENTIK_TAG=2026.2.2` — pinned. Don't let it drift to `:latest`. Telemetry / update-check / error-reporting disabled.
- Ports: 9000 (http), 9443 (https) on the LXC.
- Embedded outpost lives at `/outpost.goauthentik.io/*` on the Authentik host — the forward-auth endpoint Caddy points at.
- Stack is **not** git-tracked yet. If/when wiring auto-deploy: mirror the `mule-image` pattern (webhook receiver outside the app repo at `/opt/authentik-deploy/`). Repo `dtoro/authentik-conf` is reserved but not created.
## Forward-auth pattern (every gated app)
- **One Proxy Provider per app.** Authentik enforces a UNIQUE constraint `application.provider_id`, so one Provider = one Application. "Domain-level" only means they share the cookie domain. Each provider in "Forward auth (domain level)" mode, External host `https://auth.hubris.network`, Cookie domain `hubris.network`. First one was `hubris-forward-auth` (Paperless).
- **Authentication flow:** MUST be `default-authentication-flow` (NOT `default-source-authentication` — that's for IdP federation; gives `FlowNonApplicableException` + 404 on the authorize endpoint).
- **Authorization flow:** `default-provider-authorization-implicit-consent` (or explicit).
- **Application Launch URL** MUST be the full public URL `https://<sub>.hubris.network/` — outpost matches incoming `X-Forwarded-Host` against it.
- Each Application MUST have at least one **policy/group/user binding** — zero bindings = outpost returns 404 on access.
- **Restart Authentik after binding new apps to the outpost:**
```
pct exec 124 -- docker compose -f /opt/authentik/docker-compose.yml restart server worker
```
Caddy snippet `(authentik)` lives at the top of `/etc/caddy/Caddyfile`. Points at `http://192.168.8.180:9000` directly (NOT `https://auth.hubris.network`) to avoid hairpin TLS round-trip stripping `X-Forwarded-Host`. Must explicitly set `header_up X-Forwarded-Host {host}` in the forward-auth block. Used by gated sites with `import authentik`.
### Per-app username override pattern (Authentik)
Used when the app's local user ID doesn't match the user's Authentik username (e.g., Nextcloud's `admin` ≠ Authentik's `dtoro`).
1. On the Authentik user: add attribute `<app>_uid: <target_local_username>` (YAML, Directory → Users → Edit → Attributes).
2. Customization → Property Mappings → Create → **Scope Mapping** (not SAML):
- Name: `<app>-uid-override`, Scope name: `<app>-uid`, Expression:
```python
return {"nc_uid": user.attributes.get("<app>_uid", user.username)}
```
- **Use a custom claim key** (e.g. `nc_uid`), not `preferred_username` — the default `profile` scope mapping emits `preferred_username` and will overwrite yours depending on ordering.
3. Attach the new scope to the provider (Providers → app → Scopes).
4. On the app side, point its OIDC UID-mapping setting at the custom claim.
For Nextcloud:
```
occ user_oidc:provider <name> --mapping-uid=nc_uid
occ user_oidc:provider <name> --scope="openid profile email <app>-uid"
```
### Bypass forward-auth for API paths (mobile apps)
If the app has its own token-based API auth and a mobile client, API paths must bypass forward-auth — mobile apps can't follow the browser login redirect. Pattern in the Caddyfile site block:
```
paperless.hubris.network {
tls { dns ionos {env.IONOS_AUTH_API_TOKEN} }
@api path /api/*
handle @api {
reverse_proxy 192.168.8.130:8000
}
handle {
import authentik
reverse_proxy 192.168.8.130:8000
}
}
```
API paths to bypass per app:
- [Paperless](103-paperless.md): `/api/*` (Bearer)
- [Sonarr / Radarr / Lidarr / etc.](122-arriman.md): `/api/*` (X-Api-Key)
- [qBittorrent](122-arriman.md): `/api/*` (session cookie from `/api/v2/auth/login`)
- [SABnzbd](122-arriman.md): `/api?*` (apikey query param) — match `/api*` for query-string APIs
- Homarr: no mobile client
- [Portainer](105-apps.md): mobile uses same session auth as web; no bypass typically needed
### Backend trust of Authentik headers (skip the app's own login after SSO)
- [**Paperless**](103-paperless.md): `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Restart `paperless-webserver paperless-task-queue paperless-scheduler paperless-consumer`. Django auto-creates matching users on first SSO login; promote to superuser via existing admin UI.
- Apps without header-auth support: users log in twice (SSO + app login). Acceptable but degraded UX.
## Per-app integration map
| App | Type | Notes |
| ----------------------------------------- | ---------------- | ----- |
| [Paperless (103)](103-paperless.md) | Forward-auth + REMOTE_USER | `/api/*` bypass |
| [Nextcloud (114)](114-nextcloud.md) | Native OIDC | `nc_uid` override; local dnsmasq required (Guzzle bypasses `/etc/hosts`) |
| [mulita (120)](120-mule-images.md) | Native OIDC | `extra_hosts` override in compose |
| [Booklore (105)](105-apps.md) | Native OIDC | Redirect URI `/oauth2-callback`; `extra_hosts` |
| [Portainer (105)](105-apps.md) | Native OAuth2 | `portainer_uid` custom claim; `--trusted-origins` flag |
| [WriteFreely (105)](105-apps.md) | Native OIDC | `[oauth.generic]` block; `extra_hosts` |
| [qBittorrent (122)](122-arriman.md) | Forward-auth via IP whitelist | Reverse-proxy support enabled in qBit |
| [Artifacto (105)](105-apps.md) | Forward-auth + gateway-secret auto-login | Public `/p/*` paths bypass |
| [Home Assistant VM (108)](../vms/108-haos.md) | HACS `christiaangoossens/hass-oidc-auth` | `automatic_user_linking: true`, `default_redirect: true`. Supervisor DNS via `ha dns options`. |
## Netbird IdP integration — LANDED 2026-05-21
The combined `netbirdio/netbird-server` image was replaced with the canonical vanilla stack (`mgmt + signal + relay + dashboard` 0.71.3) on the VPS so external OIDC actually works. Authentik is now the netbird dashboard's IdP. Full migration context in [mesh.md changelog](../infrastructure/mesh.md#changelog).
**Active provider & app:**
- Provider `NetBird` (OAuth2/OpenID), **Client type: `Public`** (PKCE-only — `Confidential` would break the dashboard SPA's token exchange).
- Client ID: `netbird-dashboard`. Client Secret is in `/opt/management.json` `PKCEAuthorizationFlow.ProviderConfig.ClientSecret` on the VPS (TODO: sops-encrypt as `secrets/netbird-authentik-oidc.yaml`).
- Application `NetBird`, slug `netbird`, launch URL `https://netbird.hubris.network/`.
- Redirect URIs: `https://netbird.hubris.network/peers`, `/nb-auth`, `/nb-silent-auth`, plus `https://netbird.hubris.network/` for post-logout.
- Scopes enabled on the provider: `openid`, `profile`, `email`.
- Discovery URL: `https://auth.hubris.network/application/o/netbird/.well-known/openid-configuration` — netbird mgmt fetches this on startup; logs `loaded OIDC configuration from the provided IDP configuration endpoint`.
**Login flow:** netbird dashboard PKCE → Authentik authorize → redirect back to `/nb-auth` → JS token exchange at Authentik's `/token` endpoint → mgmt validates the bearer against Authentik's JWKS.
**First-time owner promotion gotcha** (write-down for future operators):
When a new Authentik user logs in for the first time against an account that already has peers, netbird mgmt adds them as `role=user, blocked=1, pending_approval=1`. The OLD account-owner (the one in `store.db` from before the IdP swap) can't be reached anymore, so there's no admin to approve. Recovery is a direct sqlite update on `mgmt_data`:
```
docker stop netbird-mgmt
sqlite3 /var/lib/docker/volumes/opt_mgmt_data/_data/store.db \
"UPDATE users SET role='owner', blocked=0, pending_approval=0 WHERE id='<new authentik sub>';"
docker start netbird-mgmt
```
The Authentik sub-claim is the value of the `id` column on the newly-created user row (look for `role=user, blocked=1, pending_approval=1`).
### Device Code grant — configured (2026-05-21)
`netbird up` (interactive, without `--setup-key`) works against Authentik. The recipe:
1. **Flow** `default-device-code-flow` (designation: `Stage Configuration`) with 4 stage bindings in order:
- 10: `default-authentication-identification` (username/email lookup)
- 20: `default-authentication-password` (password validation)
- 30: `default-authentication-login` (attach authenticated user to session)
- 40: `default-provider-authorization-explicit-consent`'s Consent Stage (`default-provider-authorization-consent`) — the "Authorize NetBird?" approval
2. **Brand** (System → Brands → edit the brand serving `auth.hubris.network`): set **Device code flow** field to `default-device-code-flow`.
3. No provider-side change is required — Authentik 2026.x routes `/device` via the brand's device-code flow, not via the OAuth2/OpenID provider's `Authorization flow`.
**Why this matters for the lab**: Authentik 2026.x doesn't ship a default device-code flow. Without this configuration, the URL `https://auth.hubris.network/device` renders blank (the `/device` endpoint is unrouted), so `netbird up` device-codes expire without consent → only `--setup-key` works for onboarding. The above unblocks interactive onboarding.
**Verifying** from a browser tab: visit `https://auth.hubris.network/device`. You should see a form with one **Code** input + Continue button. Then `netbird up` (no setup-key) end-to-end:
- CLI prints `verification_uri_complete: https://auth.hubris.network/device?code=...`
- Open URL → identification (skipped if logged in) → password (re-auth check) → consent ("Authorize NetBird?") → Continue
- CLI completes registration with `Connected`
### Self-service onboarding (not yet — future-session)
`auth.hubris.network` is only reachable from inside the netbird mesh (split-horizon DNS). A brand-new client that isn't on the mesh yet can't OIDC-login → setup-key is the only path. To enable self-service onboarding via Authentik from the public internet:
- Add a Traefik route on the VPS for `auth.hubris.network` that forwards via the netbird-routed `192.168.8.0/24` to LXC 124.
- DNS already points `auth.hubris.network → 82.165.190.79` (IONOS wildcard).
Tracked in homelab memory as a queued follow-up.
### Old pre-work to remove
The previous `Provider for Netbird` + `netbird` app from 2026-04-22 (client ID `xZwVTFCsxWdBM3uIGS15wAAcVvsJiTtWdxVCEela`, service account `netbird-service`) is now obsolete — replaced by `netbird-dashboard` above. Safe to delete from Authentik admin UI; nothing currently uses the old client ID. The service account + API token can also be removed unless we wire IdpManagerConfig in mgmt later (currently `ManagerType: none`).
## DNS responsibility
dnsmasq runs alongside Authentik on this LXC, listening on `192.168.8.180:53` + `127.0.0.1:53`, serving every `*.hubris.network` subdomain → `192.168.8.175`. **There is no wildcard** — every site needs an explicit `address=` entry. See [DNS split-horizon](../infrastructure/dns.md).
## Related
- [DNS split-horizon](../infrastructure/dns.md)
- [Caddy (121)](121-caddy.md)
- [Mesh migration](../infrastructure/mesh.md)
- Every gated app under [containers/index](index.md)
## Changelog
### 2026-05-21 — Netbird IdP swap landed (Phase 6 done)
VPS migrated from combined netbird-server to vanilla mgmt+signal+relay+dashboard 0.71.3 (see [mesh.md](../infrastructure/mesh.md)), enabling Authentik as the dashboard IdP via PKCE. New Provider/App = `netbird-dashboard`, replacing the deferred pre-work. Device Code Stage still missing — interactive `netbird up` fails consent; setup-keys are the workaround until that's added.
### 2026-04-28 — wiki entry created
Initial documentation.
### 2026-04-22 — Phase 6 (Netbird IdP swap) deferred
Combined netbird-server image couldn't take an external IdP. Pre-work in Authentik kept for later (now superseded by 2026-05-21 above). Netbird mgmt host instead joined its own mesh as a peer (`100.122.165.149`) for split-horizon DNS access.
### 2026-04-22 — Artifacto, mulita, WriteFreely, Portainer wired
Native OIDC for mulita / WriteFreely / Portainer; gateway-secret auto-login pattern for Artifacto.
### 2026-04-21 — deployed; Phases 15 complete
LXC 124 provisioned, stack at `/opt/authentik`, public URL via Caddy, Paperless + Booklore + Nextcloud + Home Assistant wired. dnsmasq for split-horizon DNS lives on the same LXC.

108
containers/126-plato.md Normal file
View File

@@ -0,0 +1,108 @@
# 126 — `plato`
Docker host for [Plato](https://git.hubris.network/dtoro/Plato) — a cross-linked notes workspace (SvelteKit SPA embedded into a Go HTTP server, SQLite-backed). LAN+mesh only, no public ingress.
## At a glance
- **Hostname:** `plato`
- **IP:** `192.168.8.190`
- **Privilege:** privileged
- **Resources:** 2 cores / 2 GiB RAM / 8 GiB rootfs / 1 GiB swap
- **Mounts:** host `/mnt/library/documents/plato` ↔ container `/opt/plato/data`
- **Public hostname:** [`plato.hubris.network`](../infrastructure/dns.md) → [caddy (121)](121-caddy.md) → `192.168.8.190:8080`
## Stack
Single-container deploy. The repo's `Dockerfile` is a three-stage build (Node → Go → distroless/static-debian12:nonroot, ~23 MiB final image). The container exposes `:8080` and writes its SQLite db to `/data`.
- **Checkout:** `/opt/plato/app` (clone of `http://192.168.8.121:3000/dtoro/Plato.git`, using the cached gitea PAT in `/root/.git-credentials` — same pattern as [caddy (121)](121-caddy.md)).
- **Data:** host `/mnt/library/documents/plato` (owned `65532:65532` to match the distroless nonroot UID) bind-mounted into the LXC at `/opt/plato/data`, then bound into the container at `/data` via a `docker-compose.override.yml`:
```yaml
services:
plato:
volumes: !override
- /opt/plato/data:/data
restart: unless-stopped
```
- **`.env`** at `/opt/plato/app/.env` (optional, untracked) — LLM provider keys (`OPENROUTER_API_KEY`, `ANTHROPIC_API_KEY`, etc.) and `PLANTUML_BASE_URL` override. Absent by default; LLM features stay greyed out, PlantUML defaults to the public service.
- **Run / update:** push to `dtoro/Plato` (auto-deploys, see below) or `cd /opt/plato/app && git pull && docker compose up -d --build` for a manual rebuild.
## Auto-deploy
Push to `dtoro/Plato` `main` triggers a rebuild — same Shape B pattern as [Artifacto / mule-image](../infrastructure/auto-deploy.md). Webhook receiver at `/opt/plato-deploy/`, systemd unit `plato-deploy-webhook.service`, port `9799`, gitea hook id 8.
- Receiver: `http://192.168.8.190:9799/deploy`, signed payload (HMAC-SHA256, secret in `/etc/plato-deploy/secret`).
- Logs: `journalctl -u plato-deploy-webhook -f`.
- Health: `curl http://127.0.0.1:9799/health`.
- Manual deploy: `/opt/plato-deploy/deploy.sh`.
- Gitea's `app.ini` `ALLOWED_HOST_LIST` was extended with `192.168.8.190` to allow this delivery.
## Fresh-DB bootstrap workaround
The `schema` constant in `backend/internal/views/store.go` (as of commit `e0542c0`) creates the `views` table without `project_id`, then immediately runs `CREATE UNIQUE INDEX … ON views(project_id, lower(title))`. On a fresh DB this fails (no such column) and Plato crash-loops with `open views store: SQL logic error: no such column: project_id`. `ensureProjectIDColumn()` adds the column on subsequent migrations, but the schema apply happens first.
Until the upstream fix lands, pre-seed the DB before first start:
```
docker compose stop
rm -f /opt/plato/data/plato.db
python3 - <<'PY'
import sqlite3
c = sqlite3.connect('/opt/plato/data/plato.db')
c.executescript("""
CREATE TABLE views (
id TEXT PRIMARY KEY,
type TEXT NOT NULL DEFAULT 'document',
title TEXT NOT NULL,
aliases TEXT NOT NULL DEFAULT '[]',
content TEXT NOT NULL DEFAULT '',
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL,
project_id TEXT NOT NULL DEFAULT ''
);
CREATE UNIQUE INDEX views_project_title_lower ON views(project_id, lower(title));
CREATE INDEX views_project_id ON views(project_id);
""")
c.commit()
PY
chown 65532:65532 /opt/plato/data/plato.db
docker compose up -d
```
Once the column exists, every subsequent boot's `IF NOT EXISTS` clauses no-op. Once Plato is fixed upstream (remove the `CREATE UNIQUE INDEX` line from the boot `schema` constant — `ensureTitleIndexPerProject()` already re-creates it after the migration), this preseed becomes unnecessary.
## Why privileged
Matches the docker-host convention used by [120 mule-images](120-mule-images.md) and [122 arriman](122-arriman.md). Distroless nonroot's UID `65532` on the host bind mount maps directly through; unprivileged would shift the UID by the idmap offset and the container couldn't write `/data` without extra plumbing.
## Caddy
```
plato.hubris.network {
tls {
dns ionos {env.IONOS_AUTH_API_TOKEN}
}
reverse_proxy 192.168.8.190:8080
}
```
No Authentik forward-auth — Plato has no auth model yet; access control is "be on the LAN or the mesh".
## DNS
dnsmasq entry on [124-authentik](124-authentik.md):
```
address=/plato.hubris.network/192.168.8.175
```
## Related
- [Caddy (121)](121-caddy.md)
- [DNS (split-horizon)](../infrastructure/dns.md)
- [Media permissions](../infrastructure/media-permissions.md)
## Changelog
### 2026-05-13 — auto-deploy wired
Shape B pipeline added (`/opt/plato-deploy/`, port `9799`, gitea hook id 8). `ALLOWED_HOST_LIST` in gitea `app.ini` extended with `192.168.8.190`. See [auto-deploy](../infrastructure/auto-deploy.md#plato).
### 2026-05-13 — container created, Plato deployed
LXC 126 stood up on Debian 12 standard, privileged, docker-ce installed. Plato cloned from `dtoro/Plato`, built and started. Caddy site and dnsmasq split-horizon entry added. Recycled the IP/ID slot freed earlier the same day by the [decommissioned Seafile experiment (LXC 125)](index.md#recently-destroyed-kept-for-archaeology). Hit the [fresh-DB bootstrap bug](#fresh-db-bootstrap-workaround) on first boot; worked around by pre-seeding the SQLite schema.

View File

@@ -1,7 +1,7 @@
# 127 — `mule-photos-new`
Side-by-side **PhotoPrism M0 test** of the `dtoro/mule-image` `new` branch
at `photos-new.hubris.network`. Production [LXC 120](../120-mule-images.md) keeps
at `photos-new.hubris.network`. Production [LXC 120](120-mule-images.md) keeps
running on the legacy stack at `photos.hubris.network` until M5 cutover.
## At a glance
@@ -11,7 +11,7 @@ running on the legacy stack at `photos.hubris.network` until M5 cutover.
- **Resources:** 6 cores / 8 GiB RAM / 40 GiB rootfs / 1 GiB swap
- **Features:** `nesting=1,fuse=1,keyctl=1`
- **Mounts:** *(none — see scratch copy below)*
- **Public hostname:** [`photos-new.hubris.network`](../../infrastructure/dns.md) → [caddy (121)](../121-caddy.md) → split (PhotoPrism `:2342`, sidecar `:8000`, Vite `:5173`)
- **Public hostname:** [`photos-new.hubris.network`](../infrastructure/dns.md) → [caddy (121)](121-caddy.md) → split (PhotoPrism `:2342`, sidecar `:8000`, Vite `:5173`)
## Stack (`/opt/mule-image`)
@@ -65,7 +65,7 @@ unprivileged LXCs can't see through.
## Auth — Authentik OIDC
PhotoPrism's "Sign in with OIDC" button delegates to [Authentik (124)](../106-auth-outpost.md).
PhotoPrism's "Sign in with OIDC" button delegates to [Authentik (124)](124-authentik.md).
- **Provider/Application slug:** `mule-photos-new`
- **Issuer:** `https://auth.hubris.network/application/o/mule-photos-new/`
@@ -111,7 +111,7 @@ Mirrors the LXC 120 pattern.
Push to the `new` branch on [git.hubris.network/dtoro/mule-image](http://git.hubris.network/dtoro/mule-image) → webhook fires → rebuild. The legacy LXC 120 watches `main` and is unaffected.
**Gitea gotcha:** the receiver IP must be in `[webhook] ALLOWED_HOST_LIST`
in `/etc/gitea/app.ini` on [LXC 104](../104-gitea.md). LXC 127's
in `/etc/gitea/app.ini` on [LXC 104](104-gitea.md). LXC 127's
`192.168.8.181` was missing on first bring-up; every push delivered
status 0 with the message `webhook can only call allowed HTTP servers`.
Adding the IP and `systemctl restart gitea` is enough — same list is
@@ -145,10 +145,10 @@ curl -sk --resolve photos-new.hubris.network:443:192.168.8.175 \
```
> **Decommissioned 2026-05-22.** The PhotoPrism + sidecar + SvelteKit stack
> validated here was promoted into production on [LXC 120](../120-mule-images.md)
> validated here was promoted into production on [LXC 120](120-mule-images.md)
> via the `Mulimage 2.0` merge (`dtoro/mule-image` commit `70dc1b6`). This
> page is retained for archaeology; everything below is historic. See the
> 2026-05-22 entry in [120-mule-images.md](../120-mule-images.md#changelog) for
> 2026-05-22 entry in [120-mule-images.md](120-mule-images.md#changelog) for
> the cutover detail.
## Changelog

50
containers/index.md Normal file
View File

@@ -0,0 +1,50 @@
# LXC containers — index
All containers live on [`hubris`](../hosts/hubris.md). Each row links to the per-container page.
| ID | Name | IP | Priv | Cores | RAM | Disk | Mounts | Public hostname | Status |
| --- | ---------------- | --------------- | ---- | ----- | ----- | ----- | --------------------- | ------------------------------------- | -------- |
| 101 | [jellyfin](101-jellyfin.md) | 192.168.8.206 | unpriv (idmap) | 2 | 4 GiB | 16 GiB | `/mnt/library` | `media.hubris.network` | running |
| 103 | [paperless](103-paperless.md) | 192.168.8.130 | priv | 2 | 3 GiB | 8 GiB | `/mnt/library` | `paperless.hubris.network` | running |
| 104 | [gitea](104-gitea.md) | 192.168.8.121 | priv | 1 | 1 GiB | 8 GiB | `/mnt/library` | `git.hubris.network` | running |
| 105 | [apps](105-apps.md) | 192.168.8.205 | priv | 2 | 4 GiB | 30 GiB | `/mnt/library` | `docker` / `books` / `artifacto` / `blog` | running |
| 114 | [nextcloud](114-nextcloud.md) | 192.168.8.224 | priv | 4 | 6 GiB | 25 GiB | `/mnt/library` | `cloud.hubris.network` | running |
| 118 | [elementsynapse](118-elementsynapse.md) | 192.168.8.239 | unpriv | 1 | 2 GiB | 8 GiB | — | `matrix.hubris.network` | running |
| 119 | [sophia](119-sophia.md) | 192.168.8.157 | priv | 2 | 1 GiB | 10 GiB | `/mnt/library` | — | running |
| 120 | [mule-images](120-mule-images.md) | 192.168.8.136 | priv | 6 | 12 GiB | 60 GiB | `/mnt/library` + `/dev/dri` (iGPU passthrough) | `photos.hubris.network` | running |
| 121 | [caddy](121-caddy.md) | 192.168.8.175 | unpriv | 1 | 512 MiB | 6 GiB | — | (terminates all `*.hubris.network`) | running |
| 122 | [arriman](122-arriman.md) | 192.168.8.132 | priv | 4 | 8 GiB | 24 GiB | `/mnt/library` | `jellyseerr` / `qbit` / `sab` | running |
| 123 | [claudio-bot](123-claudio-bot.md) | 192.168.8.230 | unpriv | 1 | 512 MiB | 8 GiB | — | — | running |
| 124 | [authentik](124-authentik.md) | 192.168.8.180 | priv | 2 | 4 GiB | 20 GiB | — | `auth.hubris.network` | running |
| 126 | [plato](126-plato.md) | 192.168.8.190 | priv | 2 | 2 GiB | 8 GiB | `/mnt/library/documents/plato` | `plato.hubris.network` | running |
## Recently destroyed (kept for archaeology)
| ID | Name | Destroyed | Reason |
| --- | ---------------- | --------------- | --------------------------------------------- |
| 127 | mule-photos-new | 2026-05-22 | PhotoPrism + sidecar + SvelteKit stack promoted to LXC 120 via Mulimage 2.0 merge (`70dc1b6`); M0 test LXC retired. Caddy + dnsmasq + gitea webhook + NC webhook listeners all cleaned up in the same cutover. |
| 100 | arr (yunohost) | ~2026-04-28 | Migrated to docker stack on [arriman](122-arriman.md); planned retention window expired |
| 106 | flaresolverr | ~2026-04-28 | Folded into the arriman docker compose |
| 116 | heaper | 2026-05-14 | Decommissioned by user; data subtree at `/mnt/library/heaper` (224 MiB) retained |
| 109 | syncthing | 2026-05-14 | Decommissioned by user; `/mnt/library/syncthing` was already empty |
| 125 | seafile | 2026-05-13 | Seafile Pro evaluation, user disliked the product; teardown also removed `files.hubris.network` from caddy + dnsmasq |
| 107 | marimo | between 2026-04-21 and 2026-04-28 | Decommissioned |
| 110 | photoprism | between 2026-04-21 and 2026-04-28 | Replaced by [mulita](120-mule-images.md) |
| 111 | karakeep | between 2026-04-21 and 2026-04-28 | Decommissioned |
| 112 | immich | between 2026-04-21 and 2026-04-28 | Replaced by [mulita](120-mule-images.md) |
| 115 | reticulum | between 2026-04-21 and 2026-04-28 | Decommissioned |
> Several `.conf.bak` files survive under `/etc/pve/lxc/` if you need to recover any of the configs.
## Conventions
- All net0 are `bridge=vmbr0`, `ip=dhcp` except [124 (authentik)](124-authentik.md) which is statically `192.168.8.180/24`. IPs are stable via the LAN router's DHCP reservations.
- `onboot=1` on every container — the host brings them up after `pve-guests.service`.
- Bind mounts are declared as `mp0: /mnt/library,mp=/mnt/library`. Containers that don't mount `/mnt/library` don't need it.
- Most containers are privileged. Unprivileged ones (`101`, `118`, `121`, `123`) require an idmap block in their conf to participate in the [media GID 10000](../infrastructure/media-permissions.md) standard.
## Related
- [Hubris host](../hosts/hubris.md)
- [Media permissions](../infrastructure/media-permissions.md)
- [Caddy](121-caddy.md) — terminates every public hostname
- [DNS](../infrastructure/dns.md) — split-horizon entries for each subdomain

View File

@@ -1,70 +0,0 @@
# Docker Compose for Oikos development
# Usage: docker compose up -d postgres (just the DB)
# make dev (full dev stack)
services:
postgres:
image: timescale/timescaledb:2.17.2-pg16
environment:
POSTGRES_DB: oikos
POSTGRES_USER: oikos
POSTGRES_PASSWORD: ${OIKOS_DB_PASSWORD:-oikos_dev}
ports:
- "5432:5432"
volumes:
- pg-data:/var/lib/postgresql/data
healthcheck:
test: ["CMD", "pg_isready", "-U", "oikos"]
interval: 5s
timeout: 5s
retries: 5
# One-shot: run migrations then exit
migrate:
build:
context: .
dockerfile: compose/oikos/Dockerfile
depends_on:
postgres:
condition: service_healthy
environment:
OIKOS_DATABASE_URL: postgres://oikos:${OIKOS_DB_PASSWORD:-oikos_dev}@postgres:5432/oikos?sslmode=disable
command: ["migrate"]
restart: "no"
# One-shot: ingest seeds then exit
seed:
build:
context: .
dockerfile: compose/oikos/Dockerfile
depends_on:
migrate:
condition: service_completed_successfully
environment:
OIKOS_DATABASE_URL: postgres://oikos:${OIKOS_DB_PASSWORD:-oikos_dev}@postgres:5432/oikos?sslmode=disable
OIKOS_SEEDS_DIR: /app/seeds
command: ["seed"]
restart: "no"
# API server (Phase 2)
api:
build:
context: .
dockerfile: compose/oikos/Dockerfile
profiles: ["dev", "full"]
depends_on:
seed:
condition: service_completed_successfully
environment:
OIKOS_DATABASE_URL: postgres://oikos:${OIKOS_DB_PASSWORD:-oikos_dev}@postgres:5432/oikos?sslmode=disable
OIKOS_API_LISTEN: ":8090"
OIKOS_ENV: dev
OIKOS_DEBUG: "true"
ports:
- "8090:8090"
command: ["api"]
stop_signal: SIGTERM
stop_grace_period: 30s
volumes:
pg-data:

View File

@@ -1,23 +0,0 @@
# ADR 0001 — Go with single-binary role packaging
Status: accepted (2026-07-07) · Plan: rev 3, R3-4
## Context
The OS has three long-running roles (api, scheduler+actuator+learning,
notifier) plus one-shot jobs (migrate, seed, export). Rev 2 planned three
binaries with three Dockerfiles.
## Decision
One Go binary `oikos` with role subcommands (`oikos api | scheduler |
notifier | all | migrate | seed | export`), one multi-stage Dockerfile, one
image tagged `oikos:<git-sha>`. Compose runs the image N times with
different commands (Loki/Temporal pattern). Go over Python for static
typing, small static binaries (CGO_ENABLED=0, distroless), and goroutines
for concurrent probes.
## Consequences
- One build, guaranteed version consistency across roles, trivial local dev
(`oikos all`), simpler rollback (retag one image).
- Full rewrite of ~4,400 Python lines (logic carries over per plan reuse map).
- All roles share a dependency set; image is slightly larger than per-role
minimal images — accepted.

View File

@@ -1,25 +0,0 @@
# ADR 0002 — PostgreSQL + TimescaleDB as the only datastore
Status: accepted (2026-07-07) · Plan: rev 3
## Context
The OS needs a graph (entities/relationships), operational tables
(signals/executions/approvals), a learning corpus, time-series metrics,
audit and event logs. Alternatives: dedicated graph DB (Neo4j), dedicated
TSDB (Prometheus/VictoriaMetrics), or one Postgres.
## Decision
One PostgreSQL 16 instance with the TimescaleDB extension
(timescale/timescaledb:2-pg16). Graph traversal via recursive CTEs
(cycle-safe blast_radius); time-series via hypertables + continuous
aggregates + retention policies; events via table + LISTEN/NOTIFY.
## Consequences
- One backup/restore/DR story, one connection pool, transactional
consistency between graph and operational writes (e.g. event emission in
the same transaction as state change).
- Postgres is the accepted SPOF — mitigated by daily pg_dump + WAL PITR +
off-host copies + monthly restore drills; streaming replication is the
future path if needed.
- Homelab graph scale (hundreds of nodes) is far below where a dedicated
graph DB pays for itself.

View File

@@ -1,23 +0,0 @@
# ADR 0003 — DB-native ontology with YAML seed manifests
Status: accepted (2026-07-07) · Plan: rev 3, R3-1
## Context
Rev 1 kept inventory/ontology/policy as YAML files parsed at runtime.
Agents need graph queries (blast radius), transactional mutations with
audit, and a future UI needs to edit the model without file round-trips.
## Decision
The DB is the runtime source of truth. entity_types form an is-a hierarchy
(parent_type, is_abstract); relationship endpoint constraints may name
abstract types and validation walks the hierarchy. YAML files under seeds/
bootstrap the DB (idempotent, content-hashed via seed_versions) and serve
DR; `GET /api/v1/export` regenerates them for version control (round-trip
byte-stable, tested in CI).
## Consequences
- Ontology changes are API calls (policy-gated), not redeploys.
- Seeds can drift from DB between exports — export is part of the routine
(commit after meaningful model edits).
- Abstract types let policy rules and relationships bind once at the right
altitude (e.g. `compute-entity provides service`).

View File

@@ -1,21 +0,0 @@
# ADR 0004 — Contract-first OpenAPI API
Status: accepted (2026-07-07) · Plan: rev 3, R3-2/R3-3
## Context
Future UIs, a CLI client, and an MCP surface must stay in sync with the
API. Code-first (Gin + generated docs) drifts.
## Decision
api/openapi.yaml (OpenAPI 3.1) is the source of truth. Server stubs via
oapi-codegen (strict server, chi router); clients generated for Go (CLI)
and TypeScript (future UI). Conventions: RFC 9457 problem+json errors,
{items, next_cursor} envelopes, cursor pagination, Idempotency-Key on
unsafe POSTs, ETag/If-Match optimistic concurrency, scopes
(operator/viewer/agent) annotated per operation. CI fails on spec/handler
drift. MCP tools wrap the same service layer.
## Consequences
- UI development needs only the running API (spec served at /openapi.yaml).
- Handler changes require spec changes first — deliberate friction.
- Breaking changes ship as /api/v2 side by side; v1 is additive-only.

View File

@@ -1,18 +0,0 @@
# ADR 0005 — UUIDv7 + slug entity identity
Status: accepted (2026-07-07) · Plan: rev 3, R3-5 (resolves audit D1)
## Context
Rev 2 used TEXT primary keys ('host:hubris') — renames break FKs, and
date-string signal IDs are race-prone.
## Decision
Primary keys are UUIDv7 (time-ordered, generated in Go). Every entity also
carries a unique human slug ('host:hubris'); (type, name) is unique too.
The API accepts UUID or slug everywhere; slugs may change (rename), UUIDs
never do.
## Consequences
- Renames are metadata updates; history and edges survive.
- UUIDv7's time-ordering keeps B-tree inserts append-mostly.
- Seeds and exports use slugs (human-diffable); ingest resolves to UUIDs.

View File

@@ -1,23 +0,0 @@
# ADR 0006 — Learning is proposal-only (no self-authorization)
Status: accepted (2026-07-07) · Plan: rev 3 (resolves audit S3/S4/SA2)
## Context
The learning loop (feedback → patterns → skills) informs the classifier
that decides auto-act vs escalate. If learning could expand its own
autonomy, poisoned feedback (flapping services, biased probes) could
unlock destructive auto-act.
## Decision
The learning engine cannot write to governance (policy/autonomy) tables —
enforced structurally: its DB role has no grants on them. Pattern
activation (validated → active) and any autonomy expansion require operator
approval. Confidence is the Wilson lower bound capped by evidence_count/5;
anomalous feedback bursts quarantine the pattern; no skill ever
auto-promotes an action into destructive autonomy (hard-coded). Lowering
autonomy (kill-switch) is always immediate, never gated.
## Consequences
- Cold start is slow by design — the agent escalates until trust is earned.
- The operator is the only path to more autonomy; the audit trail shows
every grant.

View File

@@ -1,27 +0,0 @@
# ADR 0007 — Threat model and trust zones
Status: accepted (2026-07-07) · Plan: rev 3, Security model section
## Context
The control plane can restart services and (eventually) mutate config
fleet-wide. Compromise of any one container must not equal compromise of
the fleet.
## Decision
Trust zones as Docker networks: net-front (Caddy→api only), net-data
(Postgres), net-ops (SSH egress, actuator only). Hermes holds no SSH keys;
the actuator uses a restricted key (command=/from= in authorized_keys)
until the /executions gateway fully brokers actions. Caddy is an explicit
trust root but the API independently validates OIDC JWTs — network origin
is defense-in-depth, never the auth (this enables the LAN break-glass API
binding; the Hermes gateway remains mesh-only). Policy changes are
dual-controlled with before/after hash auditing and a startup
hash-vs-known-good check. Approval tokens are single-use HMAC, hashed at
rest, TTL-bound.
## Consequences
- Documented residual risks: plaintext LAN break-glass hop (emergency use),
Postgres as shared dependency of all roles, macOS host itself unmanaged
by the OS.
- Rotation cadences: actuator SSH key 6mo, machine tokens 90d, webhook
HMAC 1y — scheduler raises expiry signals 2 weeks ahead.

View File

@@ -1,20 +0,0 @@
# ADR 0008 — Forward-only migrations
Status: accepted (2026-07-07) · Plan: rev 3 (resolves audit D5/O1)
## Context
Down-migrations are rarely tested and lie about reversibility once data
has flowed. Rollback needs a strategy that works with real data.
## Decision
golang-migrate, embedded (//go:embed), up-only. Migrations run in a
one-shot init container with a DDL-only DB user before app roles start.
Within one deploy window migrations are additive-only (new columns
nullable, new tables optional) so previous-SHA images tolerate the new
schema. Rollback = redeploy previous image tag; if the migration itself is
the problem, pg_restore the automatic pre-deploy dump. Mistakes roll
forward via compensating migrations.
## Consequences
- No down.sql to write or test; the pre-deploy dump is the real safety net.
- Destructive schema changes (drop/rename) take two deploys by design.

View File

@@ -1,19 +0,0 @@
# ADR 0009 — SSE over WebSocket for the event stream
Status: accepted (2026-07-07) · Plan: rev 3, R3-14
## Context
Live updates (signals, executions, approvals) push server→client only.
Rev 2 specified WebSocket.
## Decision
Server-Sent Events at GET /api/v1/events/stream: plain HTTP (proxies
through Caddy without upgrade handling), native browser EventSource with
auto-reconnect, Last-Event-ID resume backed by the events table. Bounded
per-subscriber buffers with drop-oldest; heartbeat comments every 15s.
Delivery is best-effort — GET /events backfills. Transactional emission +
post-commit LISTEN/NOTIFY feed the stream.
## Consequences
- No bidirectional channel; if one is ever needed (interactive terminals),
add WebSocket alongside — this ADR covers the event feed only.

View File

@@ -1,20 +0,0 @@
# ADR 0010 — Infisical secrets with SOPS DR fallback
Status: accepted (2026-07-07) · Plan: rev 3, Phase 5 (resolves audit S9)
## Context
SOPS+age is file-based: no runtime API, no machine identities, no
rotation tracking, and every consumer needs the age key.
## Decision
Infisical in the Docker stack; services fetch via machine identities;
secrets never in env files or plain config (config hierarchy: defaults →
file → env → Infisical, secrets only). Bootstrap root of trust: Infisical
master key in the mac-mini Keychain, backed up offline. One age key is
retained and all secrets are exported to a SOPS-encrypted fallback file
until an Infisical restore drill has passed; the fallback is refreshed on
rotation.
## Consequences
- Chicken-and-egg is explicit: the Keychain + offline copy are the root.
- SOPS retirement is gated on a passed restore drill, not on the calendar.

View File

@@ -1,18 +0,0 @@
# Architecture Decision Records
MADR-style records for Oikos. One decision per file, numbered, never edited
after acceptance — superseding decisions get a new ADR that links back.
Statuses: proposed | accepted | superseded-by-NNNN.
| ADR | Title |
|---|---|
| [0001](0001-go-single-binary.md) | Go with single-binary role packaging |
| [0002](0002-postgres-timescale-only-datastore.md) | PostgreSQL + TimescaleDB as the only datastore |
| [0003](0003-db-native-ontology-yaml-seeds.md) | DB-native ontology with YAML seed manifests |
| [0004](0004-openapi-first.md) | Contract-first OpenAPI API |
| [0005](0005-uuidv7-plus-slug-identity.md) | UUIDv7 + slug entity identity |
| [0006](0006-learning-proposal-only.md) | Learning is proposal-only (no self-authorization) |
| [0007](0007-threat-model.md) | Threat model and trust zones |
| [0008](0008-forward-only-migrations.md) | Forward-only migrations |
| [0009](0009-sse-over-websocket.md) | SSE over WebSocket for the event stream |
| [0010](0010-infisical-with-sops-fallback.md) | Infisical secrets with SOPS DR fallback |

36
go.mod
View File

@@ -1,36 +0,0 @@
module github.com/dtoro/oikos
go 1.26.3
require (
github.com/getkin/kin-openapi v0.140.0
github.com/go-chi/chi/v5 v5.3.1
github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/jsonschema-go v0.4.3
github.com/google/uuid v1.6.0
github.com/jackc/pgx/v5 v5.10.0
github.com/modelcontextprotocol/go-sdk v1.6.1
github.com/oapi-codegen/runtime v1.4.2
golang.org/x/crypto v0.53.0
golang.org/x/sync v0.21.0
gopkg.in/yaml.v3 v3.0.1
)
require (
github.com/apapsch/go-jsonmerge/v2 v2.0.0 // indirect
github.com/go-openapi/jsonpointer v0.22.5 // indirect
github.com/go-openapi/swag/jsonname v0.25.5 // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/oasdiff/yaml v0.1.0 // indirect
github.com/oasdiff/yaml3 v0.0.13 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 // indirect
github.com/segmentio/asm v1.1.3 // indirect
github.com/segmentio/encoding v0.5.4 // indirect
github.com/yosida95/uritemplate/v3 v3.0.2 // indirect
golang.org/x/oauth2 v0.35.0 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/text v0.38.0 // indirect
)

88
go.sum
View File

@@ -1,88 +0,0 @@
github.com/RaveNoX/go-jsoncommentstrip v1.0.0/go.mod h1:78ihd09MekBnJnxpICcwzCMzGrKSKYe4AqU6PDYYpjk=
github.com/apapsch/go-jsonmerge/v2 v2.0.0 h1:axGnT1gRIfimI7gJifB699GoE/oq+F2MU7Dml6nw9rQ=
github.com/apapsch/go-jsonmerge/v2 v2.0.0/go.mod h1:lvDnEdqiQrp0O42VQGgmlKpxL1AP2+08jFMw88y4klk=
github.com/bmatcuk/doublestar v1.1.1/go.mod h1:UD6OnuiIn0yFxxA2le/rnRU1G4RaI4UvFv1sNto9p6w=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI=
github.com/dlclark/regexp2 v1.11.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8=
github.com/getkin/kin-openapi v0.140.0 h1:JFn675aXRFjyiZKa/BFWploGldQlI0gobp4J5k0EZ2g=
github.com/getkin/kin-openapi v0.140.0/go.mod h1:lISrB64F0CPcuDJ3LdtPTMJBY8VENjR9wJBdrcT6J3g=
github.com/go-chi/chi/v5 v5.3.1 h1:3j4HZLGZQ3JpMCrPJF/Jl3mYJfWLKBfNJ6quurUGCf8=
github.com/go-chi/chi/v5 v5.3.1/go.mod h1:R+tYY2hNuVUUjxoPtqUdgBqevM9s9njzkTLutVsOCto=
github.com/go-openapi/jsonpointer v0.22.5 h1:8on/0Yp4uTb9f4XvTrM2+1CPrV05QPZXu+rvu2o9jcA=
github.com/go-openapi/jsonpointer v0.22.5/go.mod h1:gyUR3sCvGSWchA2sUBJGluYMbe1zazrYWIkWPjjMUY0=
github.com/go-openapi/swag/jsonname v0.25.5 h1:8p150i44rv/Drip4vWI3kGi9+4W9TdI3US3uUYSFhSo=
github.com/go-openapi/swag/jsonname v0.25.5/go.mod h1:jNqqikyiAK56uS7n8sLkdaNY/uq6+D2m2LANat09pKU=
github.com/go-openapi/testify/v2 v2.4.0 h1:8nsPrHVCWkQ4p8h1EsRVymA2XABB4OT40gcvAu+voFM=
github.com/go-openapi/testify/v2 v2.4.0/go.mod h1:HCPmvFFnheKK2BuwSA0TbbdxJ3I16pjwMkYkP4Ywn54=
github.com/golang-jwt/jwt/v5 v5.3.1 h1:kYf81DTWFe7t+1VvL7eS+jKFVWaUnK9cB1qbwn63YCY=
github.com/golang-jwt/jwt/v5 v5.3.1/go.mod h1:fxCRLWMO43lRc8nhHWY6LGqRcf+1gQWArsqaEUEa5bE=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/google/jsonschema-go v0.4.3 h1:/DBOLZTfDow7pe2GmaJNhltueGTtDKICi8V8p+DQPd0=
github.com/google/jsonschema-go v0.4.3/go.mod h1:r5quNTdLOYEz95Ru18zA0ydNbBuYoo9tgaYcxEYhJVE=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
github.com/jackc/pgx/v5 v5.10.0 h1:VhSvgU2jSli8o3AqIEOTJr7rZwAEUVo4E4XhR94Zfr0=
github.com/jackc/pgx/v5 v5.10.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/juju/gnuflag v0.0.0-20171113085948-2ce1bb71843d/go.mod h1:2PavIy+JPciBPrBUjwbNvtwB6RQlve+hkpll6QSNmOE=
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0=
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/modelcontextprotocol/go-sdk v1.6.1 h1:0zOSupjKUxPKSocPT1Wtago+mUHU2/uZ4xSOY0FGReU=
github.com/modelcontextprotocol/go-sdk v1.6.1/go.mod h1:kzm3kzFL1/+AziGOE0nUs3gvPoNxMCvkxokMkuFapXQ=
github.com/oapi-codegen/nullable v1.1.0 h1:eAh8JVc5430VtYVnq00Hrbpag9PFRGWLjxR1/3KntMs=
github.com/oapi-codegen/nullable v1.1.0/go.mod h1:KUZ3vUzkmEKY90ksAmit2+5juDIhIZhfDl+0PwOQlFY=
github.com/oapi-codegen/runtime v1.4.2 h1:GMxFVYLzoYLua+/KvzgSphkyK1lLTReQI9Vf4hvATKE=
github.com/oapi-codegen/runtime v1.4.2/go.mod h1:GwV7hC2hviaMzj+ITfHVRESK5J2W/GefVwIND/bMGvU=
github.com/oasdiff/yaml v0.1.0 h1:0bqZjfKc/8S9urj4JuwepX41WX9EoA6ifhU3SV06cXg=
github.com/oasdiff/yaml v0.1.0/go.mod h1:kOlRmMdL2X3vucLCEQO5u61SU22RysnfXvcttrZA1O0=
github.com/oasdiff/yaml3 v0.0.13 h1:06svmvOHOVBqF81+sY2EUScvUI/iS/vl2VIeUUxZQwg=
github.com/oasdiff/yaml3 v0.0.13/go.mod h1:y5+oSEHCPT/DGrS++Wc/479ERge0zTFxaF8PbGKcg2o=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/rogpeppe/go-internal v1.15.0 h1:D0RCU5rMAp+SpgkiNdrjfJ+LX4J1M32V2NeCY7EJ6hc=
github.com/rogpeppe/go-internal v1.15.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 h1:KRzFb2m7YtdldCEkzs6KqmJw4nqEVZGK7IN2kJkjTuQ=
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU=
github.com/segmentio/asm v1.1.3 h1:WM03sfUOENvvKexOLp+pCqgb/WDjsi7EK8gIsICtzhc=
github.com/segmentio/asm v1.1.3/go.mod h1:Ld3L4ZXGNcSLRg4JBsZ3//1+f/TjYl0Mzen/DQy1EJg=
github.com/segmentio/encoding v0.5.4 h1:OW1VRern8Nw6ITAtwSZ7Idrl3MXCFwXHPgqESYfvNt0=
github.com/segmentio/encoding v0.5.4/go.mod h1:HS1ZKa3kSN32ZHVZ7ZLPLXWvOVIiZtyJnO1gPH1sKt0=
github.com/spkg/bom v0.0.0-20160624110644-59b7046e48ad/go.mod h1:qLr4V1qq6nMqFKkMo8ZTx3f+BZEkzsRUY10Xsm2mwU0=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/yosida95/uritemplate/v3 v3.0.2 h1:Ed3Oyj9yrmi9087+NczuL5BwkIc4wvTb5zIM+UJPGz4=
github.com/yosida95/uritemplate/v3 v3.0.2/go.mod h1:ILOh0sOhIJR3+L/8afwt/kE++YT040gmv5BQTMR2HP4=
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
golang.org/x/oauth2 v0.35.0 h1:Mv2mzuHuZuY2+bkyWXIHMfhNdJAdwW3FuWeCPYN5GVQ=
golang.org/x/oauth2 v0.35.0/go.mod h1:lzm5WQJQwKZ3nwavOZ3IS5Aulzxi68dUSgRHujetwEA=
golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.46.0 h1:noSf2Fq6F8DBgS+LysIkx7rIExoNHJsxOAtPp4rthXw=
golang.org/x/sys v0.46.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.44.0 h1:0rLvDRCtNj0gZkyIXhCyOb2OAzEhLVqc4B+hrsBhrmc=
golang.org/x/term v0.44.0/go.mod h1:7ze4MdzUzLXpSAoFP1H0bOI9aXDqveSvatT5vKcFh2Y=
golang.org/x/text v0.38.0 h1:sXmwo9DwP3OK9EZ7PqAdaooSGozfl/3a6/xJcbzPRhE=
golang.org/x/text v0.38.0/go.mod h1:YXZt3QhHUKYT53r2lLKFIVi6Ao1jdzrTR/KQ09qyxF4=
golang.org/x/tools v0.45.0 h1:18qN3FAooORvApf5XjCXgsuayZOEtXf6JK18I3+ONa8=
golang.org/x/tools v0.45.0/go.mod h1:LuUGqqaXcXMEFEruIVJVm5mgDD8vww/z/SR1gQ4uE/0=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=

View File

@@ -5,7 +5,6 @@ name: apps
kind: lxc
os: linux
role: docker-apps
state: active
host: hubris
pve_id: 105
lan_ip: 192.168.8.205
@@ -34,29 +33,23 @@ services_hosted:
- name: artifacto
backend: apps
url: https://artifacto.hubris.network
doc_page: knowledge/wiki/containers/105-apps.md
config_repo: dtoro/Artifacto
- name: homelab_mcp
backend: apps
port: 9810
systemd_unit: homelab-mcp
public_host: mcp.hubris.network
endpoint: https://mcp.hubris.network/mcp
doc_page: knowledge/wiki/infrastructure/homelab-context.md
config_repo: dtoro/Homelab-Docs
endpoint: https://mcp.hubris.network/sse
note: MCP server. Read-only context + management. Reachable on the LAN via Caddy and from off-LAN via
Netbird (192.168.8.0/24 is a network resource routed through hubris).
risk_notes: "agents' primary read surface \u2014 outage degrades every agent to grepping the clone"
- name: secrets_issuance
backend: apps
port: 9820
systemd_unit: secrets-issuance
public_host: secrets.hubris.network
endpoint: https://secrets.hubris.network/issue
doc_page: .agents/operations/agent-enrollment.md
config_repo: dtoro/Homelab-Docs
note: Issues per-client age private keys. Gated at source-IP layer (mesh + LAN subnets in MESH_SUBNETS).
risk_notes: "identity issuance \u2014 any change is security-sensitive; key operations are destructive-class"
age_pubkey: age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/105-apps.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,10 +5,9 @@ name: arriman
kind: lxc
os: linux
role: arr-stack
state: active
host: strong
host: hubris
pve_id: 122
lan_ip: 192.168.8.245
lan_ip: 192.168.8.132
mesh:
tailscale:
fqdn: arr
@@ -18,7 +17,7 @@ mesh_globals:
- netbird
- tailscale
mounts:
- /mnt/media_local
- /mnt/library
public_hosts:
- jellyseerr.hubris.network
- qbit.hubris.network
@@ -29,10 +28,7 @@ services_hosted:
- name: arr_stack
backend: arriman
note: jellyseerr / qbit / sab on docker compose
doc_page: knowledge/wiki/containers/122-arriman.md
notes:
- Migrated from hubris to strong 2026-07-05 (Phase 2). Library on ludo-lvm.
- Contains homarr, radarr, sonarr, lidarr, sabnzbd, qbittorrent, bazarr, flaresolverr, prowlarr, jellyseerr
- qBittorrent auth subnet whitelist expanded to 192.168.8.0/24 (for seanime + Caddy access)
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/122-arriman.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

32
hosts/authentik.yaml Normal file
View File

@@ -0,0 +1,32 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: authentik
kind: lxc
os: linux
role: idp
host: hubris
pve_id: 124
lan_ip: 192.168.8.180
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
public_host: auth.hubris.network
runs:
- authentik
- dnsmasq
services_hosted:
- name: authentik
url: https://auth.hubris.network
backend: authentik
- name: dnsmasq
backend: authentik
note: split-horizon DNS, /etc/dnsmasq.d/hubris-split.conf
notes:
- 'Also hosts split-horizon dnsmasq: /etc/dnsmasq.d/hubris-split.conf'
see_also:
- containers/124-authentik.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: caddy
kind: lxc
os: linux
role: reverse-proxy
state: active
host: hubris
pve_id: 121
lan_ip: 192.168.8.175
@@ -24,12 +23,10 @@ services_hosted:
backend: caddy
role: reverse-proxy
note: terminates all *.hubris.network
doc_page: knowledge/wiki/containers/121-caddy.md
config_repo: dtoro/caddy-conf
risk_notes: "wide blast radius \u2014 every *.hubris.network route rides on it (see oikos/policy.yaml\
\ service_overrides)"
notes:
- Terminates all *.hubris.network
- /etc/caddy is a git checkout of dtoro/caddy-conf
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/121-caddy.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,20 +1,22 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: auth-outpost
name: claudio-bot
kind: lxc
os: linux
role: authentik-gateway
state: active
role: matrix-agent
host: hubris
pve_id: 106
lan_ip: 192.168.8.6
pve_id: 123
lan_ip: 192.168.8.230
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
notes:
- Runs Authentik outpost (reverse-proxy/SSO enforcement) for protected services
mcp_endpoint: https://mcp.hubris.network/mcp
- Reads /opt/homelab-context/ on startup
age_pubkey: age1xmkeq968areza2necqyq0065dpeegngzyr6dhagh0n6pl33lccfqe5mqn9
see_also:
- containers/123-claudio-bot.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,29 +0,0 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: dns
kind: lxc
os: linux
role: dns-server
state: active
host: hubris
pve_id: 107
lan_ip: 192.168.8.2
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
runs:
- dns
services_hosted:
- name: dns
backend: dns
note: Technitium DNS, split-horizon zone
doc_page: knowledge/wiki/containers/107-dns.md
risk_notes: "LAN-wide resolver \u2014 misconfig breaks name resolution for every client"
notes:
- Technitium DNS, split-horizon zone for *.hubris.network
- Primary DNS for 192.168.8.0/24 LAN (inventory.services.dns references this)
mcp_endpoint: https://mcp.hubris.network/mcp
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,12 +5,12 @@ name: elementsynapse
kind: lxc
os: linux
role: matrix-server
state: active
host: strong
host: hubris
pve_id: 118
lan_ip: 192.168.8.242
lan_ip: 192.168.8.239
mesh:
tailscale: {}
tailscale:
fqdn: elementsynapse
mesh_globals:
primary: netbird
accepted:
@@ -23,9 +23,7 @@ services_hosted:
- name: matrix
url: https://matrix.hubris.network
backend: elementsynapse
doc_page: knowledge/wiki/containers/118-elementsynapse.md
risk_notes: "alert/approval channel for Oikos \u2014 outage silences agent escalation"
notes:
- Migrated from hubris to strong 2026-07-05 (Phase 1 of strong migration plan).
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/118-elementsynapse.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: gitea
kind: lxc
os: linux
role: git-server
state: active
host: hubris
pve_id: 104
lan_ip: 192.168.8.121
@@ -27,10 +26,9 @@ services_hosted:
url: https://git.hubris.network
backend: gitea
backend_url: http://192.168.8.121:3000
doc_page: knowledge/wiki/containers/104-gitea.md
config_repo: dtoro/gitea-customizations
risk_notes: hosts all config repos + deploy webhooks; outage blocks auto-deploy and sync
notes:
- Bare repos live at /mnt/library/repos/dtoro/*.git
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/104-gitea.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,25 +0,0 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: grimmory
kind: lxc
os: linux
role: book-library
state: active
host: strong
pve_id: 130
lan_ip: 192.168.8.247
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
mounts:
- /mnt/media_local
public_host: books.hubris.network
notes:
- Docker host for Grimmory (community fork of Booklore). Created 2026-06-29.
- Migrated from hubris to strong 2026-07-05 (Phase 2d). Books on ludo-lvm.
age_pubkey: age1uellsemnjrzgfg9fxw4jefpy05laxzggwnwhh6ny3wl7alyp6v8q0muxet
mcp_endpoint: https://mcp.hubris.network/mcp
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: haos
kind: vm
os: linux
role: home-automation
state: active
host: hubris
pve_id: 108
lan_ip: 192.168.8.101
@@ -22,6 +21,7 @@ runs:
services_hosted:
- name: haos
backend: haos
doc_page: knowledge/wiki/vms/108-haos.md
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- vms/108-haos.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,26 +0,0 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: house
kind: lxc
os: linux
role: family-planner
state: active
host: strong
pve_id: 129
lan_ip: 192.168.8.244
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
public_host: house.hubris.network
notes:
- Docker host for Yuvomi (family planner). Created 2026-06-26.
- Migrated from hubris to strong 2026-07-05 (Phase 1 of strong migration plan).
- Runs Yuvomi container + WebDAV doc bridge to paperless
- 192.168.8.212 was the hubris IP before migration (briefly picked up by teddycloud via DHCP; teddycloud
has since been given a static IP, see hosts.teddycloud)
age_pubkey: age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
mcp_endpoint: https://mcp.hubris.network/mcp
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,17 +1,13 @@
# `hubris` — Proxmox host
Proxmox VE host running 1 VM and 13 LXC containers — the whole homelab's
workloads still live here. As of 2026-07-01, hubris is node 1 of the 2-node
`Homelab` cluster (see [Cluster](#cluster)); the second node is
[strong](strong.md), which hosts nothing yet.
Single-node Proxmox VE running 1 VM and 13 LXC containers. The whole homelab.
## At a glance
- **Role:** Proxmox VE 9.1.2 hypervisor (kernel `6.14.11-4-pve`)
- **Hardware:** GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS (Phoenix APU), 12 vCPU / ~28 GiB RAM, 2× Samsung 990 EVO Plus NVMe (one SSD primary, one for `library` LVM). 2× Realtek RTL8125 NICs (`r8169`).
- **BIOS:** 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
- **Uplink:** `vmbr1` (slave: `eno1`) → SODOLA switch → Fritz!Box 7590. DHCP-reserved `192.168.178.10/24`, gateway `192.168.178.1`.
- **Homelab bridge:** `vmbr0` — portless internal bridge, `192.168.8.77/24` + `192.168.8.1/24` alias (LXC default gateway). All 16 LXCs and the HAOS VM are on `vmbr0`. Proxmox routes between `vmbr0` and `vmbr1`; Fritz!Box has a static route `192.168.8.0/24 → 192.168.178.10`.
- **WiFi:** disabled 2026-06-02 — `wlp3s0` removed from `/etc/network/interfaces`, wpa config deleted. Was used as a failover to the now-retired Slate AX AP.
- **BIOS:** 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../investigations/2026-04-21-hubris-crash-loop.md).
- **LAN (primary):** `192.168.8.77/24` on bridge `vmbr0` (slave: `eno1`), gateway `192.168.8.1`. Default route metric 0.
- **WiFi (failover):** `192.168.8.141/24` on `wlp3s0` (MediaTek MT7922, AX), DHCP from the same router. Default route metric 200. See [Phase 1 WiFi failover](#phase-1-wifi-failover) below.
- **Mesh:** Netbird `wt0` `100.122.38.109/16`. Resolver: `100.122.38.109` (the local netbird daemon, which forwards to LAN/upstream and learns `*.hubris.network` answers via that path). See [mesh](../infrastructure/mesh.md).
- **UI:** `https://proxmox.hubris.network` (via [caddy](../containers/121-caddy.md)) or `https://192.168.8.77:8006`.
@@ -25,37 +21,13 @@ workloads still live here. As of 2026-07-01, hubris is node 1 of the 2-node
`/mnt/library` holds the shared media + data pool: `anime`, `audiobooks`, `books`, `comics`, `documents`, `downloads`, `heaper`, `homecloud`, `images`, `marimo`, `movies`, `music`, `notes`, `podcasts`, `repos`, `roms`, `sophia`. Bind-mounted into every container that needs it. Permissions standard: [media GID 10000](../infrastructure/media-permissions.md).
## Cluster
Member of `Homelab`, a 2-node Proxmox cluster with [strong](strong.md)
(cluster/OS hostname `strong`), formed 2026-07-01.
- **Corosync ring0:** hubris's internal `192.168.8.77` (the `vmbr0` address).
strong reaches it via the existing Fritz!Box static route
(`192.168.8.0/24 → 192.168.178.10`) — no dedicated corosync link, just the
household LAN. Fine for a home cluster; not latency-isolated.
- **Quorum:** 2 nodes, 1 vote each, no QDevice tiebreaker. Quorum needs both
votes — if either node is down (reboot, maintenance, network hiccup), the
survivor's running guests keep working but `/etc/pve` goes read-only:
no start/stop/create/edit until quorum returns. Decided to skip a QDevice
for now; revisit if hubris's periodic reboots (BIOS/thermal work, see
Quirks below) make this painful in practice.
- **Storage:** `local` / `local-lvm` are the standard per-node default IDs
(every node has its own, not actually shared). The `library` lvmthin pool
is explicitly restricted to `nodes hubris` in `/etc/pve/storage.cfg` since
it's a physical thinpool that only exists on this host's hardware.
- strong currently hosts no LXCs/VMs — it exists solely as a cluster
member so far. See [strong.md](strong.md) and the [library-SSD
migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)
for what comes next (physical drive move, service migration — not started).
## Tenants
### VMs
- [108 — `haos-16.3`](../vms/108-haos.md) — Home Assistant OS, 4 GiB / 32 GiB
### LXC containers
See [containers/index](../containers/index.md). 10 active on hubris (101, 118, 122, 129, 130 migrated to [strong](strong.md) 2026-07-05).
See [containers/index](../containers/index.md). 13 active (109 syncthing destroyed 2026-05-14).
## Boot-time tuning (load-bearing)
@@ -64,6 +36,17 @@ See [containers/index](../containers/index.md). 10 active on hubris (101, 118, 1
- **Crash capture:** `/etc/sysctl.d/60-crash-capture.conf` (panic on oops/hardlockup/softlockup/rcu, auto-reboot 10 s), `/etc/modprobe.d/softdog.conf` (`soft_panic=1 soft_margin=60`), `/etc/systemd/system.conf.d/watchdog.conf` (`RuntimeWatchdogSec=15s`). Pstore traces collected to `/var/lib/systemd/pstore/` by `systemd-pstore.service`. **Caveat:** silent CPU lockups leave pstore empty.
- **`rasdaemon`** (Debian pkg) collects MCE / memory / PCIe AER / thermal events to `/var/lib/rasdaemon/ras-mc_event.db`. Query with `ras-mc-ctl --summary` / `--errors`. (mcelog is retired in Debian 13 — don't go looking for it.)
## Phase 1 WiFi failover
Host is dual-homed on LAN (`eno1`/`vmbr0`) and WiFi (`wlp3s0`) so management/SSH stay reachable when LAN drops. **Guests are not yet failed over** — the LXC fleet remains on `vmbr0`/`eno1`. Phase 2 will migrate guest networking off the bridge so the homelab survives full LAN loss.
- WiFi creds in `/etc/wpa_supplicant/wpa_supplicant-wlp3s0.conf` (hashed PSK, mode 600). SSID lives in `/etc/network/interfaces` as `wpa-conf`.
- Both interfaces sit on the same `192.168.8.0/24`; cross-talk avoided with `arp_ignore=1` + `arp_announce=2` on `eno1`/`vmbr0`/`wlp3s0` (set via `post-up` in `/etc/network/interfaces`).
- A second default route at metric 200 is added on `wlp3s0` (post-up). LAN wins while up.
- **Carrier-based failover:** `vmbr0`'s carrier follows the LXC veth members, so it stays `1` even when `eno1` loses link. `ignore_routes_with_linkdown` is therefore not enough on its own. `wan-failover.service` (`/usr/local/sbin/wan-failover.sh`) watches `/sys/class/net/eno1/carrier` via `ip monitor link` and removes/restores the `vmbr0` default route on transitions. Logs to `journalctl -t wan-failover`.
- Failover verified 2026-04-28: `ip link set eno1 down` → outbound HTTP keeps working via WiFi; `ip link set eno1 up` → vmbr0 default restored.
- Reachable on `192.168.8.77` (LAN) and `192.168.8.141` (WiFi); SSH works on either.
## Network performance baseline (2026-05-14)
| Path | Throughput | Notes |
@@ -89,7 +72,7 @@ See [monitoring](../infrastructure/monitoring.md), [backups](../infrastructure/b
## Quirks
- `/etc/pve` is fuse — the Proxmox cluster filesystem, now genuinely cluster-synced (2-node) rather than the single-node-but-still-fuse case this note used to describe.
- `/etc/pve` is fuse — normal for the Proxmox cluster filesystem, even on a single-node install.
- ZFS is **not** in use; storage is LVM-thin + ext4.
- Two Realtek 8125 NICs use the in-tree `r8169` driver, not the OOT `r8125`.
- Hardware is thermally marginal. NVMe sensors live near warn temp under load. Thermal pads installed on the SSDs 2026-04-23; host relocated to a better-ventilated spot 2026-04-29.
@@ -100,8 +83,6 @@ See [monitoring](../infrastructure/monitoring.md), [backups](../infrastructure/b
- `root@hubris` (self, RSA) — local
- `d.toro.v@pm.me` (ed25519) — user's iMac, added 2026-04-22
- `root@strong` (RSA) — strong's cluster-join key, added 2026-07-01 so
`pvecm add` could authenticate without a password prompt
OpenSSH on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:22022` and bypasses `authorized_keys` (OIDC/browser). See [SSH access](../infrastructure/ssh-access.md) for the dual-server gotcha.
@@ -113,20 +94,13 @@ OpenSSH on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:220
- [Media permissions](../infrastructure/media-permissions.md)
- [Monitoring](../infrastructure/monitoring.md)
- [Backups (disabled)](../infrastructure/backups.md)
- [Operations cheatsheet](../../../.agents/operations/commands.md)
- [Investigation: 2026-04-21 crash loop](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md)
- [strong — Proxmox host](strong.md)
- [Operations cheatsheet](../operations/commands.md)
- [Investigation: 2026-04-21 crash loop](../investigations/2026-04-21-hubris-crash-loop.md)
## Changelog
### 2026-07-01 — strong joined as a 2nd cluster node ("Homelab")
User reformatted `strong` (formerly a Linux dev workstation, `192.168.178.181`) to Proxmox VE 9.2.3. Cluster/OS hostname on that box is `strong` (left as-is from install). Bootstrapped root SSH on strong from a one-time console password (installed hubris's existing trusted key set: `root@hubris`, `d.toro.v@pm.me`), then generated a keypair on strong and pre-authorized it here (`root@strong`) so `pvecm add 192.168.8.77 --use_ssh 1` (run from strong) could join without an interactive password prompt. No cabling/routing changes needed — strong reaches hubris's corosync address (`192.168.8.77`) via the existing Fritz!Box static route. Cluster now 2 nodes, quorate, **no QDevice** (explicit choice — see [Cluster](#cluster) above for the quorum tradeoff this implies). strong hosts no guests yet; this is Phase 1 of the [library-SSD migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md), nothing further from that plan has been executed.
### 2026-06-02 — Slate AX retired; SODOLA switch added; network restructured
Replaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed. Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md) and [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).
### 2026-05-14 — LXC 109 (syncthing) decommissioned
User destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy 109 --purge` cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing` was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up. Entry moved to the "recently destroyed" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology); references stripped from [README](../../../README.md), [media-permissions](../infrastructure/media-permissions.md), [vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md).
User destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy 109 --purge` cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing` was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up. Entry moved to the "recently destroyed" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology); references stripped from [README](../README.md), [media-permissions](../infrastructure/media-permissions.md), [vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md).
### 2026-05-14 — network performance baseline captured
First explicit speed snapshot: WAN ↓113.5 / ↑19.9 Mbit (24.6 ms), `eno1` 1 Gb full-duplex negotiated, intra-host `vmbr0` ~34.7 Gbit/s host↔LXC and ~34.8 Gbit/s LXC↔LXC (single TCP stream, zero retransmits). `iperf3` + `speedtest-cli` installed on host. Noted `eno1` `rx_errors` at 1.62 M (~1.7 % of 96 M RX packets in 14 d uptime) plus 10.9 k `align_errors` — flagged for follow-up; expect to recheck the trend in ~1 week, suspect patch cable / switch port first if still climbing. See new "Network performance baseline" section above.
@@ -138,7 +112,7 @@ User destroyed the heaper LXC. No `116.conf.bak` left behind in `/etc/pve/lxc/`.
`/etc/sysctl.d/99-bbr.conf` switches `net.ipv4.tcp_congestion_control` from `cubic` to `bbr` and `net.core.default_qdisc` from `fq_codel` to `fq`. Also bumps `rmem_max`/`wmem_max` to 64 MiB and widens `tcp_rmem`/`tcp_wmem`. `tcp_bbr` module pinned at boot via `/etc/modules-load.d/bbr.conf`. Triggered by Nextcloud client downloads from a WiFi laptop pulling ~2 MB/s despite a 152 Mbps link — server-side baseline through Caddy with BBR is ~400 MB/s single-stream loopback, so any client-perceived single-stream improvement is pure congestion-control win. Touches every LXC's outbound TCP since they all share this kernel.
### 2026-04-29 — relocated to better-ventilated spot
User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 3435.5 °C, ACPI zone 4749 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 6061 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation).
User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 3435.5 °C, ACPI zone 4749 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 6061 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation).
### 2026-04-28 — Phase 1 WiFi failover
Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failover, metric 200) on the GL-AXT1800-714-5G AP. Installed `wpasupplicant`+`iw`; added `wlp3s0` stanza to `/etc/network/interfaces` with `wpa-conf`; ARP isolation sysctls in `post-up`. Built `wan-failover.service` to remove the vmbr0 default route on `eno1` carrier loss, since the bridge's carrier doesn't follow `eno1` (the LXC veths keep it `1`). LXC/VM guests are still LAN-only — Phase 2 will migrate them.
@@ -147,10 +121,10 @@ Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failov
This wiki created. Live state at this date: 14 LXCs running (109 syncthing stopped), 1 VM, kernel `6.14.11-4-pve`, uptime 3 d 0 h post drive-removal A/B test. Compared to memory snapshot from a week ago, **destroyed**: LXC 100 (yunohost arr), 106 (flaresolverr), 107 (marimo), 110 (photoprism), 111 (karakeep), 112 (immich), 115 (reticulum). 100 + 106 destroyed per the planned 2026-04-21 \*arr migration retention; the others removed since.
### 2026-04-23 — SSD cooling + thermal pads installed
Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 6061 °C, nvme1 3840 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict).
Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 6061 °C, nvme1 3840 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict).
### 2026-04-22 — drive removal A/B test
Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md).
### 2026-04-22 — `cpu-epp.service` ordering bug fixed
Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `pve-guests.service`, so the hottest boot window (20+ guests starting on `performance`) preceded EPP application. Now `After=sysinit.target` + `Before=pve-guests.service`.
@@ -159,4 +133,4 @@ Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `
`60-crash-capture.conf`, softdog `soft_panic=1`, RuntimeWatchdog 15 s. `rasdaemon` installed and enabled. Pure silicon hangs still leave no trace; this catches everything else.
### 2026-04-21 — `cpu-epp.service` deployed
Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the [crash-loop incident](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the [crash-loop incident](../investigations/2026-04-21-hubris-crash-loop.md).

View File

@@ -5,7 +5,6 @@ name: hubris
kind: proxmox-host
os: linux
role: hypervisor
state: active
lan_ip: 192.168.8.77
mesh:
netbird:
@@ -29,8 +28,8 @@ services_hosted:
url: https://proxmox.hubris.network
backend: hubris
port: 8006
doc_page: knowledge/wiki/hosts/hubris.md
risk_notes: "hypervisor UI \u2014 changes here affect every guest on the node"
age_pubkey: age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- hosts/hubris.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,10 +5,9 @@ name: jellyfin
kind: lxc
os: linux
role: media-server
state: active
host: strong
host: hubris
pve_id: 101
lan_ip: 192.168.8.246
lan_ip: 192.168.8.206
mesh:
tailscale:
fqdn: jellyfin
@@ -18,7 +17,7 @@ mesh_globals:
- netbird
- tailscale
mounts:
- /mnt/media_local
- /mnt/library
public_host: media.hubris.network
runs:
- jellyfin
@@ -26,14 +25,7 @@ services_hosted:
- name: jellyfin
url: https://media.hubris.network
backend: jellyfin
doc_page: knowledge/wiki/containers/101-jellyfin.md
risk_notes: native Authentik OIDC via SSO-Auth plugin, no Caddy forward-auth gate; VAAPI transcode depends
on GPU passthrough on strong
notes:
- Jellyfin 10.11.11 with VAAPI hardware acceleration (Radeon 680M iGPU on strong)
- 4 cores / 8 GiB RAM / 1 GiB swap
- SSO-Auth plugin v4.0.0.4 with Authentik OIDC (no Caddy forward-auth gate)
- GPU passed via dev0+dev1: /dev/dri/renderD128 + card0
- Migrated from hubris to strong 2026-07-05 (Phase 2). Library on ludo-lvm.
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/101-jellyfin.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,19 +1,17 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: rclone
kind: lxc
name: ludo-mini
kind: workstation
os: linux
role: backup
state: active
role: dev
mesh:
netbird:
fqdn: rclone.netbird.selfhosted
fqdn: ludo-mini.netbird.selfhosted
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
age_pubkey: age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
mcp_endpoint: https://mcp.hubris.network/mcp
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,8 +5,6 @@ name: mac-mini
kind: workstation
os: macos
role: dev
state: active
lan_ip: 192.168.178.182
mesh:
netbird:
fqdn: mac-mini-234-17.netbird.selfhosted
@@ -19,6 +17,5 @@ ssh:
user: dtoro
notes:
- Only macOS in the fleet. Bootstrap uses launchd.
age_pubkey: age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs
mcp_endpoint: https://mcp.hubris.network/mcp
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: mule-images
kind: lxc
os: linux
role: photo-management
state: active
host: hubris
pve_id: 120
lan_ip: 192.168.8.136
@@ -26,7 +25,7 @@ services_hosted:
- name: photos
url: https://photos.hubris.network
backend: mule-images
doc_page: knowledge/wiki/containers/120-mule-images.md
config_repo: dtoro/mule-image
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/120-mule-images.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: netbird-vps
kind: external
os: linux
role: netbird-mgmt
state: active
mesh:
netbird:
ip: 100.122.165.149
@@ -17,16 +16,6 @@ mesh_globals:
- tailscale
ssh:
user: root
runs:
- authentik
services_hosted:
- name: authentik
url: https://auth.hubris.network
backend: netbird-vps
doc_page: knowledge/wiki/containers/106-auth-outpost.md
note: core runs on the VPS since 2026-05-31; LAN forward-auth outpost is auth-outpost (LXC 106) at 192.168.8.6:9000.
Previous backend value "authentik" referenced the retired embedded-outpost host (LXC 124).
risk_notes: "SSO provider \u2014 outage locks login to OIDC/forward-auth services"
notes:
- "Public IONOS VPS \u2014 hosts the vanilla netbird mgmt+signal+relay+dashboard stack + host coturn (see\
\ infrastructure/vps-hardening.md + infrastructure/mesh.md changelog 2026-05-21)."
@@ -36,5 +25,5 @@ notes:
- Configs rendered by `homelab render-vps-configs` from vps/turnserver.conf.tmpl + vps/management.json.tmpl,
with secrets decrypted from secrets/turn-shared-secret.yaml + secrets/netbird-authentik-oidc.yaml on
hubris.
mcp_endpoint: https://mcp.hubris.network/mcp
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: nextcloud
kind: lxc
os: linux
role: file-sync
state: active
host: hubris
pve_id: 114
lan_ip: 192.168.8.224
@@ -26,6 +25,7 @@ services_hosted:
- name: nextcloud
url: https://cloud.hubris.network
backend: nextcloud
doc_page: knowledge/wiki/containers/114-nextcloud.md
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/114-nextcloud.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: nfs-export
kind: lxc
os: linux
role: storage-export
state: active
host: hubris
pve_id: 102
lan_ip: 192.168.8.200
@@ -14,5 +13,7 @@ mesh_globals:
accepted:
- netbird
- tailscale
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/102-nfs-export.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: paperless
kind: lxc
os: linux
role: document-archive
state: active
host: hubris
pve_id: 103
lan_ip: 192.168.8.130
@@ -26,7 +25,7 @@ services_hosted:
- name: paperless
url: https://paperless.hubris.network
backend: paperless
doc_page: knowledge/wiki/containers/103-paperless.md
risk_notes: "document archive \u2014 treat data as irreplaceable; DB operations are destructive-class"
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/103-paperless.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

28
hosts/plato.yaml Normal file
View File

@@ -0,0 +1,28 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: plato
kind: lxc
os: linux
role: app
host: hubris
pve_id: 126
lan_ip: 192.168.8.190
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
mounts:
- /mnt/library/documents/plato
public_host: plato.hubris.network
runs:
- plato
services_hosted:
- name: plato
url: https://plato.hubris.network
backend: plato
see_also:
- containers/126-plato.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,7 +5,6 @@ name: republic-laptop
kind: workstation
os: linux
role: primary-dev
state: active
mesh:
netbird:
fqdn: republic-laptop.netbird.selfhosted
@@ -16,5 +15,5 @@ mesh_globals:
- tailscale
ssh:
user: dtoro
mcp_endpoint: https://mcp.hubris.network/mcp
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,26 +0,0 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: romm
kind: lxc
os: linux
role: rom-manager
state: active
host: strong
pve_id: 134
lan_ip: 192.168.8.249
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
mounts:
- /mnt/media_local
public_host: roms.hubris.network
notes:
- Docker host for RomM (romm.app) self-hosted ROM manager. Created 2026-07-05.
- MariaDB sidecar at /opt/romm/docker-compose.yml.
- ROMs on ludo-lvm media volume at /mnt/media_local/roms.
- 1 core / 2 GiB RAM / 16 GiB rootfs (ludo-lvm).
mcp_endpoint: https://mcp.hubris.network/mcp
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,29 +0,0 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: seanime
kind: lxc
os: linux
role: anime-media-server
state: active
host: strong
pve_id: 133
lan_ip: 192.168.8.248
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
mounts:
- /mnt/media_local/anime
public_host: seanime.hubris.network
notes:
- Seanime anime media server for online streaming + local library scanning
- Created 2026-07-05. Binary at /opt/seanime/bin/seanime, systemd service.
- Connected to qBittorrent on arriman (192.168.8.245:8080)
- 8 online streaming extensions installed (HiAnime, AniWatch, KickAssAnime, etc.)
- /anime mounted from strong ludo-lvm (/mnt/media_local/anime)
- Caddy: "https://seanime.hubris.network \u2192 192.168.8.248:43211"
- qBittorrent auth subnet whitelist expanded to 192.168.8.0/24 for seanime access
mcp_endpoint: https://mcp.hubris.network/mcp
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -5,10 +5,9 @@ name: sophia
kind: lxc
os: linux
role: workshop
state: active
host: hubris
pve_id: 119
lan_ip: 192.168.8.109
lan_ip: 192.168.8.157
mesh:
tailscale:
fqdn: sophia
@@ -19,5 +18,7 @@ mesh_globals:
- tailscale
mounts:
- /mnt/library
mcp_endpoint: https://mcp.hubris.network/mcp
see_also:
- containers/119-sophia.md
mcp_endpoint: https://mcp.hubris.network/sse
secrets_issuance_endpoint: https://secrets.hubris.network/issue

View File

@@ -1,32 +0,0 @@
# Generated by mcp/build_host_files.py from inventory.yaml.
# Do NOT edit by hand — your changes will be overwritten.
# Source of truth: ../inventory.yaml
name: strong
kind: proxmox-host
os: linux
role: hypervisor
state: active
lan_ip: 192.168.178.181
mesh_globals:
primary: netbird
accepted:
- netbird
- tailscale
ssh:
user: root
notes:
- Reformatted from Linux workstation ("ludo-mini" in this wiki, still the machine's nickname) to Proxmox
VE 9.2.3 on 2026-07-01. Renamed the inventory/wiki identity from ludo-mini to strong on the same day
so it matches the OS/cluster hostname everywhere (bootstrap looks up hosts/$(hostname).yaml, so a mismatch
would break enrollment).
- "Joined hubris's \"Homelab\" cluster same day. 2-node, no QDevice tiebreaker yet \u2014 see hosts/hubris.md\
\ quorum note."
- Netbird not yet installed (fresh OS wiped prior enrollment); reachable today only via the household
LAN / existing Fritz static route to 192.168.8.0/24. Re-enroll in mesh as a follow-up if off-LAN access
to this host itself (not just its future guests) is needed.
- "First step of the planned library-SSD migration \u2014 see .hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md\
\ (filename kept as-is, it's a historical planning doc). Only Phase 1 (Proxmox install + cluster join)\
\ is done; no physical drive move, service migration, or GPU passthrough has happened yet."
age_pubkey: age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4
mcp_endpoint: https://mcp.hubris.network/mcp
secrets_issuance_endpoint: https://secrets.hubris.network/issue

Some files were not shown because too many files have changed in this diff Show More