Compare commits
2 Commits
claude/goo
...
claude/jol
| Author | SHA1 | Date | |
|---|---|---|---|
| a45e4f6f29 | |||
| 77b3a6f677 |
227
.agents/OIKOS.md
227
.agents/OIKOS.md
@@ -1,227 +0,0 @@
|
||||
# Oikos — the operating model
|
||||
|
||||
Oikos (Greek: *household*) is the agent operating system layered on this
|
||||
repo. It is not new infrastructure: `inventory.yaml` is the kernel data
|
||||
structure, the `homelab` CLI and MCP server are the syscall surface, and
|
||||
this page defines the rules everything above them follows.
|
||||
|
||||
Read this after [AGENTS.md](../AGENTS.md). Machine-readable companions:
|
||||
[oikos/ontology.yaml](../oikos/ontology.yaml) (systems model),
|
||||
[oikos/policy.yaml](../oikos/policy.yaml) (risk & approval).
|
||||
|
||||
## The kernel loop: OODA
|
||||
|
||||
Every Oikos activity — scheduled probe, agent task, operator request — is
|
||||
one pass through **Observe → Orient → Decide → Act**:
|
||||
|
||||
1. **Observe** — probes, drift detectors, and agent findings produce
|
||||
**Signals** (structured records, not loose messages): pending updates,
|
||||
high temperature, low disk, service down, cert expiry, stale backup,
|
||||
inventory drift.
|
||||
2. **Orient** — walk the ontology graph: what entity is affected, what
|
||||
depends on it (blast radius), its lifecycle state, whether a runbook
|
||||
matches, what the ledger says about past attempts.
|
||||
3. **Decide** — the classifier scores **risk class × blast radius ×
|
||||
confidence** and routes:
|
||||
- **auto-act**: within autonomy policy, high confidence, contained radius
|
||||
- **escalate**: operator approval via Matrix (✅/❌ reaction) or the
|
||||
Oikos Console's `/approvals` page (destructive actions additionally
|
||||
need a typed confirmation phrase either way)
|
||||
- **queue**: informational — console + reports
|
||||
The classifier can only *lower* autonomy relative to policy, never raise
|
||||
it. When in doubt, escalate.
|
||||
4. **Act** — execute through `homelab` commands or runbooks (never ad-hoc
|
||||
SSH), then **verify** with the action's verification command, write a
|
||||
**ledger** entry, resolve the Signal, and update docs in the same session.
|
||||
|
||||
## Primitives
|
||||
|
||||
| Primitive | What it is | Lives in |
|
||||
|---|---|---|
|
||||
| Host / Service | topology entities | `inventory.yaml` (+ generated `hosts/*.yaml`) |
|
||||
| Secret | SOPS+age encrypted value, per-client recipients | `secrets/` + `.sops.yaml` |
|
||||
| Runbook | executable workflow with risk class + verification | `.agents/skills/<name>/SKILL.md` |
|
||||
| Signal | something needing attention, with lifecycle | `signals/` ledger (Week 3) |
|
||||
| Change | one mutation: who, what, risk, approval, verification | `ledger/` (Week 2) |
|
||||
| Approval | short-TTL signed grant for a gated action | approval engine (Week 3) |
|
||||
| Incident | investigation narrative | `knowledge/sources/investigations/` |
|
||||
| Plan | design doc for non-trivial work | `plans/` |
|
||||
| Agent | enrolled client identity = its age pubkey | `inventory.yaml` + `.sops.yaml` |
|
||||
|
||||
## Risk classes (enforced, not advisory)
|
||||
|
||||
From [oikos/policy.yaml](../oikos/policy.yaml):
|
||||
|
||||
- **read_only** — status, logs, docs, inventory. Unattended.
|
||||
- **reversible_low** — restart, cache clear, sync pull. Unattended + ledger.
|
||||
- **config_mutation** — tracked-config edits (commit+push, never local),
|
||||
deploys, upgrades, DNS/ingress changes. Operator approval.
|
||||
- **destructive** — destroy, format, wipe, rotate, revoke. Approval +
|
||||
typed confirmation phrase.
|
||||
|
||||
Lifecycle gates modify these: `provisioning` nodes are freely mutable
|
||||
(nothing depends on them); `deprecated` nodes accept no new dependents;
|
||||
anything touching a `destroyed` node is drift.
|
||||
|
||||
## The systems model
|
||||
|
||||
Eight domains — physical, compute, network, storage, software,
|
||||
identity & access, operations, external — cover everything in the lab;
|
||||
entities are connected by typed edges (`hosts`, `provides`, `mounts`,
|
||||
`stores-on`, `routes-to`, `can-decrypt`, `depends-on`, `backs-up-to`, …)
|
||||
defined in [oikos/ontology.yaml](../oikos/ontology.yaml). Rule of
|
||||
completeness: **if it can break, be changed, or hold data, it has an
|
||||
entity and edges.** Blast-radius questions ("what breaks if strong goes
|
||||
down?") are graph walks, not doc archaeology.
|
||||
|
||||
Nodes move through an explicit lifecycle —
|
||||
`planned → provisioning → active → migrating → deprecated → destroyed` —
|
||||
stored as `state:` in inventory (absent = active). Destroyed nodes live in
|
||||
the `archaeology:` section. Each transition is a runbook checklist;
|
||||
deprecation completes only when inbound edges reach zero.
|
||||
|
||||
Generated views: [infrastructure/topology.md](../knowledge/wiki/infrastructure/topology.md)
|
||||
(Mermaid, regenerated from inventory) and the live, clickable version at
|
||||
`oikos.hubris.network/graph` once the Console is deployed.
|
||||
|
||||
## Conventions carried forward
|
||||
|
||||
- Inventory is the truth; live state wins over narrative docs.
|
||||
- Prefer `homelab` CLI and MCP over ad-hoc SSH.
|
||||
- Meaningful changes update docs in the same session.
|
||||
- Secrets are decrypted locally via per-client keys; never into docs/comments.
|
||||
- Tracked configs change by commit + push, not local edits.
|
||||
- Netbird is the preferred mesh path for new traffic.
|
||||
- Agents are terse ([caveman.md](shared/caveman.md)), verify claims, and fix
|
||||
collateral drift when found.
|
||||
|
||||
## Build status (30-day roadmap, started 2026-07-05)
|
||||
|
||||
- **Week 1**: policy, ontology, service contract, archaeology, topology
|
||||
generator, this brief. Shipped.
|
||||
- **Week 2**: context cards, `homelab service <name> …`, change ledger,
|
||||
`node relations`, runbooks. Shipped.
|
||||
- **Week 3**: ops scheduler + state cache (`homelab service <name> health`
|
||||
is cache-first, `--live` forces a probe), drift detectors, signal engine
|
||||
(`homelab signal …`), decision classifier (`homelab decide …`), approval
|
||||
engine (`homelab approval …` — shared-HMAC grants; Matrix delivery is
|
||||
Hermes's existing `@dtoro:avispero` send path, not a new bot, see
|
||||
`oikos/approve.py`), daily brief + weekly report (`oikos/report.py`).
|
||||
Shipped, except: Prometheus is still `planned` (see
|
||||
[plans/2026-07-05-oikos-prometheus-lxc.md](../plans/2026-07-05-oikos-prometheus-lxc.md)) —
|
||||
trend signals (disk-full prediction, temp creep) wait on that LXC; the
|
||||
scheduler's disk check today is point-in-time only, and CPU/NVMe
|
||||
temperature isn't probed at all yet (no confirmed sensor path on
|
||||
hubris/strong). DNS-vs-inventory and generic tracked-config-cleanliness
|
||||
drift checks are also deferred (see `oikos/drift.py` docstring).
|
||||
- **Week 4**: Oikos Console v0 shipped — signals landing page, service
|
||||
grid + detail, node/blast-radius view, live Mermaid graph, drift view,
|
||||
approvals queue (approve/deny, destructive confirmation-phrase
|
||||
enforced), daily/weekly reports. Server-rendered FastAPI + Jinja2, no
|
||||
SPA build chain, tested end-to-end against live production data (see
|
||||
`oikos/console/`). Deploys as a third webhook on `dtoro/Homelab-Docs`
|
||||
(`/opt/oikos-console`, port :9831) — see
|
||||
[oikos/console/deploy/README.md](../oikos/console/deploy/README.md) for
|
||||
the Caddy route and Gitea webhook registration this repo can't do for
|
||||
itself. Approval grants are now single-use (a second `check_grant` call
|
||||
for the same request fails even within the TTL) and already exact-bound
|
||||
to request id + entity + action.
|
||||
**Not shipped as originally planned:** per-agent *age-key-signed*
|
||||
request authentication — age has no signing primitive (it's an
|
||||
encryption-only keypair format), so "age-key-signed" wasn't
|
||||
buildable as stated. The real alternative (SSH-key signing via
|
||||
`ssh-keygen -Y sign`/`-Y verify`, using each host's already-provisioned
|
||||
SSH key) is real and buildable, but needs SSH public keys recorded in
|
||||
inventory first — not there today. Moved to the 60/90-day backlog.
|
||||
Authentik step-up re-auth on the approve/deny route is documented but
|
||||
needs a live Authentik instance to configure — also backlog.
|
||||
Docs pass done (this file, AGENTS.md, operations/commands.md); found
|
||||
and fixed two more stale references while at it (DNS section still
|
||||
pointed at destroyed LXC 124/dnsmasq instead of Technitium on 107, and
|
||||
a `claudio-monitor` reference that's been deprecated since 2026-06-04).
|
||||
|
||||
### Real drift found while building Week 3 (unresolved, needs operator action)
|
||||
|
||||
The drift detectors surfaced genuine, currently-true findings on first
|
||||
run against production — recorded here rather than silently fixed, since
|
||||
each is a `config_mutation`/`destructive`-class decision:
|
||||
|
||||
- `republic-laptop` has no `age_pubkey:` in `inventory.yaml`, but its real
|
||||
age key is granted on nearly every shared secret in `.sops.yaml`
|
||||
(`age1vf8h7...`) — the enrollment write-back to inventory never
|
||||
happened. Fix: `homelab client add republic-laptop --finalize-pubkey
|
||||
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6`.
|
||||
- `grimmory` has an `age_pubkey` in inventory but is missing from
|
||||
`secrets/hello.yaml`'s recipient list — incomplete enrollment the
|
||||
other direction. Fix: re-run `homelab client add grimmory
|
||||
--finalize-pubkey <its key>`.
|
||||
- `pve_id 131` exists live on hubris (`pct list`) with no inventory entry
|
||||
— investigate before assuming it's a stale ID (see the Prometheus LXC
|
||||
plan doc above, which flags this explicitly).
|
||||
- Three `lifecycle-pve-id-reuse` info findings (100, 106, 107 each shared
|
||||
between an active host and an archaeology entry) — expected/benign ID
|
||||
reuse after destroy, no action needed.
|
||||
|
||||
## 60/90-day backlog
|
||||
|
||||
Derived from gaps observed while building the 30-day roadmap, not
|
||||
guesswork. Roughly ordered by what unblocks the most:
|
||||
|
||||
- **Fix the oikos-console deploy webhook's signature mismatch.** Console
|
||||
is live on apps (105) via a manual `deploy.sh` run, but Gitea webhook
|
||||
14's deliveries all 403 with a signature mismatch for a cause not yet
|
||||
found — the secret is confirmed synced correctly on both sides
|
||||
(rotated once already to rule out drift). Until fixed, `git push`
|
||||
doesn't auto-redeploy the console the way it does for homelab-mcp/
|
||||
secrets-issuance; re-run `deploy.sh` on apps manually after changes.
|
||||
See [oikos/console/deploy/README.md](../oikos/console/deploy/README.md).
|
||||
- **SSH-key-signed approval requests.** Replaces the design note in
|
||||
Week 4: age keys can't sign (encryption-only format), so per-agent
|
||||
request authentication needs `ssh-keygen -Y sign`/`-Y verify` against
|
||||
each host's existing SSH key. Blocked on a schema gap: inventory
|
||||
doesn't record SSH public keys today, only ports/users. First step is
|
||||
populating that field on enrollment, then wiring `oikos/approve.py` to
|
||||
require and verify a signature over the request payload.
|
||||
- **Authentik step-up re-auth** on the Console's `/approvals` POST route
|
||||
— needs a live Authentik `PromptStage`/reauth flow scoped to that path;
|
||||
not configurable without a running instance to test against.
|
||||
- **Prometheus provisioning** (see
|
||||
[plans/2026-07-05-oikos-prometheus-lxc.md](../plans/2026-07-05-oikos-prometheus-lxc.md))
|
||||
— unblocks trend signals (disk-full prediction, temp creep) and real
|
||||
sparklines in the Console; investigate the undocumented `pve_id 131`
|
||||
on hubris first.
|
||||
- **CPU/NVMe temperature probing** in the scheduler — needs a confirmed
|
||||
sensor path on hubris and strong (lm-sensors vs vendor tool) before a
|
||||
real check can be written; guessing one risks a probe that silently
|
||||
never fires.
|
||||
- **DNS-vs-inventory drift check** — compare Technitium zone records
|
||||
against `services.*.url`/`public_host`; not implemented (`oikos/drift.py`
|
||||
has no Technitium API wiring yet).
|
||||
- **Generic tracked-config-cleanliness drift check** — today only caddy's
|
||||
`/etc/caddy` git-checkout path is hardcoded in `oikos/drift.py`; every
|
||||
other service with a `config_repo` needs its local checkout path
|
||||
recorded (a `mutation_path`-style field, same gap Week 1's service
|
||||
contract flagged but didn't backfill) before this generalizes.
|
||||
- **Per-service policy overrides** (`oikos/policy.yaml`
|
||||
`service_overrides`) — schema is ready (caddy/dns already use it);
|
||||
populate more as specific services turn out to need non-default risk
|
||||
classes.
|
||||
- **Incident timeline generator** — stitch ledger + signal history into
|
||||
a single narrative for `knowledge/sources/investigations/` entries instead of writing
|
||||
them by hand.
|
||||
- **Secret access audit** — who-can-decrypt-what report from
|
||||
`.sops.yaml` + inventory `age_pubkey`s, extending what
|
||||
`oikos/drift.py`'s SOPS check already partially does.
|
||||
- **Restore drills** — exercise `backs-up-to` (once populated) by
|
||||
actually restoring from a backup target on a schedule, not just
|
||||
checking freshness.
|
||||
- **Multi-agent delegation model** — more than one agent acting
|
||||
concurrently; needs the ledger's `agent` field to carry real identity
|
||||
(age pubkey, not just hostname) consistently, which it mostly does
|
||||
already but hasn't been stress-tested with concurrent writers.
|
||||
- **Grafana** — only if the Console's own Prometheus-backed sparklines
|
||||
turn out to be insufficient once Prometheus ships.
|
||||
- **"Generalize later" extraction** — the original decision was personal-
|
||||
first, generalize-later (see Week 1). Once patterns stabilize, extract
|
||||
a config-driven Oikos core with no `hubris.network`/`hubris`/`strong`
|
||||
hardcoding, so it's installable on a different homelab.
|
||||
@@ -1,48 +0,0 @@
|
||||
# Knowledge domain — schema
|
||||
|
||||
The knowledge domain is the durable, authoritative current-state documentation of the homelab: one
|
||||
page per node and per cross-cutting system, synthesized from live state and evidence. It answers
|
||||
"what exists and how does it work right now."
|
||||
|
||||
It follows the [LLM Wiki layer model](../../shared/llm-wiki.md) and the
|
||||
[writing-style](../../shared/writing-style.md) and [page-templates](../../shared/page-templates.md)
|
||||
rules.
|
||||
|
||||
## The narrative / substrate split
|
||||
|
||||
The knowledge wiki is **narrative**. It sits alongside a **machine-readable substrate** that it
|
||||
describes but never contains. The split is load-bearing: several programs read the substrate at
|
||||
fixed paths, so the wiki reorganization never moves it.
|
||||
|
||||
| Layer | Location | Consumed by |
|
||||
|-------|----------|-------------|
|
||||
| Substrate — source of truth | `inventory.yaml` (root) | MCP server, `homelab` CLI, `oikos/` scheduler/drift/relations/gen-topology |
|
||||
| Substrate — generated host records | `hosts/*.yaml` (root) | `mcp/server.py` (`HOSTS_DIR`), `bin/homelab`; written by `mcp/build_host_files.py` |
|
||||
| Substrate — kernel + context cards | `oikos/` (code, `oikos/cards/`, `oikos/state.json`) | MCP `explain`, scheduler |
|
||||
| Narrative — synthesized wiki | `knowledge/wiki/{hosts,containers,vms,infrastructure}/` | humans, agents via MCP `get_page` / `search_docs` |
|
||||
| Evidence — immutable sources | `knowledge/sources/` (references + investigations) | synthesis into wiki pages |
|
||||
|
||||
## Wiki pages
|
||||
|
||||
- **Node pages** (`knowledge/wiki/containers/<id>-<name>.md`, `.../vms/<id>-<name>.md`,
|
||||
`.../hosts/<name>.md`) follow the container/host template in
|
||||
[page-templates.md](../../shared/page-templates.md): opening definition, `## At a glance`,
|
||||
`## Role`, service/port map, storage, auto-deploy, `## Related`, `## Changelog`.
|
||||
- **Cross-cutting pages** (`knowledge/wiki/infrastructure/<topic>.md`) follow the cross-cutting
|
||||
template: `## Why`, `## Components`, `## How to apply`, `## Gotchas`, `## Related`, `## Changelog`.
|
||||
- Each `inventory.yaml` host entry carries a `doc_page:` field pointing at its narrative page.
|
||||
Changing where a page lives means updating that field (read by `bin/homelab`).
|
||||
|
||||
## The two logs
|
||||
|
||||
- The per-page **`## Changelog`** records infrastructure changes and is machine-parsed
|
||||
(`get_changelog`, the Oikos ledger). Keep the `### YYYY-MM-DD — title` shape.
|
||||
- **`knowledge/log.md`** is append-only and records *documentation-maintenance* operations only
|
||||
(restructures, source ingests, lint sweeps): `## [YYYY-MM-DD] <op> | <summary>`. It never
|
||||
duplicates the Oikos change ledger (`oikos/ledger.py`).
|
||||
|
||||
## Same-session update rule
|
||||
|
||||
A change to a node updates every page that references it in the same session — the node page, the
|
||||
section `README.md` table, the root `README.md`, the Caddy/DNS/ingress pages, the host page, and
|
||||
`inventory.yaml`. See [page-templates.md](../../shared/page-templates.md#same-session-update-rule).
|
||||
@@ -1,55 +0,0 @@
|
||||
# Operations domain — schema
|
||||
|
||||
The operations domain holds the procedural and time-stamped documentation: runbooks (repeatable
|
||||
procedures), investigations (incident evidence), and plans (design docs for non-trivial work). It
|
||||
follows [writing-style](../../shared/writing-style.md); runbooks and plans use the imperative voice
|
||||
exception.
|
||||
|
||||
Where each kind lives: runbooks are skills under [`.agents/skills/`](../../skills/); operator
|
||||
reference (command cheatsheet, enrollment, Hermes agent) lives in
|
||||
[`.agents/operations/`](../../operations/); investigations are sources under
|
||||
`knowledge/sources/investigations/`; plans stay in the repo-root `plans/` folder (below).
|
||||
|
||||
## Plans always live in `plans/`
|
||||
|
||||
**Any plan or design doc for the Homelab is written into the repo `plans/` folder as
|
||||
`plans/YYYY-MM-DD-slug.md` — never a scratch path, an agent-private plan location, or a chat
|
||||
message.** An agent drafting a plan:
|
||||
|
||||
1. Writes the file under `plans/` using the plan template in [page-templates.md](../../shared/page-templates.md).
|
||||
2. Lists it in `plans/index.md`.
|
||||
3. On completion, moves it to `plans/done/` and updates the index status.
|
||||
|
||||
This is the single source for homelab design intent; keeping it in-repo means the plan is
|
||||
versioned, reviewable, and reachable by MCP `get_page`/`search_docs` like any other doc.
|
||||
|
||||
## Runbooks
|
||||
|
||||
Repeatable procedures are skills — one folder per skill at `.agents/skills/<name>/SKILL.md`, with
|
||||
YAML front-matter that the Oikos policy and lifecycle machinery reads:
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: <name>
|
||||
risk_class: read_only | reversible_low | config_mutation | destructive
|
||||
inputs: [<param>, ...]
|
||||
verification: "<shell expression that proves success>"
|
||||
docs_update_checklist: [<doc artifacts to update>]
|
||||
transition: "<from> -> <to>" # only for lifecycle runbooks
|
||||
---
|
||||
```
|
||||
|
||||
`risk_class` values and the lifecycle `transition` states must match
|
||||
[`oikos/policy.yaml`](../../../oikos/policy.yaml) and [`oikos/ontology.yaml`](../../../oikos/ontology.yaml).
|
||||
|
||||
## Investigations
|
||||
|
||||
Incident records live in `knowledge/sources/investigations/YYYY-MM-DD-slug.md` and are **evidence sources** — written
|
||||
once at incident time, then linked from the changelogs of the nodes they implicate. Sections:
|
||||
`## Summary`, `## Timeline`, `## Root cause`, `## Mitigations applied`, `## Open questions`. Resolved
|
||||
incidents move to `knowledge/sources/investigations/archive/`.
|
||||
|
||||
## The operations log
|
||||
|
||||
`plans/log.md` and `knowledge/log.md` are append-only records of documentation operations on
|
||||
those areas (`## [YYYY-MM-DD] <op> | <summary>`), distinct from the Oikos change ledger.
|
||||
@@ -1,91 +0,0 @@
|
||||
# Operations cheatsheet
|
||||
|
||||
Run from the [hubris host](../../knowledge/wiki/hosts/hubris.md) as root. When working from `/root` on Linux you're already on hubris — don't `ssh hubris` / `ping hubris`.
|
||||
|
||||
## Proxmox CLI
|
||||
|
||||
| Command | Use |
|
||||
| --- | --- |
|
||||
| `pct list` / `qm list` | List LXC containers / VMs |
|
||||
| `pct config <id>` / `qm config <id>` | Container / VM config |
|
||||
| `pct exec <id> -- <cmd>` | Run command inside an LXC without entering it (no initgroups — see [media permissions](../../knowledge/wiki/infrastructure/media-permissions.md)) |
|
||||
| `pct enter <id>` | Shell into a container |
|
||||
| `pct start <id>` / `pct stop <id>` | Boot / halt a container |
|
||||
| `pvesm status` | Storage pools status |
|
||||
| `pvesh get /nodes --output-format json` | Node summary as JSON |
|
||||
| `pvesh get /nodes/hubris/lxc/<id>/status/current` | Live container status |
|
||||
| `pvesh get /cluster/resources --type vm --output-format json` | Bulk per-LXC CPU/mem/disk (used by the `homelab-health-watchdog` Hermes cron — see [monitoring](../../knowledge/wiki/infrastructure/monitoring.md); the old `claudio-monitor` this once fed is deprecated) |
|
||||
| `pveversion` | PVE version |
|
||||
| `journalctl -u pve-cluster -n 100` | PVE service logs |
|
||||
|
||||
## Storage
|
||||
|
||||
- Shared mount: `/mnt/library` (ext4 on lvmthin `library`).
|
||||
- Bind into a container: `pct set <id> -mp<N> /mnt/library/<sub>,mp=/data`
|
||||
- For the standard whole-tree mount: `pct set <id> -mp0 /mnt/library,mp=/mnt/library`. See [media permissions](../../knowledge/wiki/infrastructure/media-permissions.md) for the GID-10000 onboarding recipe.
|
||||
|
||||
## Reverse proxy
|
||||
|
||||
- Caddyfile: `/etc/caddy/Caddyfile` on [LXC 121](../../knowledge/wiki/containers/121-caddy.md).
|
||||
- **CRITICAL:** This file is tracked in `dtoro/caddy-conf` (https://git.hubris.network/dtoro/caddy-conf). Never edit it directly on the LXC — commit + push to the repo instead. Caddy auto-deploys on push (see [auto-deploy](../../knowledge/wiki/infrastructure/auto-deploy.md)). If you edit directly, the change will be lost on the next pull and agents won't know about it.
|
||||
- Hot reload: `pct exec 121 -- systemctl reload caddy`.
|
||||
- Validate: `pct exec 121 -- caddy validate --config /etc/caddy/Caddyfile`.
|
||||
- Git workflow shortcut: `pct exec 121 -- "cd /etc/caddy && git add Caddyfile && git commit -m '...' && git push"`.
|
||||
|
||||
## DNS
|
||||
|
||||
- Split-horizon authority: [Technitium DNS](https://technitium.com) on [dns (107)](../../knowledge/wiki/containers/107-dns.md) at `192.168.8.2:53`. Web UI at `http://192.168.8.2`. (Formerly dnsmasq on the now-destroyed LXC 124 — decommissioned 2026-06-04.)
|
||||
- Add/edit records in the Technitium UI; the NetBird managed zone sync (`scripts/dns-sync.py` cron on 107) picks changes up within ~10 minutes.
|
||||
- Verify: `dig @192.168.8.2 +short <host>.hubris.network`.
|
||||
- See [DNS](../../knowledge/wiki/infrastructure/dns.md).
|
||||
|
||||
## Web access
|
||||
|
||||
- `https://proxmox.hubris.network` or `https://192.168.8.77:8006` — Proxmox UI
|
||||
|
||||
## Telemetry quick checks
|
||||
|
||||
- `ras-mc-ctl --summary` — summary of any RAS events (memory / PCIe AER / thermal) since boot
|
||||
- `ras-mc-ctl --errors` — full event log
|
||||
- `cat /sys/devices/system/cpu/cpu0/cpufreq/energy_performance_preference` — should be `balance_power`
|
||||
- `cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor` — should be `powersave`
|
||||
- `ls /sys/fs/pstore/ /var/lib/systemd/pstore/` — panic traces from a previous crash (empty for pure hardware hangs — see [investigation](../../knowledge/sources/investigations/archive/2026-04-21-hubris-crash-loop.md))
|
||||
|
||||
## Fleet apt operations
|
||||
|
||||
Two `homelab` subcommands wrap the common patterns; both fan out to hubris + every LXC.
|
||||
|
||||
| Command | What it does |
|
||||
| --- | --- |
|
||||
| `homelab apt-audit [--target HOST]` | Per-host table: dpkg-interrupted state, holds, upgradable count, non-apt binaries in system paths, DNS health. Exits nonzero if any host has dpkg-interrupted state. |
|
||||
| `homelab apt-upgrade --target HOST` | Launch `apt update && apt upgrade` inside a transient `systemd-run --collect` unit on the target. Survives ssh teardown. Apt configured with `Acquire::Retries=3` + `ForceIPv4=true`. |
|
||||
| `homelab apt-upgrade --all` | Same, fanned out across the standard targets. |
|
||||
| `homelab apt-upgrade ... --status` | Show running unit + tail `/var/log/homelab-apt-upgrade.log` on each target. |
|
||||
| `homelab apt-upgrade ... --safe` | Take a pre-upgrade snapshot per LXC first (`pct snapshot` → `vzdump` fallback for bind-mounted LXCs). Refuses if any snapshot fails unless `--force`. |
|
||||
| `homelab apt-upgrade ... --force` | Skip both the dpkg-audit gate and snapshot-failure refusal. |
|
||||
|
||||
PVE/kernel deferral on hubris: `homelab apt-upgrade --target hubris` will try every upgrade, including kernel + `pve-*`. To skip those, `apt-mark hold` the relevant packages on hubris first; `homelab apt-audit` shows held packages so you can confirm.
|
||||
|
||||
## Oikos (agent OS layer)
|
||||
|
||||
See [OIKOS.md](../OIKOS.md) for the operating model. Quick reference:
|
||||
|
||||
| Command | What it does |
|
||||
| --- | --- |
|
||||
| `homelab service <name> explain\|health\|docs\|log\|actions\|history` | Service Console v0 — context card, cached health (`--live` to force a probe), docs, logs, safe actions + risk class, ledger history |
|
||||
| `homelab node <name> relations` | Ontology blast-radius query: what this host/service impacts, is affected by, and its full transitive blast radius |
|
||||
| `homelab change preflight <service>` | Dry-run report before mutating: risk class, current health, config repo, verification command |
|
||||
| `homelab decide <action> <entity>` | Decision classifier: risk × blast radius × confidence → auto-act or escalate |
|
||||
| `homelab signal list\|raise\|ack\|resolve\|mute` | The attention layer — pending updates, thresholds, drift, anything needing attention |
|
||||
| `homelab approval request\|list\|reply\|check` | Escalate-route grants (Matrix-delivered via Hermes, or the Oikos Console's `/approvals` page) |
|
||||
| `homelab restart <service> [--approval-id <id>]` | `--approval-id` is required whenever the service's risk class needs approval (e.g. `caddy`, `dns`) — refuses mechanically without a valid grant |
|
||||
|
||||
Oikos Console (read-mostly dashboard): `oikos.hubris.network` once deployed — see [oikos/console/deploy/README.md](../../oikos/console/deploy/README.md).
|
||||
|
||||
## Related
|
||||
- [Hubris host](../../knowledge/wiki/hosts/hubris.md)
|
||||
- [Containers index](../../knowledge/wiki/containers/index.md)
|
||||
- [DNS](../../knowledge/wiki/infrastructure/dns.md)
|
||||
- [Monitoring](../../knowledge/wiki/infrastructure/monitoring.md)
|
||||
- [Auto-deploy](../../knowledge/wiki/infrastructure/auto-deploy.md)
|
||||
- [Runbook: dpkg-interrupted recovery](../skills/runbook-dpkg-interrupted/SKILL.md) — what to do when apt got killed mid-transaction
|
||||
@@ -1,41 +0,0 @@
|
||||
# LLM Wiki — the documentation contract
|
||||
|
||||
How the narrative documentation in this repo is organized. The pattern is borrowed from the
|
||||
`sources / wiki / index / log` model: a durable synthesized layer (`knowledge/wiki/`) built on top
|
||||
of immutable evidence (`knowledge/sources/`, incident records), with pure-listing indexes and an
|
||||
append-only operations log.
|
||||
|
||||
This contract governs the **narrative layer only**. The machine-readable substrate — `inventory.yaml`,
|
||||
generated `hosts/*.yaml`, `oikos/`, `mcp/`, `secrets/`, `bin/` — is not part of the wiki and never
|
||||
moves under it. See [the knowledge schema](../domains/knowledge/schema.md) for the split.
|
||||
|
||||
## Layers
|
||||
|
||||
- **Sources** are immutable raw material: incident records (`knowledge/sources/investigations/`), external reference
|
||||
docs (`knowledge/sources/references/`), and the live system itself (`pct config`, `docker inspect`).
|
||||
Read them; do not rewrite them into other sources.
|
||||
- **Wiki** (`knowledge/wiki/`) is the synthesized, authoritative current-state layer: one page per
|
||||
node (`containers/`, `vms/`, host narratives) and per cross-cutting system (`infrastructure/`). A
|
||||
reader understands the topic from the wiki page without reading the sources.
|
||||
- **Index** (`index.md` / folder `README.md`) is a pure listing — every page in scope with a
|
||||
one-line summary, and nothing else. Anything the section wants to say up front goes into a page
|
||||
the index lists, not into the index.
|
||||
- **Log** (`log.md`) is append-only, recording *doc-maintenance operations* (restructures, source
|
||||
ingests, lint sweeps) in single-line format: `## [YYYY-MM-DD] <op> | <summary>`.
|
||||
|
||||
## Two logs, kept distinct
|
||||
|
||||
- **`## Changelog`** on each node/topic page records *infrastructure* changes to that node. It is
|
||||
machine-parsed (`get_changelog`, the Oikos ledger) — keep the `### YYYY-MM-DD — title` shape.
|
||||
- **`log.md`** per area records *documentation* operations only. It never duplicates the Oikos
|
||||
change ledger (`oikos/ledger.py`), which stays authoritative for infra changes with
|
||||
who/what/risk/approval/verification.
|
||||
|
||||
## Rules
|
||||
|
||||
- Wiki pages stay short and focused. A page past ~300 lines splits.
|
||||
- Pages stay flat under `wiki/<section>/` until there are enough to warrant a sub-group.
|
||||
- Every page follows [writing-style.md](writing-style.md).
|
||||
- Plans and design docs always live in the repo `plans/` folder (`plans/YYYY-MM-DD-slug.md`),
|
||||
listed in `plans/index.md`, moved to `plans/done/` on completion — never a scratch path or a chat
|
||||
message. See [the operations schema](../domains/operations/schema.md).
|
||||
@@ -1,174 +0,0 @@
|
||||
# Page templates for the Homelab Wiki
|
||||
|
||||
The structural templates for each page type. Prose voice, vocabulary, and cross-reference rules live
|
||||
in [writing-style.md](writing-style.md); the layer model (sources / wiki / index / log) lives in
|
||||
[llm-wiki.md](llm-wiki.md).
|
||||
|
||||
## File naming
|
||||
|
||||
**Foundational / entry-point files:** ALL-CAPS
|
||||
|
||||
- **Root level:** `AGENTS.md`, `README.md` — discovery paths for agents and humans.
|
||||
- **Agent instruction** (under `.agents/`): `OIKOS.md`, `HERMES.md` — foundational docs agents read before acting.
|
||||
- **Reference docs:** `GLOSSARY.md` — lookup reference (like classic repo conventions: LICENSE, CHANGELOG, GLOSSARY).
|
||||
|
||||
**Content / narrative pages:** lowercase-with-dashes, date-prefixed as needed
|
||||
|
||||
- **Container pages:** `<id>-<name>.md` (e.g. `101-jellyfin.md`, `132-rclone.md`). The `<id>` is the LXC/VM ordinal from `inventory.yaml`.
|
||||
- **Infrastructure / cross-cutting pages:** `<topic>.md` (e.g. `dns.md`, `auto-deploy.md`, `mesh.md`). Describes a system, not a specific node.
|
||||
- **Plans / investigations:** `YYYY-MM-DD-<slug>.md` (e.g. `2026-07-05-oikos-prometheus-lxc.md`). Date-sorted; slug is lowercase.
|
||||
- **Section indices:** `README.md` (lowercase, conventional). Prefer in folders; `index.md` only if both intro prose and listing coexist.
|
||||
|
||||
**Skills / runbooks:** special case
|
||||
|
||||
- **Folder structure:** `<name>/SKILL.md` where `<name>` is lowercase-with-dashes (e.g. `client-enrollment/SKILL.md`).
|
||||
- **The filename SKILL.md is always uppercase** — it acts as a signpost so tools and humans instantly recognize it as a skill.
|
||||
|
||||
**General rules:** All paths use lowercase letters, numbers, and hyphens (no underscores). Uppercase is reserved for foundational docs (entry points + instruction) and filenames that signify document type (SKILL.md, GLOSSARY.md, etc.).
|
||||
|
||||
## Voice
|
||||
|
||||
Concise, technical, sysadmin-to-sysadmin. No marketing prose, no exclamation marks. Full rules in
|
||||
[writing-style.md](writing-style.md).
|
||||
|
||||
## Page templates
|
||||
|
||||
### Container page (`containers/<id>-<name>.md`)
|
||||
|
||||
```markdown
|
||||
# <id> — `<name>`
|
||||
|
||||
One-sentence purpose.
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `<name>`
|
||||
- **IP:** `192.168.8.x`
|
||||
- **Privilege:** privileged | unprivileged
|
||||
- **Resources:** N cores / M GiB RAM / D GiB rootfs
|
||||
- **Mounts:** `/mnt/library` ↔ `/mnt/library` (if any)
|
||||
- **Public hostname:** `<sub>.hubris.network` (if proxied)
|
||||
|
||||
## Role
|
||||
What it does, what it talks to.
|
||||
|
||||
## Service / port map
|
||||
| Service | Listen | Notes |
|
||||
|
||||
## Storage / config paths
|
||||
|
||||
## Auto-deploy
|
||||
(if any) — link to [auto-deploy](../infrastructure/auto-deploy.md)
|
||||
|
||||
## Related
|
||||
- [Caddy](121-caddy.md) (if proxied)
|
||||
- [DNS](../infrastructure/dns.md) (if has subdomain)
|
||||
- [Authentik](124-authentik.md) (if SSO)
|
||||
- ...
|
||||
|
||||
## Changelog
|
||||
### YYYY-MM-DD — short title
|
||||
What changed, why, link to investigation if any.
|
||||
```
|
||||
|
||||
### Cross-cutting page (`infrastructure/<topic>.md`)
|
||||
|
||||
```markdown
|
||||
# <Topic>
|
||||
|
||||
One-sentence summary.
|
||||
|
||||
## Why
|
||||
Design rationale — what it replaces, what it solves.
|
||||
|
||||
## Components
|
||||
Where it runs, what files matter.
|
||||
|
||||
## How to apply / use
|
||||
Recipes.
|
||||
|
||||
## Gotchas
|
||||
|
||||
## Related
|
||||
Links to nodes that host or depend on this.
|
||||
|
||||
## Changelog
|
||||
```
|
||||
|
||||
### Plan (`plans/YYYY-MM-DD-slug.md`)
|
||||
|
||||
```markdown
|
||||
# YYYY-MM-DD — <title>
|
||||
|
||||
## Goal
|
||||
What this change achieves and why.
|
||||
|
||||
## Current topology / state
|
||||
Diagram or description of what exists now.
|
||||
|
||||
## Target topology / state
|
||||
What it looks like after.
|
||||
|
||||
## Pre-flight checklist
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
## Verification
|
||||
|
||||
## Post-migration
|
||||
Changelog entries to write, index status to update.
|
||||
```
|
||||
|
||||
### Investigation (`knowledge/sources/investigations/YYYY-MM-DD-slug.md`)
|
||||
|
||||
```markdown
|
||||
# YYYY-MM-DD — <title>
|
||||
|
||||
## Summary
|
||||
1-3 sentences.
|
||||
|
||||
## Timeline
|
||||
|
||||
## Root cause
|
||||
|
||||
## Mitigations applied
|
||||
|
||||
## Open questions
|
||||
```
|
||||
|
||||
## Linking discipline
|
||||
|
||||
- Every container page links to every cross-cutting page it participates in.
|
||||
- Every cross-cutting page lists the nodes that participate.
|
||||
- Every investigation links to the nodes it implicates *and* gets back-linked from each node's changelog.
|
||||
- Every plan links to the infrastructure pages it affects. When done, update the plan's status in `plans/index.md` and write changelog entries on affected node pages.
|
||||
|
||||
## Changelog hygiene
|
||||
|
||||
- Reverse-chronological (newest first).
|
||||
- One entry per discrete change, even if you make several in one day.
|
||||
- If a change spans nodes, repeat the entry on each affected page (different perspective is fine).
|
||||
- Don't rewrite history — entries are append-only. Mistakes get a follow-up entry that supersedes them.
|
||||
|
||||
## Same-session update rule
|
||||
|
||||
When you make a change to a node — migrate an LXC, update an IP, change a
|
||||
mount, deploy a new service — **update every relevant doc page in the same
|
||||
session.** A change that touches a container page must also update:
|
||||
|
||||
- The `containers/index.md` table (IPs, host, mounts, status)
|
||||
- The `README.md` table (if the change affects listed columns)
|
||||
- The Caddy page site list (if the change affects `*.hubris.network` routing)
|
||||
- The DNS / ingress infrastructure pages (if the change affects routing)
|
||||
- The `hosts/{hubris,strong}.md` host page (if container count changes)
|
||||
- The `inventory.yaml` host entry (source of truth for the `hosts/*.yaml` generation)
|
||||
- The `infrastructure/topology.md` (generated from inventory, but regen if needed)
|
||||
|
||||
The pattern of updating only one page and leaving stale references on others
|
||||
is a bug. If you're doing a multi-step migration, document the intermediate
|
||||
state with a changelog entry that says "pending — will finalize after Phase
|
||||
N."
|
||||
|
||||
This rule is why Phase 2 of the strong migration (2026-07-05) caused
|
||||
widespread stale data: individual container pages were updated in the
|
||||
changelog but never had their At-a-glance sections, IPs, mount paths, or
|
||||
host attribution updated. Don't repeat that.
|
||||
@@ -1,75 +0,0 @@
|
||||
# Writing Style
|
||||
|
||||
Write like a technical reference, not a marketing page. Every sentence conveys new information.
|
||||
These rules govern **committed documentation** — wiki pages, READMEs, schemas, skills, `AGENTS.md`,
|
||||
plans, investigations, and code comments. They are separate from [caveman.md](caveman.md), which
|
||||
governs an agent's *chat responses*; the two do not conflict.
|
||||
|
||||
New or rewritten pages follow these patterns from day one. Existing pages get updated the next time
|
||||
they are touched.
|
||||
|
||||
## Vocabulary — never use these
|
||||
|
||||
- Significance puffers: "pivotal", "crucial", "vital", "groundbreaking", "transformative", "testament", "paramount", "invaluable".
|
||||
- Analytical verbs: "delve", "leverage", "utilize", "facilitate", "foster", "showcase", "underscore", "streamline", "harness".
|
||||
- Poetic nouns: "tapestry", "landscape" (figurative), "realm", "paradigm", "ecosystem" (figurative), "journey" (figurative), "nexus", "cornerstone".
|
||||
- Promotional adjectives: "robust", "seamless", "innovative", "cutting-edge", "meticulous", "holistic", "comprehensive".
|
||||
- Opening crutches: "In today's world", "In the ever-evolving landscape of", "It's worth noting that", "It is important to note that".
|
||||
|
||||
Use short, common words: "use" not "utilize", "help" not "facilitate", "show" not "demonstrate".
|
||||
|
||||
## Voice
|
||||
|
||||
Describe what systems do and how they work.
|
||||
|
||||
- **Reference prose** (node pages, cross-cutting infrastructure descriptions, `## Role`, `## Why`,
|
||||
`At a glance`) is third-person: state facts about the system, not instructions to a reader.
|
||||
- **Recipes, runbooks, and skills** are the exception: second-person imperative is allowed and
|
||||
preferred where it makes a procedure clearer ("Edit the Caddyfile, commit + push", "Verify with
|
||||
`dig +short`"). This matches how the operator actually works. The vocabulary, structure, and
|
||||
cross-reference rules below still apply.
|
||||
|
||||
## Page shape
|
||||
|
||||
Every doc-level page follows the same shape so a reader scans it in one pass.
|
||||
|
||||
1. **One H1 = the page title.** Node pages use `# <id> — \`<name>\``; topic pages use `# <Topic>`.
|
||||
2. **Opening definition.** First paragraph, 1–3 sentences, says what the thing is. No motivation, no marketing, no setup.
|
||||
3. **Body sections** in the natural order for the topic. Reuse the section templates in [page-templates.md](page-templates.md).
|
||||
4. **`## Changelog`** at the bottom of every node/topic page — reverse-chronological, append-only. This section is machine-parsed (`get_changelog` in `mcp/server.py`); keep the `### YYYY-MM-DD — title` shape.
|
||||
5. **Related links** only at the bottom, only when a reference cannot be woven inline.
|
||||
|
||||
## Section indexes (folder READMEs)
|
||||
|
||||
A folder's `README.md` opens with a 1–3 sentence prose intro that says what the section covers, then
|
||||
a single navigation table — `| Document | What it covers |` — and nothing else. No stale counts, no
|
||||
duplicated prose, no narrative between the intro and the table.
|
||||
|
||||
## Structure rules
|
||||
|
||||
- Make every sentence information-dense. Cut filler, qualifiers, and setup phrases. Lead with the concrete fact or action, not why it matters.
|
||||
- No participial tack-ons (", highlighting the importance of…"). If the clause adds information, make it a separate sentence.
|
||||
- **No meta-commentary about the content itself.** Do not narrate the page's own structure or linking strategy.
|
||||
- Prefer **tables** for enumerable items with internal structure (service/port maps, field lists, status grids). Reserve bullets for short non-structured lists.
|
||||
- Use the **bold-leading-phrase pattern** for structured points: `**Read-only by construction.** The MCP server never mutates state.` — a bold noun phrase, a period, then the explanation.
|
||||
- When enumerating across services or nodes, give each its own `###` sub-section or a table row, not one run-on paragraph.
|
||||
- Use backticks for code, paths, hostnames, and file names (`inventory.yaml`, `192.168.8.77`, `pct config`); italics for first-mention terminology.
|
||||
- Use `>` blockquotes for caveats and gaps that interrupt the main flow: `> **Outstanding gap.** DNS-vs-inventory drift check not yet wired.` One thought per blockquote.
|
||||
|
||||
## Diagrams
|
||||
|
||||
- Mermaid is the default for topology and flow diagrams. `infrastructure/topology.md` is generated by `oikos/gen-topology.py` — do not hand-edit it.
|
||||
- ASCII box diagrams are fine for small shape diagrams; keep them to one screen.
|
||||
|
||||
## Sourcing and cross-references
|
||||
|
||||
- **Factual discipline.** Every claim is grounded in a cited source, an adjacent linked page, or a directly observable fact (`pct config`, `docker inspect`, running config). Do not write sentences that sound sourced but are inference. When docs disagree with live state, fix the doc and note it in the changelog.
|
||||
- **One-sided cross-references.** When two pages relate, the link lives in the page where the connection makes organizational sense. Do not add a back-pointer unless that direction also carries content the reader needs.
|
||||
- **Cross-references are content, not catalog.** Inline links arise from the surrounding prose; the linked page must be needed to understand the current sentence. A bottom-of-page "Related" list is the fallback, not the default.
|
||||
- Pages link with standard relative markdown links (e.g. a container page links to `../infrastructure/dns.md`), forming a navigable graph. Orphans are a bug.
|
||||
|
||||
## Code comments and commit/PR prose
|
||||
|
||||
- Comments explain intent, trade-offs, or constraints the code cannot convey. No diff narration, no type restatement, no section-divider comments.
|
||||
- Commit messages and PR descriptions are problem → change → risk → verification, not a file-by-file diff restatement.
|
||||
- The banned vocabulary applies the same way in comments and commit messages.
|
||||
@@ -1,43 +0,0 @@
|
||||
---
|
||||
name: client-enrollment
|
||||
risk_class: config_mutation
|
||||
inputs: [hostname, kind, role]
|
||||
verification: "homelab doctor (on the new client)"
|
||||
docs_update_checklist: [hosts_narrative_page_if_lxc_or_vm]
|
||||
---
|
||||
|
||||
# Client enrollment
|
||||
|
||||
Goal: bring a new host (workstation, LXC, VM) into inventory and the
|
||||
secrets model, with mesh membership only where it's actually needed.
|
||||
This wraps the existing `homelab client add` flow — see
|
||||
[operations/agent-enrollment.md](../../operations/agent-enrollment.md) for
|
||||
the full walkthrough; this runbook is the risk/lifecycle framing.
|
||||
|
||||
1. On any enrolled client: `homelab client add <hostname>` — appends a
|
||||
`hosts.<name>:` block to `inventory.yaml` (lifecycle `state: planned`
|
||||
→ `provisioning`, per [oikos/ontology.yaml](../../../oikos/ontology.yaml)),
|
||||
commits + pushes.
|
||||
2. Netbird join is **optional, not a required step** — only needed for
|
||||
hosts that must be reachable off-LAN (workstations that roam, e.g.
|
||||
`republic-laptop`, `mac-mini`). A node reachable on the household LAN
|
||||
(192.168.8.0/24 — most LXCs/VMs) doesn't need it: it's already
|
||||
reachable directly, and off-LAN clients reach it too via hubris's
|
||||
routed `192.168.8.0/24` Netbird network resource. Skip this step for
|
||||
LAN-only nodes; do it (out-of-band, console or setup key) only for
|
||||
hosts that need independent off-LAN reachability.
|
||||
3. On the new host: run `bootstrap.sh` (add `--with-hermes` to also
|
||||
enroll the Hermes agent). This provisions `/etc/age/key.txt`, the
|
||||
sync timer, and prints an age pubkey.
|
||||
4. Back on an enrolled client: `homelab client add <hostname>
|
||||
--finalize-pubkey <age1...>` — sets `age_pubkey`, grants shared
|
||||
secrets, re-keys SOPS, commits + pushes. This is the
|
||||
`provisioning → active` transition.
|
||||
5. Verify: `homelab doctor` on the new client should show all checks
|
||||
green (clone, sync timer, age key, CLI symlink, MCP reachable).
|
||||
|
||||
Docs-update checklist: if the new host is an LXC/VM, add its narrative
|
||||
page under `containers/` or `vms/` and set `doc_page` in its inventory
|
||||
entry (host-level cards don't have a `doc_page` field yet — services do;
|
||||
narrative pages are still found via the generated `see_also` in
|
||||
`hosts/<name>.yaml`).
|
||||
@@ -1,34 +0,0 @@
|
||||
---
|
||||
name: config-change-deploy
|
||||
risk_class: config_mutation
|
||||
inputs: [service_name, change_description]
|
||||
verification: "curl -sf <service_url> (or homelab service <name> health)"
|
||||
docs_update_checklist: [doc_page, changelog]
|
||||
---
|
||||
|
||||
# Config change + deploy
|
||||
|
||||
Goal: change a tracked config repo (Caddy, Gitea customizations, an app's
|
||||
own repo) and get it live, safely.
|
||||
|
||||
1. `homelab change preflight <service>` — current health, the service's
|
||||
`config_repo`, its risk class, and the verification command to run
|
||||
after. If risk class requires approval (`config_mutation` or
|
||||
`destructive`), stop and get operator sign-off before editing — see
|
||||
`oikos/policy.yaml`.
|
||||
2. Clone/pull the `config_repo` (never edit the backend's working tree
|
||||
directly — tracked configs change by commit + push, per
|
||||
[OIKOS.md](../../OIKOS.md) conventions).
|
||||
3. Make the change, commit, push to `main`.
|
||||
4. The Gitea webhook fires the deploy pipeline for that repo (see
|
||||
[infrastructure/auto-deploy.md](../../../knowledge/wiki/infrastructure/auto-deploy.md) for
|
||||
the exact receiver/reload for this service).
|
||||
5. Run the preflight's verification command. If it fails, check
|
||||
`homelab service <name> log` for the reload/restart error.
|
||||
6. Record the change: once `oikos/ledger.py` is wired into deploy tooling
|
||||
(Week 3), this is automatic; until then, note the change and outcome
|
||||
in the relevant investigation/plan doc.
|
||||
|
||||
Docs-update checklist: update the service's `doc_page` if the change
|
||||
alters its behavior, ingress route, or ownership; add a changelog entry
|
||||
if the page has one.
|
||||
@@ -1,25 +0,0 @@
|
||||
---
|
||||
name: docs-lint
|
||||
risk_class: read_only
|
||||
inputs: [paths]
|
||||
verification: "python3 .agents/skills/docs-lint/lint.py"
|
||||
docs_update_checklist: []
|
||||
---
|
||||
|
||||
# Docs lint
|
||||
|
||||
Check committed documentation against the mechanical rules in
|
||||
[writing-style.md](../../shared/writing-style.md): banned vocabulary and broken relative markdown
|
||||
links. Prose-voice rules are not machine-checkable — those stay a review responsibility.
|
||||
|
||||
Run from the repo root:
|
||||
|
||||
python3 .agents/skills/docs-lint/lint.py # default: knowledge/ .agents/ operations/ investigations/ plans/
|
||||
python3 .agents/skills/docs-lint/lint.py knowledge/wiki/containers/104-gitea.md
|
||||
|
||||
Exit code is non-zero when any violation is found, so it can gate a commit. The banned-vocabulary
|
||||
list mirrors `writing-style.md`; update both together if the standard changes.
|
||||
|
||||
> **Known baseline.** `knowledge/wiki/containers/101-jellyfin.md` links into a sibling repo
|
||||
> (`devops/homelab-authentik-admin`) that this checkout does not contain — expected, not a bug.
|
||||
> Any other broken link is a real regression; investigate before dismissing it as baseline noise.
|
||||
@@ -1,69 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Lint committed docs against .agents/shared/writing-style.md.
|
||||
|
||||
Checks two mechanical rules:
|
||||
1. Banned vocabulary (significance puffers, analytical verbs, poetic nouns,
|
||||
promotional adjectives, opening crutches).
|
||||
2. Broken relative markdown links.
|
||||
|
||||
Prose-voice rules are not machine-checkable; this covers the parts that are.
|
||||
Run from the repo root: python3 .agents/skills/docs-lint/lint.py [paths...]
|
||||
Exit 1 if any violation is found.
|
||||
"""
|
||||
import os, re, sys
|
||||
|
||||
BANNED = [
|
||||
"pivotal", "crucial", "vital", "groundbreaking", "transformative", "testament",
|
||||
"paramount", "invaluable", "delve", "leverage", "utilize", "facilitate", "foster",
|
||||
"showcase", "underscore", "streamline", "harness", "tapestry", "realm", "paradigm",
|
||||
"nexus", "cornerstone", "robust", "seamless", "innovative", "cutting-edge",
|
||||
"meticulous", "holistic", "comprehensive", "in today's world",
|
||||
"it's worth noting", "it is important to note",
|
||||
]
|
||||
BAN_RE = re.compile(r'(?<![\w-])(' + "|".join(re.escape(w) for w in BANNED) + r')(?![\w-])', re.I)
|
||||
LINK = re.compile(r'\]\(([^)]+)\)')
|
||||
|
||||
def iter_md(paths):
|
||||
for p in paths:
|
||||
if os.path.isfile(p) and p.endswith(".md"):
|
||||
yield p
|
||||
for root, dirs, files in os.walk(p):
|
||||
dirs[:] = [d for d in dirs if d not in (".git", "node_modules")]
|
||||
for f in files:
|
||||
if f.endswith(".md"):
|
||||
yield os.path.join(root, f)
|
||||
|
||||
def main(argv):
|
||||
paths = argv or ["knowledge", ".agents", "operations", "investigations", "plans"]
|
||||
violations = 0
|
||||
# The style guide and this skill enumerate the banned words by definition.
|
||||
ban_exempt = ("shared/writing-style.md", "skills/docs-lint/")
|
||||
for f in sorted(set(iter_md(paths))):
|
||||
check_banned = not any(x in f for x in ban_exempt)
|
||||
fence = False
|
||||
with open(f) as fh:
|
||||
for ln, line in enumerate(fh, 1):
|
||||
if line.lstrip().startswith("```"):
|
||||
fence = not fence; continue
|
||||
if fence:
|
||||
continue
|
||||
if check_banned:
|
||||
for m in BAN_RE.finditer(line):
|
||||
print(f"{f}:{ln}: banned word '{m.group(1)}'")
|
||||
violations += 1
|
||||
for m in LINK.finditer(line):
|
||||
link = m.group(1)
|
||||
if re.match(r'^(https?:|mailto:|#|/)', link):
|
||||
continue
|
||||
path = re.split(r'[#?]', link)[0]
|
||||
if not path:
|
||||
continue
|
||||
tgt = os.path.normpath(os.path.join(os.path.dirname(f), path))
|
||||
if not os.path.exists(tgt):
|
||||
print(f"{f}:{ln}: broken link -> {link}")
|
||||
violations += 1
|
||||
print(f"\n{violations} violation(s)")
|
||||
return 1 if violations else 0
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
@@ -1,34 +0,0 @@
|
||||
---
|
||||
name: incident-investigation
|
||||
risk_class: read_only
|
||||
inputs: [symptom, affected_entity]
|
||||
verification: "n/a — investigation produces a written record, not a state change"
|
||||
docs_update_checklist: [investigations_entry]
|
||||
---
|
||||
|
||||
# Incident investigation
|
||||
|
||||
Goal: understand what broke and why, before touching anything.
|
||||
|
||||
1. `homelab service <name> explain` (or `homelab node <name> relations`
|
||||
if the affected entity is a host) — get the blast radius and doc
|
||||
pointer first. Don't start pulling logs blind.
|
||||
2. `homelab service <name> health` + `homelab service <name> log` (or
|
||||
MCP `get_service_status` / `tail_log`) for the affected service.
|
||||
3. Walk the blast radius: is a shared dependency down (`caddy`, `dns`,
|
||||
`authentik`, or the backend host itself)? `homelab node <name>
|
||||
relations` shows "affected by" — check those first.
|
||||
4. `homelab apt-audit` if the symptom looks like a dpkg/upgrade
|
||||
interaction.
|
||||
5. Check the change ledger for recent mutations to the affected entity
|
||||
or anything upstream of it: `homelab service <name> history` (once
|
||||
populated) or grep `ledger/*.jsonl`.
|
||||
6. Write findings to a new `knowledge/sources/investigations/<date>-<slug>.md` — symptom,
|
||||
timeline, root cause, fix applied, prevention. This is the durable
|
||||
record; don't rely on chat history.
|
||||
|
||||
Docs-update checklist: always create the investigation entry. If the
|
||||
root cause was stale/wrong inventory data (a `doc_page`, `config_repo`,
|
||||
or `backend` that didn't match reality — this happened during Week 1
|
||||
kernel work, see the `authentik` backend fix), correct `inventory.yaml`
|
||||
in the same session.
|
||||
@@ -1,36 +0,0 @@
|
||||
---
|
||||
name: lifecycle-activate-node
|
||||
risk_class: config_mutation
|
||||
inputs: [node_name]
|
||||
verification: "homelab service <name> health (if it hosts a service); homelab doctor (if it's a client)"
|
||||
docs_update_checklist: [doc_page_complete]
|
||||
transition: "provisioning -> active"
|
||||
---
|
||||
|
||||
# Lifecycle: activate a node
|
||||
|
||||
Per [oikos/ontology.yaml](../../../oikos/ontology.yaml). Requires: age key
|
||||
enrolled if it needs secrets, mesh joined if it needs off-LAN reach,
|
||||
ingress live if public, health check answering, doc page complete,
|
||||
ledger entry.
|
||||
|
||||
1. If the node is a `homelab` client: finish enrollment per
|
||||
[client-enrollment.md](../client-enrollment/SKILL.md) (`--finalize-pubkey`,
|
||||
mesh join, `homelab doctor` green).
|
||||
2. If it hosts a public service: add the `services:` entry in
|
||||
`inventory.yaml` (backend, url, doc_page, config_repo, risk_notes —
|
||||
see the Week-1 service contract fields) and wire the Caddy route in
|
||||
`dtoro/caddy-conf`.
|
||||
3. Confirm the health check answers: `homelab service <name> health` or
|
||||
a direct `curl`.
|
||||
4. Flip `state: provisioning` → `state: active` (or delete the `state:`
|
||||
field — `active` is the default) in `inventory.yaml`.
|
||||
5. Complete the doc page (stub → full narrative: role, specs, how it's
|
||||
configured, dependencies).
|
||||
6. Record the activation: `oikos/ledger.py append host:<name> activate
|
||||
config_mutation --result ok` (or let the CLI wrapper do this once
|
||||
Week 3's runbook automation lands).
|
||||
|
||||
Regenerate derived data: `python3 mcp/build_host_files.py && python3
|
||||
oikos/gen-topology.py` so `hosts/<name>.yaml`, the topology diagram, and
|
||||
the context card all reflect the new state.
|
||||
@@ -1,35 +0,0 @@
|
||||
---
|
||||
name: lifecycle-deprecate-node
|
||||
risk_class: config_mutation
|
||||
inputs: [node_name, replacement_node_or_reason]
|
||||
verification: "homelab node <name> relations — 'affected by' must be empty before completing"
|
||||
docs_update_checklist: [doc_page_deprecation_note]
|
||||
transition: "active -> deprecated"
|
||||
---
|
||||
|
||||
# Lifecycle: deprecate a node
|
||||
|
||||
Per [oikos/ontology.yaml](../../../oikos/ontology.yaml): a node keeps running
|
||||
but takes no new dependents. **Completion condition: zero remaining
|
||||
inbound `depends-on`/`routes-to` edges** — this is a hard gate, not a
|
||||
suggestion; `oikos/policy.yaml` `lifecycle_overrides.deprecated.refuse`
|
||||
lists `new-inbound-edges` as refused going forward.
|
||||
|
||||
1. Set `state: deprecated` on the node.
|
||||
2. `homelab node <name> relations` — read `affected_by`. Every entry
|
||||
there is something still relying on this node.
|
||||
3. Migrate or retire each dependent one at a time (point its `backend`/
|
||||
`config_repo`/ingress route elsewhere, or deprecate it too if it's
|
||||
being retired alongside).
|
||||
4. Re-run `homelab node <name> relations` after each dependent is moved.
|
||||
The transition to `destroyed` is only safe once `affected_by` is
|
||||
empty — check this every time, don't assume from memory.
|
||||
5. Note the deprecation on the doc page: reason, replacement (if any),
|
||||
date.
|
||||
|
||||
If step 2 shows dependents you didn't expect, stop and investigate
|
||||
before proceeding — that's exactly the kind of drift the Week-3 detector
|
||||
will catch automatically, but until then this manual check is the gate.
|
||||
|
||||
Next (once `affected_by` is empty):
|
||||
[lifecycle-destroy-node.md](../lifecycle-destroy-node/SKILL.md).
|
||||
@@ -1,42 +0,0 @@
|
||||
---
|
||||
name: lifecycle-destroy-node
|
||||
risk_class: destructive
|
||||
inputs: [node_name]
|
||||
verification: "homelab node <name> relations returns unknown-entity; pct list on the backend no longer shows it"
|
||||
docs_update_checklist: [archaeology_entry, containers_index_update]
|
||||
transition: "deprecated -> destroyed"
|
||||
---
|
||||
|
||||
# Lifecycle: destroy a node
|
||||
|
||||
**Destructive.** Requires operator approval + typed confirmation phrase
|
||||
per `oikos/policy.yaml`. Requires (ontology): backups verified, secrets
|
||||
recipients removed + re-keyed, ingress/DNS removed, archaeology entry,
|
||||
ledger entry.
|
||||
|
||||
1. Confirm the node is `deprecated` with zero `affected_by` edges
|
||||
(`homelab node <name> relations`) — do not skip this even if the
|
||||
deprecation runbook was followed recently; state can drift.
|
||||
2. If it's an enrolled client: `homelab client remove <name>` — revokes
|
||||
the age key, re-keys SOPS, removes the inventory entry. This is
|
||||
already destructive-class and confirmed in the CLI.
|
||||
3. Remove any ingress route (Caddy config repo) and DNS record still
|
||||
pointing at it.
|
||||
4. Verify backups of anything on it are retained per policy before the
|
||||
disk goes away (see `backs-up-to`).
|
||||
5. Destroy the LXC/VM (`pct destroy` / `qm destroy`).
|
||||
6. Move the `hosts.<name>:` block (if any inventory remnant survives
|
||||
`client remove`, e.g. infra-only LXCs with no age key) into
|
||||
inventory.yaml's `archaeology:` section: `pve_id`, `destroyed` date,
|
||||
`reason`. Add a row to `containers/index.md` "Recently destroyed"
|
||||
table (kept for human-readable browsing alongside the structured
|
||||
data).
|
||||
7. `oikos/ledger.py append host:<name> destroy destructive --result ok`.
|
||||
8. Regenerate: `python3 mcp/build_host_files.py && python3
|
||||
oikos/gen-topology.py` — the node drops out of `hosts/*.yaml` and
|
||||
appears in the topology doc's archaeology table.
|
||||
|
||||
If the destroy fails partway (e.g. secrets revoked but pct destroy
|
||||
errors), do not re-run step 2 — `client remove` is not idempotent
|
||||
against a second revocation attempt on the issuance server. Finish the
|
||||
remaining steps manually and note the partial state in an investigation.
|
||||
@@ -1,39 +0,0 @@
|
||||
---
|
||||
name: lifecycle-migrate-node
|
||||
risk_class: config_mutation
|
||||
inputs: [node_name, source_host, target_host]
|
||||
verification: "homelab node <name> relations (re-check blast radius); homelab service <svc> health for every hosted service"
|
||||
docs_update_checklist: [doc_page_migration_note, inventory_host_and_lan_ip]
|
||||
transition: "active -> migrating -> active"
|
||||
---
|
||||
|
||||
# Lifecycle: migrate a node
|
||||
|
||||
Modeled on the strong Phase 1+2 migration
|
||||
([plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)).
|
||||
Requires (ontology): preflight + backup-verified before migrating;
|
||||
post-verify + Caddy backends checked + mounts checked + docs updated
|
||||
before returning to `active`.
|
||||
|
||||
1. `homelab change preflight <every service the node hosts>` — capture
|
||||
current health as a baseline.
|
||||
2. Verify backups are current for anything with data at rest on the
|
||||
node (see `backs-up-to` edges once populated).
|
||||
3. Set `state: migrating` in `inventory.yaml`.
|
||||
4. Perform the migration (pct/qm move, or create-on-target +
|
||||
data-copy + destroy-source, per the specific case).
|
||||
5. Update `inventory.yaml`: new `host:`, `lan_ip`, `mesh` addresses for
|
||||
the node; update every `services:` entry whose `backend` pointed at
|
||||
it if the backend name itself changes (usually it doesn't — only the
|
||||
`host:`/`lan_ip` on the guest entry moves).
|
||||
6. Post-verify: re-run the Week-1 drift check by hand — confirm Caddy's
|
||||
backend IP for each affected service matches the new `lan_ip`
|
||||
(automatic in Week 3's drift detector), confirm mounts still resolve.
|
||||
7. `homelab service <name> health` for every service the node hosts.
|
||||
8. Set `state: active`. Add a migration note to the node's doc page
|
||||
(old host/IP → new, date, phase reference) — this repo's convention
|
||||
for every past migration (see `containers/101-jellyfin.md`,
|
||||
`containers/129-house.md`).
|
||||
|
||||
Regenerate: `python3 mcp/build_host_files.py && python3
|
||||
oikos/gen-topology.py`.
|
||||
@@ -1,33 +0,0 @@
|
||||
---
|
||||
name: lifecycle-provision-node
|
||||
risk_class: config_mutation
|
||||
inputs: [node_name, kind, storage_pool]
|
||||
verification: "grep 'state: provisioning' hosts/<name>.yaml"
|
||||
docs_update_checklist: [doc_page_stub]
|
||||
transition: "planned -> provisioning"
|
||||
---
|
||||
|
||||
# Lifecycle: provision a node
|
||||
|
||||
Per [oikos/ontology.yaml](../../../oikos/ontology.yaml) `lifecycle.transitions`.
|
||||
Policy note: `provisioning` nodes get a lifecycle override —
|
||||
`config_mutation` actions downgrade to `reversible_low` because nothing
|
||||
depends on the node yet (see `oikos/policy.yaml` `lifecycle_overrides`).
|
||||
|
||||
Requires (from ontology): inventory entry, IP reserved, storage pool
|
||||
chosen, doc page stub.
|
||||
|
||||
1. Create the LXC/VM on its target Proxmox host (`pct create` /
|
||||
`qm create`), choosing the storage pool deliberately — record it as
|
||||
the `storage:` field once populated (Week 1 schema; not yet backfilled
|
||||
for existing nodes).
|
||||
2. Add the inventory entry: `homelab client add <name>` for anything that
|
||||
will run the `homelab` CLI, or a direct `hosts.<name>:` block with
|
||||
`state: provisioning`, `kind`, `host`, `pve_id`, `lan_ip` for
|
||||
infra-only LXCs that won't self-enroll.
|
||||
3. Stub the doc page (`containers/<pve_id>-<name>.md` or
|
||||
`vms/<pve_id>-<name>.md`) — even a one-line "provisioning, see plan X"
|
||||
is enough to satisfy the transition requirement.
|
||||
4. Reserve the IP in DNS/DHCP notes if it's a fixed LAN address.
|
||||
|
||||
Next: [lifecycle-activate-node.md](../lifecycle-activate-node/SKILL.md).
|
||||
@@ -1,30 +0,0 @@
|
||||
---
|
||||
name: service-health-check
|
||||
risk_class: read_only
|
||||
inputs: [service_name]
|
||||
verification: "homelab service <name> health"
|
||||
docs_update_checklist: []
|
||||
---
|
||||
|
||||
# Service health check
|
||||
|
||||
Goal: determine whether a service is actually healthy, without ad-hoc SSH.
|
||||
|
||||
1. `homelab service <name> explain` — read the context card: backend,
|
||||
blast radius, doc pointer, risk notes.
|
||||
2. `homelab service <name> health` — live health probe (HTTP code against
|
||||
the service's `url`/`endpoint`). Once the Week-3 scheduler ships, this
|
||||
reads a cached snapshot by default; pass `--live` to force a fresh probe.
|
||||
3. If unhealthy, `homelab service <name> log` (or MCP `tail_log`) for the
|
||||
last 200 lines.
|
||||
4. Cross-check blast radius: `homelab node <name> relations` — is this
|
||||
entity's own backend host healthy? A downstream failure (e.g. `strong`
|
||||
down) will show up here before the service's own logs explain anything.
|
||||
5. If the fix is a restart: classify first (`oikos/policy.yaml` —
|
||||
`service-restart` is `reversible_low` unless the service has a
|
||||
`service_overrides` entry, e.g. `caddy`/`dns` are `config_mutation`).
|
||||
Unattended agents may act on `reversible_low` without approval.
|
||||
|
||||
Docs-update checklist: none for a pure health check. If the investigation
|
||||
reveals stale `risk_notes` or a wrong `doc_page`, fix `inventory.yaml` in
|
||||
the same session.
|
||||
@@ -1,72 +0,0 @@
|
||||
# Oikos CI (Gitea Actions). Gates the deploy webhook on a green run (plan M1).
|
||||
# Mirrors `make lint`, `make test`, and the generated-code drift guard.
|
||||
name: ci
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
|
||||
jobs:
|
||||
build-test:
|
||||
runs-on: ubuntu-latest
|
||||
services:
|
||||
postgres:
|
||||
image: timescale/timescaledb:2.17.2-pg16
|
||||
env:
|
||||
POSTGRES_DB: oikos
|
||||
POSTGRES_USER: oikos
|
||||
POSTGRES_PASSWORD: oikos_dev
|
||||
ports:
|
||||
- 5432:5432
|
||||
options: >-
|
||||
--health-cmd "pg_isready -U oikos"
|
||||
--health-interval 5s
|
||||
--health-timeout 5s
|
||||
--health-retries 10
|
||||
env:
|
||||
OIKOS_TEST_DATABASE_URL: postgres://oikos:oikos_dev@postgres:5432/oikos?sslmode=disable
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version: "1.26"
|
||||
cache: true
|
||||
|
||||
- name: go vet
|
||||
run: go vet ./...
|
||||
|
||||
- name: golangci-lint
|
||||
uses: golangci/golangci-lint-action@v6
|
||||
with:
|
||||
version: latest
|
||||
args: --timeout 5m
|
||||
continue-on-error: true # advisory until the lint baseline is clean
|
||||
|
||||
- name: govulncheck
|
||||
run: |
|
||||
go install golang.org/x/vuln/cmd/govulncheck@latest
|
||||
govulncheck ./... || true # advisory
|
||||
|
||||
- name: generated code is up to date
|
||||
run: make generate-check
|
||||
|
||||
- name: build
|
||||
run: go build ./...
|
||||
|
||||
- name: test (race + coverage)
|
||||
run: go test -race -covermode=atomic -coverprofile=coverage.out -timeout 300s ./...
|
||||
|
||||
- name: coverage gates (policy + learning ≥ 80%, others ≥ 60%)
|
||||
run: |
|
||||
go tool cover -func=coverage.out | tail -1
|
||||
# Note: policy/ and learning/ packages land in Phase 3; enforce
|
||||
# their 80% gate then. For now, report total coverage.
|
||||
|
||||
docker-build:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- name: docker build (verify image builds; no push)
|
||||
run: docker build -f compose/oikos/Dockerfile -t oikos:ci .
|
||||
9
.gitignore
vendored
9
.gitignore
vendored
@@ -1,10 +1 @@
|
||||
.DS_Store
|
||||
__pycache__/
|
||||
*.pyc
|
||||
|
||||
# Regenerated every scheduler run (every 10 min); no audit value in the
|
||||
# diff. Signals (signals/*.jsonl) ARE tracked — this is just the ephemeral
|
||||
# health-probe cache. See oikos/scheduler.py.
|
||||
oikos/state.json
|
||||
|
||||
.worktrees/
|
||||
@@ -12,7 +12,7 @@
|
||||
|
||||
## Problem statement
|
||||
|
||||
The Technitium DHCP server on [CT 107](../../knowledge/wiki/containers/107-dns.md) serves `192.168.8.100–192.168.8.240`. **Every static homelab IP except hubris (`.77`) sits inside that range:**
|
||||
The Technitium DHCP server on [CT 107](containers/107-dns.md) serves `192.168.8.100–192.168.8.240`. **Every static homelab IP except hubris (`.77`) sits inside that range:**
|
||||
|
||||
| Host | IP | Inside pool? |
|
||||
|---|---|---|
|
||||
|
||||
@@ -1,358 +0,0 @@
|
||||
# Assessment: Which nodes can move to `strong`
|
||||
|
||||
## Executive summary
|
||||
|
||||
hubris is **memory-starved**: 28 GiB RAM, 71.9 GiB allocated across 18 LXC + 2 VM
|
||||
(2.5× overcommit), 10 GiB swap in active use. strong sits **completely empty** —
|
||||
28 GiB RAM, 25 GiB free, 0 guests, 2.7 TiB unused storage. The single most
|
||||
effective decongestion move is to shift guests off hubris onto strong.
|
||||
|
||||
This document assesses every guest for move-readiness, grouped by constraints
|
||||
(library dependency, GPU, core-infra status), and proposes a phased migration
|
||||
that does **not** require the physical library-SSD move (the blocker of the
|
||||
original plan) — library access from strong is provided via NFS from hubris.
|
||||
|
||||
---
|
||||
|
||||
## Current resource state (live, 2026-07-05)
|
||||
|
||||
### hubris — overloaded
|
||||
|
||||
| Resource | Capacity | Allocated (all guests) | Actual use | Status |
|
||||
|----------|----------|----------------------|------------|--------|
|
||||
| RAM | 28 GiB | 70.2 GiB (2.5× overcommit) | 18 GiB used + 10 GiB swap | ⚠️ heavy swap pressure |
|
||||
| CPU | 12 vCPU (6c/12t) | 43 vCPU (3.6× overcommit) | ~43% scaling MHz | OK (shares) |
|
||||
| local-lvm | 856 GiB | 582 GiB allocated (29% thin) | — | OK |
|
||||
| library (lvmthin) | 3.7 TiB | — | 1.2 TiB used (34%) | OK, 2.3 TiB free |
|
||||
|
||||
### strong — empty, ready
|
||||
|
||||
| Resource | Capacity | Used | Status |
|
||||
|----------|----------|------|--------|
|
||||
| RAM | 28 GiB | 2.3 GiB (host only) | 25 GiB free |
|
||||
| CPU | 16 vCPU (8c/16t) | idle (0.08 load) | 100% free |
|
||||
| local-lvm | 856 GiB | 0 | empty |
|
||||
| ludo-lvm | 1.8 TiB | 0 | empty |
|
||||
| Guests | — | 0 LXC, 0 VM | nothing running |
|
||||
|
||||
### Network topology constraint
|
||||
|
||||
```
|
||||
Fritz!Box (192.168.178.1)
|
||||
└── SODOLA 2.5G switch
|
||||
├── hubris eno1 → vmbr1 (192.168.178.10) → vmbr0 (192.168.8.0/24)
|
||||
│ └── all 20 guests on 192.168.8.x
|
||||
└── strong vmbr0 (192.168.178.181)
|
||||
└── no internal bridge yet, guests would be on 192.168.178.x
|
||||
```
|
||||
|
||||
strong reaches `192.168.8.0/24` via the Fritz static route through hubris.
|
||||
Guests on strong get `192.168.178.x` IPs unless we add an internal bridge
|
||||
on strong (Phase 0 prerequisite — see below).
|
||||
|
||||
---
|
||||
|
||||
## Per-guest assessment
|
||||
|
||||
### Tier 1 — Move immediately (no library dependency, no core-infra)
|
||||
|
||||
These guests mount **no** `/mnt/library` and are not part of the core
|
||||
infrastructure spine (caddy/dns/auth/mcp). They are the easiest wins.
|
||||
|
||||
| ID | Name | Cores | RAM | Library? | GPU? | Notes |
|
||||
|----|------|-------|-----|----------|------|-------|
|
||||
| 118 | elementsynapse | 2 | 4 GiB | ❌ | ❌ | Matrix homeserver. Public via Caddy (`matrix.hubris.network`). Only change: Caddy backend IP. **Easiest move in the fleet.** |
|
||||
| 129 | house | 2 | 3 GiB | ❌ | ❌ | Yuvomi family planner (Docker). No public Caddy route yet (uses VPS traefik directly). Self-contained. |
|
||||
|
||||
**Combined RAM freed from hubris: 7 GiB.** No NFS, no library, no GPU.
|
||||
|
||||
### Tier 2 — Move with library NFS (high resource consumers)
|
||||
|
||||
These are the heaviest guests and the original migration plan's primary
|
||||
targets. They mount `/mnt/library` and two use the iGPU. Moving them
|
||||
requires an NFS export from hubris → strong (reverse of the original
|
||||
plan's direction, since the physical SSD hasn't moved).
|
||||
|
||||
| ID | Name | Cores | RAM | Library? | GPU? | I/O profile | Notes |
|
||||
|----|------|-------|-----|----------|------|-------------|-------|
|
||||
| 120 | mule-images | 6 | 12 GiB | ✅ mp0 | ✅ iGPU | Write-heavy (photo processing) | **#1 RAM consumer.** Strong has Radeon 680M iGPU (VAAPI works). |
|
||||
| 122 | arriman | 4 | 8 GiB | ✅ mp0 | ❌ | Write-heavy (downloads) | *arr stack + qbit + sab. Mounts library for download writes. |
|
||||
| 101 | jellyfin | 4 | 8 GiB | ✅ mp0 | ✅ iGPU | Read-heavy sequential | Media streaming + transcode. Strong 680M handles VAAPI. |
|
||||
|
||||
**Combined RAM freed: 28 GiB.** This alone would eliminate hubris's swap
|
||||
pressure entirely.
|
||||
|
||||
### Tier 3 — Could move, low urgency
|
||||
|
||||
| ID | Name | Cores | RAM | Library? | Notes |
|
||||
|----|------|-------|-----|----------|-------|
|
||||
| 130 | grimmory | 1 | 2 GiB | ✅ mp0 | Book library (Docker). Migrated from apps LXC recently. |
|
||||
| 131 | teddycloud | 1 | 1 GiB | ✅ mp0 | New (not in inventory.yaml yet). |
|
||||
| 132 | rclone | 1 | 2 GiB | ✅ mp0 (ro) | Backup container. Read-only library mount. |
|
||||
| 128 | trmnl | 1 | 768 MiB | ❌ | TRMNL middleware. No library. Could move but tiny. |
|
||||
| 119 | sophia | 2 | 1 GiB | ✅ mp0 | Workshop. Light use. |
|
||||
|
||||
### Stay on hubris (core infrastructure)
|
||||
|
||||
| ID | Name | Cores | RAM | Why it stays |
|
||||
|----|------|-------|-----|--------------|
|
||||
| 121 | caddy | 1 | 512 MiB | Reverse proxy — terminates all `*.hubris.network`. Must stay on hubris for LAN-side reachability. **Needs backend IP updates** when guests move. |
|
||||
| 107 | dns | 1 | 1 GiB | Technitium DNS, split-horizon. Core. |
|
||||
| 106 | auth-outpost | 1 | 512 MiB | Authentik SSO enforcement. Core. |
|
||||
| 105 | apps | 2 | 4 GiB | homelab MCP + secrets-issuance + artifacto. Core infra. Mounts library. |
|
||||
| 104 | gitea | 1 | 1 GiB | Git server. NFS would hurt git lock/stat perf. Mounts library (bare repos). |
|
||||
| 103 | paperless | 2 | 3 GiB | Document archive. Moderate I/O, OCR writes. Mounts library. |
|
||||
| 102 | nfs-export | 1 | 512 MiB | Exports library to zimaos via NFS. Must stay with the physical library. |
|
||||
| 100 | zimaos | 4 | 8 GiB (VM) | NAS frontend eval. Already NFS-mounts library from 102. |
|
||||
| 108 | haos | 2 | 4 GiB (VM) | Home Assistant OS. Hardware access, low latency. |
|
||||
|
||||
---
|
||||
|
||||
## Constraints & prerequisites
|
||||
|
||||
### 1. Network — strong needs an internal bridge (Phase 0)
|
||||
|
||||
strong currently has only `vmbr0` on `192.168.178.0/24`. Guests created there
|
||||
get household-LAN IPs, not homelab-subnet IPs. Two options:
|
||||
|
||||
- **Option A (recommended):** Add `vmbr1` on strong as a portless internal
|
||||
bridge with a `192.168.8.x/24` address (e.g. `192.168.8.3`). Route between
|
||||
strong's `vmbr0` and `vmbr1` the same way hubris does. Guests go on `vmbr1`
|
||||
and get `192.168.8.x` IPs — transparent to Caddy, DNS, and inter-LXC refs.
|
||||
Requires adding a static route on Fritz (or relying on hubris's existing
|
||||
route — strong would need IP forwarding + a route to 192.168.8.0/24 via vmbr1).
|
||||
|
||||
- **Option B (simpler, messier):** Put guests on `192.168.178.x` directly.
|
||||
Caddy can still reach them (hubris routes to 192.168.178.0/24). But DNS
|
||||
records, inter-LXC references, and firewall rules all assume `192.168.8.x`.
|
||||
More config churn per guest.
|
||||
|
||||
### 2. Storage — rootfs migration (no shared storage)
|
||||
|
||||
`local-lvm` is per-node (not shared). Moving an LXC requires either:
|
||||
- `vzdump` → restore on strong (clean, but needs temp disk space + downtime)
|
||||
- `rsync` the rootfs to a new LXC on strong (faster for large rootfs like 120's 100G)
|
||||
- `pct migrate` only works with shared storage — **not applicable here**
|
||||
|
||||
For VMs (100, 108): `qm migrate` also needs shared storage. Not moving VMs.
|
||||
|
||||
### 3. Library access — NFS from hubris to strong
|
||||
|
||||
Since the physical library SSD is still on hubris, strong's guests that need
|
||||
`/mnt/library` must NFS-mount it from hubris. Options:
|
||||
|
||||
- **Export from hubris host directly** (simplest): add `/mnt/library` to
|
||||
`/etc/exports` on hubris with the same squash params as LXC 102
|
||||
(`rw,all_squash,anonuid=33,anongid=10000,no_subtree_check`). Mount on strong
|
||||
at `/mnt/library`. Strong's guests bind-mount it just like hubris's guests do.
|
||||
|
||||
- **Use existing nfs-export LXC 102**: strong NFS-mounts from `192.168.8.200`
|
||||
(LXC 102). This already has the right squash config. Less host-level change.
|
||||
**This is the path of least resistance.**
|
||||
|
||||
### 4. GPU — iGPU passthrough on strong
|
||||
|
||||
strong has a Ryzen 7 PRO 6850U with Radeon 680M iGPU. For jellyfin (VAAPI
|
||||
transcoding) and mule-images (photo processing), we need:
|
||||
- `/dev/dri/renderD128` passed to the LXC (`lxc.cgroup2.devices.allow` +
|
||||
`lxc.mount.entry` or Proxmox's `dev0:` passthrough)
|
||||
- `video` / `render` group membership inside the container
|
||||
- Confirm `amdgpu` driver loads on strong's host kernel (it should — same APU family)
|
||||
|
||||
### 5. Quorum — 2-node cluster, no QDevice
|
||||
|
||||
Moving guests to strong does NOT fix the quorum issue but **reduces blast
|
||||
radius**: if hubris reboots (its known thermal instability), the guests on
|
||||
strong keep running independently. Consider adding a QDevice as a separate
|
||||
follow-up — it's orthogonal to this migration.
|
||||
|
||||
---
|
||||
|
||||
## Revised migration phases
|
||||
|
||||
The original plan's NFS-over-LAN approach has been superseded. Instead,
|
||||
**media library data moves to ludo-lvm** on strong so migrated guests access
|
||||
it as a local ext4 mount. Data is split by origin:
|
||||
|
||||
```
|
||||
hubris (stays): library SSD (3.7T, 1.2T used)
|
||||
└── /mnt/library/{documents,images,cloud,homecloud,notes,repos,sophia}
|
||||
↑ user-generated content (docs, photos, cloud sync, notes, repos, workshop)
|
||||
|
||||
strong (moves): ludo-lvm (1.8T, 0 used at start)
|
||||
└── /mnt/media_local ← 1.5T thin volume
|
||||
└── {downloads,movies,music,tv,anime,books}
|
||||
↑ non-user-generated content (media arr stack, book library)
|
||||
```
|
||||
|
||||
| Category | Stays on hubris | Moves to strong |
|
||||
|----------|----------------|-----------------|
|
||||
| Media | — | downloads (25G), movies (51G), music (29G), tv (30G), anime (206G) |
|
||||
| Books | — | books (2.6G) |
|
||||
| Docs/Photos | documents (249M), images (4K) | — |
|
||||
| Cloud sync | cloud (287G), homecloud (367G) | — |
|
||||
| Personal | notes (6.7M), repos (84M), sophia (151G) | — |
|
||||
| **Total** | **~805G** | **~344G** |
|
||||
|
||||
ludo-lvm (1.8T) fits all media + books with ~1.15T headroom for growth.
|
||||
hubris library SSD (3.7T, 1.2T used) retains the user-generated content.
|
||||
Both sides keep their data local — no cross-node NFS needed for daily I/O.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2a — Prepare ludo-lvm on strong
|
||||
|
||||
1. Create a ext4 filesystem on ludo-lvm for media:
|
||||
```bash
|
||||
lvcreate -n media -L 1.5T ludo-lvm
|
||||
mkfs.ext4 /dev/ludo-lvm/media
|
||||
```
|
||||
2. Mount at `/mnt/media_local` on strong, add to `/etc/fstab`
|
||||
3. rsync media directories from hubris → strong:
|
||||
```bash
|
||||
rsync -av --progress /mnt/library/{movies,tv,anime,downloads,music,books} strong:/mnt/media_local/
|
||||
```
|
||||
|
||||
### Phase 2b — Migrate arriman (122) to strong
|
||||
|
||||
1. Stop arriman on hubris, dump rootfs (24G)
|
||||
2. Restore on strong with IP `192.168.8.245/28` on vmbr1
|
||||
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (downloads land locally)
|
||||
4. Update Caddy: jellyseerr/qbit/sab backends → new IP
|
||||
5. Update inventory.yaml
|
||||
|
||||
### Phase 2c — Migrate jellyfin (101) to strong
|
||||
|
||||
1. Stop jellyfin on hubris, dump rootfs (16G)
|
||||
2. Restore on strong with IP `192.168.8.246/28` on vmbr1
|
||||
3. Pass `/dev/dri/renderD128` + `/dev/dri/card0` (Radeon 680M + RX 7600)
|
||||
4. Mount `/mnt/media_local` → `/mnt/library` via mp0 (media reads locally)
|
||||
5. Update Caddy: `media.hubris.network` → new IP
|
||||
6. Reinstall `sso-inject.js` in web dir (lost on every apt upgrade)
|
||||
7. Test VAAPI transcoding, SSO login, media playback
|
||||
|
||||
### Phase 2d — Migrate grimmory (130) to strong
|
||||
|
||||
1. Stop grimmory on hubris, dump rootfs (16G)
|
||||
2. Restore on strong with IP `192.168.8.247/28` on vmbr1
|
||||
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (books read locally)
|
||||
4. Update Caddy: `books.hubris.network` → new IP
|
||||
5. Update inventory.yaml
|
||||
6. Test: book browsing, calibre-web access
|
||||
|
||||
### No NFS export needed
|
||||
|
||||
With the data split by origin, hubris guests that only need user-generated
|
||||
content (documents, images, cloud, repos, sophia) still access them from the
|
||||
original library SSD — no cross-node NFS required. The two sides are
|
||||
independent.
|
||||
|
||||
**Result after Phase 2: hubris frees 26 GiB RAM (4 migrated guests) + 344G of
|
||||
library I/O burden. Strong becomes the media/books powerhouse.**
|
||||
|
||||
---
|
||||
|
||||
### Phase 3 — Migrate mule-images (120) to strong
|
||||
|
||||
Move photo management (12 GiB RAM, 6 cores, iGPU) last because it needs:
|
||||
- `/mnt/library` access (now NFS from strong — already set up in Phase 2d)
|
||||
- `/dev/dri/renderD128` (Radeon 680M — confirm VAAPI compatibility first)
|
||||
|
||||
Steps:
|
||||
1. Stop mule-images on hubris, rsync the 100G rootfs to strong (faster than vzdump)
|
||||
2. Restore on strong with IP on vmbr1
|
||||
3. Pass Radeon 680M iGPU
|
||||
4. Reconfigure library paths → `/mnt/media_local` (or keep NFS mount)
|
||||
5. Update Caddy: `photos.hubris.network` → new IP
|
||||
6. Test photo import + processing pipeline
|
||||
|
||||
---
|
||||
|
||||
### Phase 4 — Tier 3 moves (optional)
|
||||
|
||||
Migrate grimmory (130), teddycloud (131), rclone (132), trmnl (128), sophia (119)
|
||||
as needed — each frees 1–2 GiB. Not urgent; do when convenient.
|
||||
|
||||
---
|
||||
|
||||
### Phase 5 — Follow-up
|
||||
|
||||
- **QDevice**: add a tiebreaker for 2-node quorum
|
||||
- **Gaming VM**: strong's 6850U has enough cores alongside migrated LXCs
|
||||
- **Hubris library cleanup**: after all guests are confirmed working, decide
|
||||
whether to keep the original library SSD as backup or repurpose it
|
||||
|
||||
---
|
||||
|
||||
## Resource math after Phase 3 (all Tier 1 + 2 moved)
|
||||
|
||||
| | hubris | strong |
|
||||
|---|--------|--------|
|
||||
| Guests | 11 LXC + 2 VM | 5 LXC |
|
||||
| RAM allocated | ~25 GiB | ~45 GiB |
|
||||
| RAM capacity | 28 GiB | 28 GiB |
|
||||
| Overcommit | 0.9× (under-committed) | 1.6× (manageable) |
|
||||
| Library disk | Local ext4 (3.7T) → NFS client | Local ext4 on ludo-lvm (1.8T) |
|
||||
| GPU | Radeon 760M (idle) | Radeon 680M (jellyfin + mule-images) |
|
||||
|
||||
strong becomes the media/library powerhouse. hubris becomes a lean core-infra
|
||||
node (DNS, auth, git, docs, caddy, HA).
|
||||
|
||||
---
|
||||
|
||||
---
|
||||
|
||||
## Risk register
|
||||
|
||||
| Risk | Impact | Mitigation |
|
||||
|------|--------|------------|
|
||||
| NFS latency for library reads (jellyfin, arriman) | Media playback stutter, slow downloads | Test iperf between strong↔hubris first. If 2.5G link, NFS throughput is fine (~1 Gbit/s). |
|
||||
| GPU passthrough on strong (680M vs 760M) | Transcode quality/compat differences | Both are AMD VAAPI — same driver stack. Test `vainfo` inside LXC before going live. |
|
||||
| Caddy backend IP churn | Service outage if IP wrong | Update Caddyfile in git repo (caddy-conf), test each route before destroying old LXC. |
|
||||
| vzdump/restore downtime | Service unavailable during migration | Schedule off-hours. Use rsync for large rootfs (120's 100G) to minimize freeze window. |
|
||||
| 2-node quorum still fragile | If hubris goes down, strong /etc/pve goes read-only | Guests keep running. Add QDevice as follow-up. |
|
||||
| Library data integrity during NFS transition | Permission drift | NFS `all_squash,anonuid=33,anongid=10000` matches existing LXC 102 config. Verify with `ls -la /mnt/library` after mount. |
|
||||
|
||||
---
|
||||
|
||||
## Open questions for operator
|
||||
|
||||
1. **Internal bridge on strong**: proceed with `vmbr1` on `192.168.8.3/24`
|
||||
(Option A), or use `192.168.178.x` guest IPs (Option B)?
|
||||
2. **Migration method**: `vzdump`/restore (clean, downtime) vs `rsync` rootfs
|
||||
(faster for large disks, needs manual config copy)?
|
||||
3. **Phase 1 priority**: move elementsynapse + house first (quick wins), or
|
||||
go straight to Phase 2 (mule-images/jellyfin/arriman) for maximum relief?
|
||||
4. **Should we add a QDevice now** before moving anything, to protect
|
||||
management plane during the migration?
|
||||
|
||||
---
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-07-05 — Phase 2d complete (grimmory migrated; media NFS to zimaos)
|
||||
grimmory (130) → 192.168.8.247 on strong. Rsync'd /books (2.6G) to ludo-lvm.
|
||||
LXC 102 (nfs-export) now mounts strong's NFS at /mnt/media and exports it as
|
||||
a second share alongside /mnt/library. Zimaos mounts both: /media/library
|
||||
(hubris user-generated) and /media/media (strong media+books).
|
||||
See hosts/strong.md changelog.
|
||||
|
||||
### 2026-07-05 — Phase 2 complete (arriman + jellyfin migrated; library on ludo-lvm)
|
||||
arriman (122) → 192.168.8.245, jellyfin (101) → 192.168.8.246. Created 1.5T
|
||||
thin volume on ludo-lvm, rsync'd 363G of media data. Both containers use local
|
||||
ext4 mount — no NFS. Jellyfin has 680M + RX 7600 GPU passthrough.
|
||||
Caddy backends updated. See hosts/strong.md changelog.
|
||||
|
||||
### 2026-07-05 — Phase 1 complete (elementsynapse + house migrated to strong)
|
||||
Both Tier 1 guests moved: elementsynapse (118) → 192.168.8.242, house (129) → 192.168.8.244.
|
||||
Strong now has vmbr1 at 192.168.8.241/28. Hubris has proxy ARP + /32 routes for strong
|
||||
guest range. DHCP scope narrowed to 192.168.8.100-239 to avoid conflicts.
|
||||
Teddycloud (LXC 131) given static IP 192.168.8.150 due to IP conflict with
|
||||
previous DHCP allocation at 192.168.8.243.
|
||||
See hosts/strong.md changelog for full steps.
|
||||
|
||||
### 2026-07-05 — assessment created
|
||||
Built from live `pct config` + `pvesm status` + `free -h` data pulled from
|
||||
both nodes. Supersedes the storage-migration framing of the original
|
||||
library-SSD plan — this assessment treats the SSD move as optional and
|
||||
focuses on guest relocation via NFS.
|
||||
34
.sops.yaml
34
.sops.yaml
@@ -25,9 +25,7 @@ creation_rules:
|
||||
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
|
||||
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
|
||||
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
|
||||
age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4,
|
||||
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
|
||||
|
||||
- path_regex: ^secrets/gitea-pat\.yaml$
|
||||
# Write-scoped Gitea PAT (dtoro user). Same recipient list as hello.yaml
|
||||
@@ -38,14 +36,12 @@ creation_rules:
|
||||
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
|
||||
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
|
||||
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
|
||||
age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4,
|
||||
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
|
||||
|
||||
- path_regex: ^secrets/gitea-tokens\.yaml$
|
||||
# Workstations only.
|
||||
age: >-
|
||||
# placeholder — fill with age_pubkey of: republic-laptop, mac-mini, strong, hubris
|
||||
# placeholder — fill with age_pubkey of: republic-laptop, mac-mini, ludo-mini, hubris
|
||||
|
||||
- path_regex: ^secrets/webhook-hmacs\.yaml$
|
||||
# LXCs that run a webhook receiver.
|
||||
@@ -62,8 +58,7 @@ creation_rules:
|
||||
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
|
||||
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
|
||||
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
|
||||
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
|
||||
|
||||
- path_regex: ^secrets/netbird-authentik-oidc\.yaml$
|
||||
# Authentik OIDC client secret for the netbird-dashboard provider.
|
||||
@@ -74,8 +69,7 @@ creation_rules:
|
||||
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0,
|
||||
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
|
||||
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h,
|
||||
age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
|
||||
|
||||
- path_regex: ^secrets/netbird-pat\.yaml$
|
||||
# NetBird API Personal Access Token. Consumed by the dns-sync job on the
|
||||
@@ -115,22 +109,4 @@ creation_rules:
|
||||
age1vf8h7s8mqsn2q5eadgpdupsj4mwn8zguc77d85ws3xj40sl9rgksx2rxw6,
|
||||
age1z62ff2ak9zj5ctcvaxwyyhedwjvlwgm2dkn9nk3wrwk8fkavcpmsqwc2vs,
|
||||
age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
|
||||
|
||||
- path_regex: ^secrets/oikos-approval-hmac\.yaml$
|
||||
# HMAC signing key for Oikos approval-grant tokens (oikos/approve.py).
|
||||
# Recipients: apps (105, runs the approval engine alongside homelab-mcp)
|
||||
# and hubris (admin/debug decrypt). See OIKOS.md "Approval engine".
|
||||
age: >-
|
||||
age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6,
|
||||
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0
|
||||
|
||||
- path_regex: ^secrets/oikos-console-deploy-secret\.yaml$
|
||||
# Shared HMAC secret for the Gitea deploy webhook (id 14) ->
|
||||
# oikos-console-deploy.service on apps (105). Generated + registered
|
||||
# with Gitea before the apps-side install ran (see
|
||||
# oikos/console/deploy/README.md "Status") — write this exact value
|
||||
# into /etc/oikos-console-deploy/secret rather than letting
|
||||
# webhook/install.sh generate a fresh one.
|
||||
age: >-
|
||||
age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0
|
||||
# webhook noop 2026-05-20T18:16:57+02:00
|
||||
|
||||
72
AGENTS.md
72
AGENTS.md
@@ -4,20 +4,6 @@ You are running on a machine that is part of the **hubris** homelab. The full
|
||||
context is in this checkout at `/opt/homelab-context/`. This file is the entry
|
||||
point. Read it once at start, then keep working.
|
||||
|
||||
The operating model — OODA loop, risk classes, approval rules, the ontology,
|
||||
and node lifecycle — is defined in [OIKOS.md](.agents/OIKOS.md). Before any mutation,
|
||||
classify the action against `oikos/policy.yaml`; when the class requires
|
||||
approval, stop and ask the operator.
|
||||
|
||||
Agent-facing instruction is separated from human content under `.agents/`:
|
||||
`.agents/shared/` holds the conventions every agent applies
|
||||
([writing-style](.agents/shared/writing-style.md), [caveman](.agents/shared/caveman.md),
|
||||
[page-templates](.agents/shared/page-templates.md), [llm-wiki](.agents/shared/llm-wiki.md)), and
|
||||
`.agents/domains/` holds the per-domain schemas
|
||||
([knowledge](.agents/domains/knowledge/schema.md), [operations](.agents/domains/operations/schema.md)).
|
||||
The narrative wiki lives under `knowledge/wiki/`; the machine-readable substrate
|
||||
(`inventory.yaml`, `hosts/*.yaml`, `oikos/`) stays at the repo root.
|
||||
|
||||
## 1. Who you are
|
||||
|
||||
Run `hostname` (Linux) or `scutil --get LocalHostName` (macOS), then read:
|
||||
@@ -33,14 +19,13 @@ the operator to run `homelab client add <hostname>` from an existing client.
|
||||
- `/opt/homelab-context/inventory.yaml` — every host, LXC, VM, and workstation
|
||||
with their mesh addresses, roles, and service mappings. Treat this file as
|
||||
authoritative; anything you read in narrative pages should agree with it.
|
||||
- `/opt/homelab-context/knowledge/wiki/infrastructure/mesh.md` — Tailscale → Netbird state.
|
||||
- `/opt/homelab-context/infrastructure/mesh.md` — Tailscale → Netbird state.
|
||||
Both meshes are accepted today; Netbird is preferred for new traffic.
|
||||
- `/opt/homelab-context/knowledge/wiki/infrastructure/dns.md` — split-horizon DNS via
|
||||
Technitium on [dns (107)](knowledge/wiki/containers/107-dns.md). `*.hubris.network`
|
||||
resolves to 192.168.x.x on the LAN and to mesh addresses off-LAN.
|
||||
- `/opt/homelab-context/.agents/operations/commands.md` — the operator's cheatsheet
|
||||
for pct, caddy, DNS, and the Oikos command surface. Use these verbs when
|
||||
you take actions.
|
||||
- `/opt/homelab-context/infrastructure/dns.md` — split-horizon DNS via
|
||||
dnsmasq on LXC 124. `*.hubris.network` resolves to 192.168.x.x on the LAN
|
||||
and to mesh addresses off-LAN.
|
||||
- `/opt/homelab-context/operations/commands.md` — the operator's cheatsheet
|
||||
for pct, caddy, dnsmasq. Use these verbs when you take actions.
|
||||
|
||||
## 3. The MCP server
|
||||
|
||||
@@ -59,15 +44,8 @@ Available tools:
|
||||
get_service_status(service), tail_log(service, lines=200),
|
||||
list_lxcs(), get_lxc_state(lxc), ping_service(service)
|
||||
|
||||
Oikos (read-only; see OIKOS.md):
|
||||
explain(service) — compact context card, cheaper than search_docs+get_page
|
||||
preflight(service) — risk class, approval requirement, verification command
|
||||
get_relations(entity) — ontology blast-radius query (host: or service: id)
|
||||
get_change_history(entity, limit=20) — change-ledger entries
|
||||
get_state_snapshot() — last scheduler Observe-pass (health, disk, drift count)
|
||||
|
||||
Mutations are **not** exposed via MCP. Use the `homelab` CLI for those, with
|
||||
operator confirmation — see OIKOS.md's risk classes and approval flow.
|
||||
operator confirmation.
|
||||
|
||||
**When to prefer MCP over grepping the clone:** any time you need to resolve a
|
||||
name to an address, look up service status, or search the wiki by content.
|
||||
@@ -75,25 +53,17 @@ Grep is fine for browsing or when MCP is unreachable.
|
||||
|
||||
## 4. Wiki conventions
|
||||
|
||||
See [page-templates.md](.agents/shared/page-templates.md) for file naming, page
|
||||
structure, and the tone standard. Quick reference:
|
||||
|
||||
- **File naming:** Foundational docs are ALL-CAPS (AGENTS.md, OIKOS.md, GLOSSARY.md);
|
||||
containers use `<id>-<name>.md`; infrastructure uses lowercase-with-dashes; plans
|
||||
and investigations use `YYYY-MM-DD-slug.md`; skills are `<name>/SKILL.md`.
|
||||
- **Where pages live:** Narrative under `knowledge/wiki/{containers,hosts,vms,infrastructure}/`;
|
||||
incident records under `knowledge/sources/investigations/`; runbook procedures under
|
||||
`.agents/skills/<name>/SKILL.md`; operator reference under `.agents/operations/`;
|
||||
design docs under `plans/`. Cross-link liberally; orphans are bugs.
|
||||
- **Changelog format:** Every page ends with a `## Changelog` section, entries in
|
||||
reverse-chronological order:
|
||||
- Pages live under `containers/`, `hosts/`, `vms/`, `infrastructure/`,
|
||||
`investigations/`, `operations/`. Cross-link liberally; orphans are bugs.
|
||||
- Every page ends with a `## Changelog` section, entries in reverse-chrono
|
||||
order:
|
||||
|
||||
### YYYY-MM-DD — short title
|
||||
one or two lines describing what changed and why.
|
||||
|
||||
- **Live state precedence.** If you observe a discrepancy between the docs and
|
||||
running state, update the docs *in the same session* (per the same-session update
|
||||
rule in [page-templates.md](.agents/shared/page-templates.md#same-session-update-rule)).
|
||||
- Investigation files are dated and slugged: `YYYY-MM-DD-slug.md`.
|
||||
- Live state takes precedence over docs. If you observe a discrepancy, update
|
||||
the docs *in the same session* (per the same-session update rule).
|
||||
|
||||
## 5. Acting on the homelab
|
||||
|
||||
@@ -106,20 +76,16 @@ structure, and the tone standard. Quick reference:
|
||||
demand using the per-client age key at `/etc/age/key.txt`. Secrets ARE
|
||||
available in this system — `list_my_secrets()` (MCP) shows what you can
|
||||
decrypt.
|
||||
- **Mutations** (restart, edit configs, etc.): classify against
|
||||
`oikos/policy.yaml` first (`homelab decide <action> <entity>`).
|
||||
`reversible_low` actions just need the interactive confirmation prompt;
|
||||
`config_mutation`/`destructive` actions are mechanically refused without
|
||||
a valid `--approval-id` from `homelab approval request` — see OIKOS.md.
|
||||
For ad-hoc work, SSH and edit directly — but commit changes that touch
|
||||
tracked configs (caddy, gitea custom, artifacto, mule-image, etc.; see
|
||||
`knowledge/wiki/infrastructure/auto-deploy.md`).
|
||||
- **Mutations** (restart, edit configs, etc.): the `homelab` CLI's mutating
|
||||
subcommands ask for confirmation. For ad-hoc work, SSH and edit directly —
|
||||
but commit changes that touch tracked configs (caddy, gitea custom,
|
||||
artifacto, mule-image, etc.; see `infrastructure/auto-deploy.md`).
|
||||
- **Wiki updates**: same-session rule applies to any meaningful state change
|
||||
this client makes.
|
||||
|
||||
## 6. Communication mode
|
||||
|
||||
Read and apply `/opt/homelab-context/.agents/shared/caveman.md` (if present). It defines the lab's
|
||||
Read and apply `/opt/homelab-context/CAVEMAN.md` (if present). It defines the lab's
|
||||
terse-communication standard — drop filler, keep substance, use fragments.
|
||||
|
||||
## 7. Auto-setup mechanism
|
||||
|
||||
123
CONTRIBUTING.md
Normal file
123
CONTRIBUTING.md
Normal file
@@ -0,0 +1,123 @@
|
||||
# Contributing to the Homelab Wiki
|
||||
|
||||
## Voice
|
||||
|
||||
Concise, technical, sysadmin-to-sysadmin. No marketing prose, no exclamation marks.
|
||||
|
||||
## Page templates
|
||||
|
||||
### Container page (`containers/<id>-<name>.md`)
|
||||
|
||||
```markdown
|
||||
# <id> — `<name>`
|
||||
|
||||
One-sentence purpose.
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `<name>`
|
||||
- **IP:** `192.168.8.x`
|
||||
- **Privilege:** privileged | unprivileged
|
||||
- **Resources:** N cores / M GiB RAM / D GiB rootfs
|
||||
- **Mounts:** `/mnt/library` ↔ `/mnt/library` (if any)
|
||||
- **Public hostname:** `<sub>.hubris.network` (if proxied)
|
||||
|
||||
## Role
|
||||
What it does, what it talks to.
|
||||
|
||||
## Service / port map
|
||||
| Service | Listen | Notes |
|
||||
|
||||
## Storage / config paths
|
||||
|
||||
## Auto-deploy
|
||||
(if any) — link to [auto-deploy](../infrastructure/auto-deploy.md)
|
||||
|
||||
## Related
|
||||
- [Caddy](121-caddy.md) (if proxied)
|
||||
- [DNS](../infrastructure/dns.md) (if has subdomain)
|
||||
- [Authentik](124-authentik.md) (if SSO)
|
||||
- ...
|
||||
|
||||
## Changelog
|
||||
### YYYY-MM-DD — short title
|
||||
What changed, why, link to investigation if any.
|
||||
```
|
||||
|
||||
### Cross-cutting page (`infrastructure/<topic>.md`)
|
||||
|
||||
```markdown
|
||||
# <Topic>
|
||||
|
||||
One-sentence summary.
|
||||
|
||||
## Why
|
||||
Design rationale — what it replaces, what it solves.
|
||||
|
||||
## Components
|
||||
Where it runs, what files matter.
|
||||
|
||||
## How to apply / use
|
||||
Recipes.
|
||||
|
||||
## Gotchas
|
||||
|
||||
## Related
|
||||
Links to nodes that host or depend on this.
|
||||
|
||||
## Changelog
|
||||
```
|
||||
|
||||
### Plan (`plans/YYYY-MM-DD-slug.md`)
|
||||
|
||||
```markdown
|
||||
# YYYY-MM-DD — <title>
|
||||
|
||||
## Goal
|
||||
What this change achieves and why.
|
||||
|
||||
## Current topology / state
|
||||
Diagram or description of what exists now.
|
||||
|
||||
## Target topology / state
|
||||
What it looks like after.
|
||||
|
||||
## Pre-flight checklist
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
## Verification
|
||||
|
||||
## Post-migration
|
||||
Changelog entries to write, index status to update.
|
||||
```
|
||||
|
||||
### Investigation (`investigations/YYYY-MM-DD-slug.md`)
|
||||
|
||||
```markdown
|
||||
# YYYY-MM-DD — <title>
|
||||
|
||||
## Summary
|
||||
1-3 sentences.
|
||||
|
||||
## Timeline
|
||||
|
||||
## Root cause
|
||||
|
||||
## Mitigations applied
|
||||
|
||||
## Open questions
|
||||
```
|
||||
|
||||
## Linking discipline
|
||||
|
||||
- Every container page links to every cross-cutting page it participates in.
|
||||
- Every cross-cutting page lists the nodes that participate.
|
||||
- Every investigation links to the nodes it implicates *and* gets back-linked from each node's changelog.
|
||||
- Every plan links to the infrastructure pages it affects. When done, update the plan's status in `plans/index.md` and write changelog entries on affected node pages.
|
||||
|
||||
## Changelog hygiene
|
||||
|
||||
- Reverse-chronological (newest first).
|
||||
- One entry per discrete change, even if you make several in one day.
|
||||
- If a change spans nodes, repeat the entry on each affected page (different perspective is fine).
|
||||
- Don't rewrite history — entries are append-only. Mistakes get a follow-up entry that supersedes them.
|
||||
@@ -15,19 +15,6 @@ truth for:
|
||||
|
||||
When in doubt, check `/opt/homelab-context/` first.
|
||||
|
||||
## Runbooks — load, don't rediscover
|
||||
|
||||
For the canonical workflows (service health check, config change +
|
||||
deploy, client enrollment, incident investigation, and each node
|
||||
lifecycle transition), read the matching `.agents/skills/<name>/SKILL.md` before
|
||||
acting. Each skill carries its risk class, required inputs, the
|
||||
verification command, and a docs-update checklist in its frontmatter —
|
||||
classify against `oikos/policy.yaml` using that risk class before any
|
||||
mutation. Don't re-derive topology or the mutation path by grepping the
|
||||
wiki when a runbook already encodes it. See [OIKOS.md](OIKOS.md) for the
|
||||
operating model these runbooks execute inside (OODA loop, risk classes,
|
||||
approval flow, ontology).
|
||||
|
||||
## Agent type — how this file gets loaded
|
||||
|
||||
| Agent | Loading mechanism |
|
||||
50
Makefile
50
Makefile
@@ -1,50 +0,0 @@
|
||||
.PHONY: build test test-db lint generate generate-check dev migrate seed export clean tidy
|
||||
|
||||
BINARY := oikos
|
||||
GO ?= go
|
||||
|
||||
build:
|
||||
$(GO) build -o $(BINARY) -tags timetzdata ./cmd/oikos
|
||||
|
||||
test:
|
||||
$(GO) test -race -cover ./...
|
||||
|
||||
# Integration tests against the compose Postgres (starts it if needed)
|
||||
test-db:
|
||||
docker compose up -d postgres
|
||||
@sleep 3
|
||||
OIKOS_TEST_DATABASE_URL="postgres://oikos:$${OIKOS_DB_PASSWORD:-oikos_dev}@localhost:5432/oikos?sslmode=disable" \
|
||||
$(GO) test -race -count=1 ./internal/db/ ./internal/httpapi/ ./internal/mcp/
|
||||
|
||||
lint:
|
||||
$(GO) vet ./...
|
||||
@command -v golangci-lint >/dev/null 2>&1 && golangci-lint run || echo "golangci-lint not installed, skipping"
|
||||
|
||||
generate:
|
||||
$(GO) run github.com/oapi-codegen/oapi-codegen/v2/cmd/oapi-codegen@v2.4.1 \
|
||||
-config api/codegen.yaml api/openapi.yaml
|
||||
$(GO) run github.com/sqlc-dev/sqlc/cmd/sqlc@v1.29.0 generate
|
||||
|
||||
# CI drift guard: regenerate and fail if the committed output changed.
|
||||
generate-check: generate
|
||||
@git diff --exit-code -- internal/httpapi/gen internal/db/sqlcgen \
|
||||
|| (echo "generated code is stale — run 'make generate' and commit" && exit 1)
|
||||
|
||||
migrate:
|
||||
$(GO) run ./cmd/oikos migrate
|
||||
|
||||
seed:
|
||||
$(GO) run ./cmd/oikos seed
|
||||
|
||||
export:
|
||||
$(GO) run ./cmd/oikos export
|
||||
|
||||
dev:
|
||||
docker compose --profile dev up -d
|
||||
|
||||
clean:
|
||||
rm -f $(BINARY)
|
||||
$(GO) clean -testcache
|
||||
|
||||
tidy:
|
||||
$(GO) mod tidy
|
||||
232
README.md
232
README.md
@@ -1,190 +1,76 @@
|
||||
# Homelab OS
|
||||
# Homelab Wiki — `hubris`
|
||||
|
||||
Living documentation for the **hubris** Proxmox homelab + Oikos operating system.
|
||||
Living documentation for the **hubris** Proxmox homelab. Every node, every cross-cutting system, and every meaningful incident is its own page; pages are linked so you can start anywhere and walk the graph.
|
||||
|
||||
**For agents running on enrolled clients:** start with [AGENTS.md](AGENTS.md), then [OIKOS.md](.agents/OIKOS.md).
|
||||
> Last refreshed against live state: **2026-04-28**.
|
||||
|
||||
---
|
||||
## Map
|
||||
|
||||
## For Agents — Navigation & Entry Points
|
||||
### Hosts
|
||||
- [`hubris`](hosts/hubris.md) — single Proxmox VE node, GMKtec NucBox M6 Ultra, `192.168.8.77`
|
||||
|
||||
### You are running on a client enrolled in the hubris homelab
|
||||
### VMs
|
||||
- [100 — `zimaos`](vms/100-zimaos.md) — ZimaOS 1.6.1, NAS frontend (evaluation)
|
||||
- [108 — `haos-16.3`](vms/108-haos.md) — Home Assistant OS
|
||||
|
||||
1. **First:** Read [AGENTS.md](AGENTS.md) once. It explains who you are, the topology, available tools, conventions, and how to act.
|
||||
2. **Before any mutation:** Read [OIKOS.md](.agents/OIKOS.md). It defines the operating model, risk classes, approval flow, and the ontology you'll consult.
|
||||
3. **For specific workflows:** Load the matching skill from `.agents/skills/<name>/SKILL.md` (e.g., [service-health-check](.agents/skills/service-health-check/SKILL.md)).
|
||||
4. **When in doubt:** Use MCP tools (`search_docs`, `get_page`, `explain`, `get_changelog`) — they're cheaper and more reliable than grepping.
|
||||
### LXC containers
|
||||
See the full table in [`containers/index.md`](containers/index.md). Quick links:
|
||||
|
||||
### Key References for Agents
|
||||
| ID | Name | IP | Role |
|
||||
| --- | ---------------- | --------------- | --------------------------------------------- |
|
||||
| 101 | [jellyfin](containers/101-jellyfin.md) | 192.168.8.206 | Media server |
|
||||
| 102 | [nfs-export](containers/102-nfs-export.md) | 192.168.8.200 | NFSv4 re-export of /mnt/library for ZimaOS |
|
||||
| 103 | [paperless](containers/103-paperless.md) | 192.168.8.130 | Document mgmt |
|
||||
| 104 | [gitea](containers/104-gitea.md) | 192.168.8.121 | Git server |
|
||||
| 105 | [apps](containers/105-apps.md) | 192.168.8.205 | Docker host (Artifacto / PlantUML / Portainer / WriteFreely) |
|
||||
| 114 | [nextcloud](containers/114-nextcloud.md) | 192.168.8.224 | Personal cloud |
|
||||
| 118 | [elementsynapse](containers/118-elementsynapse.md) | 192.168.8.239 | Matrix Synapse |
|
||||
| 119 | [sophia](containers/119-sophia.md) | 192.168.8.157 | Sophia |
|
||||
| 120 | [mule-images](containers/120-mule-images.md) | 192.168.8.136 | Mule-image / mulita photos |
|
||||
| 121 | [caddy](containers/121-caddy.md) | 192.168.8.175 | Reverse proxy |
|
||||
| 122 | [arriman](containers/122-arriman.md) | 192.168.8.132 | Docker host (\*arr stack) |
|
||||
| 124 | [authentik](containers/124-authentik.md) | 192.168.8.180 | SSO + split-horizon DNS |
|
||||
| 130 | [grimmory](containers/130-grimmory.md) | 192.168.8.213 | Digital library (Grimmory — fork of Booklore) |
|
||||
|
||||
- **What am I?** → `/opt/homelab-context/hosts/<hostname>.yaml` (read on first run)
|
||||
- **Live topology** → `inventory.yaml` + `hosts/*.yaml` (canonical, always wins)
|
||||
- **Risk & approval** → [oikos/policy.yaml](oikos/policy.yaml) (enforced, not advisory)
|
||||
- **Runbooks & workflows** → [.agents/skills/](.agents/skills/) (risk class + verification checklist included)
|
||||
- **State of Oikos** → [OIKOS.md build status](.agents/OIKOS.md#build-status-30-day-roadmap) (scheduled probes, drift detectors, signals, approval engine)
|
||||
### Cross-cutting infrastructure
|
||||
- [DNS — split-horizon](infrastructure/dns.md)
|
||||
- [Ingress — Caddy + VPS traefik](infrastructure/ingress.md)
|
||||
- [Mesh — Tailscale → Netbird migration](infrastructure/mesh.md)
|
||||
- [Monitoring — Hermes health watchdog](infrastructure/monitoring.md)
|
||||
- [Media permissions — `media` GID 10000](infrastructure/media-permissions.md)
|
||||
- [SSH access](infrastructure/ssh-access.md)
|
||||
- [Backups — restic on external drive (disabled)](infrastructure/backups.md)
|
||||
- [Auto-deploy — gitea-webhook pipelines](infrastructure/auto-deploy.md)
|
||||
- [VPS hardening — IONOS / netbird control plane](infrastructure/vps-hardening.md)
|
||||
- [Homelab context distribution](infrastructure/homelab-context.md) — cross-client `/opt/homelab-context` + MCP + secrets-issuance
|
||||
|
||||
### When to Use MCP vs Files vs Shell
|
||||
### Investigations
|
||||
Time-stamped incident notes / experiments in [`investigations/`](investigations/index.md).
|
||||
|
||||
| Task | Use | Tool |
|
||||
|------|-----|------|
|
||||
| Resolve hostname → address | MCP | `get_host(name)` or `list_services()` |
|
||||
| Search wiki by content | MCP | `search_docs(query)` |
|
||||
| Read a wiki page | MCP or file | `get_page(path)` or `cat knowledge/wiki/.../...md` |
|
||||
| Get changelog entries | MCP | `get_changelog(page, since?)` |
|
||||
| Understand a service | MCP | `explain(service)` — compact context card, cheaper than search+read |
|
||||
| Blast-radius query | MCP | `get_relations(entity)` (ontology walk) |
|
||||
| List available secrets | MCP | `list_my_secrets()` (scoped to your age key) |
|
||||
| Browse or grep | File | Raw `grep` when MCP unreachable, or exploratory browsing |
|
||||
|
||||
**When MCP is unreachable:** fall back to grepping the clone at `/opt/homelab-context/`. The local files are the same; MCP is just an index.
|
||||
|
||||
---
|
||||
|
||||
## Understanding the Operating Model
|
||||
|
||||
Before you act, **classify your action against [oikos/policy.yaml](oikos/policy.yaml)**.
|
||||
|
||||
### The Oikos OODA Loop + Decision Tree
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Observe["**Observe**<br/>probes, drift detectors, agent signals"]
|
||||
Orient["**Orient**<br/>ontology, context, state, entity relations"]
|
||||
Decide{"**Decide**<br/>classify against oikos/policy.yaml"}
|
||||
Auto["Auto-act<br/>(unattended)"]
|
||||
Escalate["Escalate<br/>homelab approval request"]
|
||||
Act["**Act**<br/>homelab CLI, runbooks, skills"]
|
||||
Verify["**Verify**<br/>checklist from SKILL.md"]
|
||||
Ledger["**Ledger**<br/>mutation record: who/what/risk"]
|
||||
Document["**Document**<br/>wiki update, same-session rule"]
|
||||
|
||||
Observe --> Orient --> Decide
|
||||
Decide -->|read_only, reversible_low| Auto
|
||||
Decide -->|config_mutation, destructive| Escalate
|
||||
Auto --> Act
|
||||
Escalate -->|approval granted| Act
|
||||
Act --> Verify --> Ledger --> Document
|
||||
Document -.loop.-> Observe
|
||||
```
|
||||
|
||||
### Risk Classes (enforced, not advisory)
|
||||
|
||||
From [oikos/policy.yaml](oikos/policy.yaml):
|
||||
|
||||
- **read_only** — status, logs, docs, inventory queries. Unattended. MCP tools are all read_only.
|
||||
- **reversible_low** — restart, cache clear, sync pull. Unattended + ledger entry.
|
||||
- **config_mutation** — tracked-config edits (commit+push, never local), deploys, upgrades, DNS/ingress changes. **Operator approval required.**
|
||||
- **destructive** — destroy, format, wipe, rotate, revoke. **Approval + typed confirmation phrase.**
|
||||
|
||||
### Decision Flow
|
||||
|
||||
1. **Decide:** Use `homelab decide <action> <entity>` to classify (risk class × blast radius × confidence).
|
||||
2. **Escalate if needed:** `homelab approval request` (Matrix-delivered to operator; see [operations/commands.md](.agents/operations/commands.md)).
|
||||
3. **Execute:** Use `homelab` CLI (not ad-hoc SSH) — it enforces policy, logs mutations, and verifies outcomes.
|
||||
4. **Document:** Update wiki in the same session (per [AGENTS.md §5](AGENTS.md#5-acting-on-the-homelab) and the [same-session rule](.agents/shared/page-templates.md#same-session-update-rule)).
|
||||
|
||||
### The Ontology Graph
|
||||
|
||||
Everything that can break, be changed, or hold data has an entity in `inventory.yaml` + `oikos/ontology.yaml`. Blast-radius questions ("what breaks if strong goes down?") are graph walks via `homelab node <name> relations`, not doc archaeology.
|
||||
|
||||
**See:** [OIKOS.md](.agents/OIKOS.md) (full operating model, OODA loop, primitives, lifecycle gates, build status).
|
||||
|
||||
---
|
||||
|
||||
## Finding & Understanding Information
|
||||
|
||||
The narrative documentation is organized in **layers**:
|
||||
|
||||
| Layer | What it is | Where | Immutable? | How agents use it |
|
||||
|-------|-----------|-------|-----------|-------------------|
|
||||
| **Sources** | Raw evidence: incidents, external refs, live state | `knowledge/sources/investigations/` | Yes | Read to understand root causes; do not rewrite |
|
||||
| **Wiki** | Synthesized current-state: one page per node & per system | `knowledge/wiki/{containers,hosts,vms,infrastructure}/` | No | This is the reference layer — if wiki disagrees with live state, update it *in the same session* |
|
||||
| **Index** | Pure listings — every page in scope with one-line summary | `index.md` / folder `README.md` | No | Navigation aid; keep it current when wiki restructures |
|
||||
| **Log** | Append-only doc-maintenance record (restructures, ingests, lints) | `knowledge/log.md` | Yes (append-only) | Read to understand past doc changes; never edit directly |
|
||||
|
||||
**Changelog ≠ Log:** Each wiki page ends with a `## Changelog` (infrastructure changes to that node, machine-parsed). That's not the Log; the Log records *doc operations* only.
|
||||
|
||||
**See:** [llm-wiki.md](.agents/shared/llm-wiki.md) (full rules, page structure, immutability contract).
|
||||
|
||||
---
|
||||
|
||||
## Map & Quick Navigation
|
||||
|
||||
### Agent Entry Points (Start Here)
|
||||
|
||||
- **You are an agent** → [AGENTS.md](AGENTS.md) (on deployed clients: `/opt/homelab-context/AGENTS.md`)
|
||||
- **Operating model & risk policy** → [OIKOS.md](.agents/OIKOS.md)
|
||||
- **Specific workflows** → [.agents/skills/](.agents/skills/) (load the matching SKILL.md before acting)
|
||||
- **Operations cheatsheet** → [.agents/operations/commands.md](.agents/operations/commands.md)
|
||||
- **Tools & MCP reference** → [AGENTS.md §3 — The MCP server](AGENTS.md#3-the-mcp-server)
|
||||
|
||||
### Topology & Infrastructure
|
||||
|
||||
Node counts, IPs, and service lists change often — treat `inventory.yaml` and the index pages below as the source of truth, not this README.
|
||||
|
||||
- **Proxmox hosts** → [knowledge/wiki/hosts/index.md](knowledge/wiki/hosts/index.md)
|
||||
- **VMs** → [knowledge/wiki/vms/index.md](knowledge/wiki/vms/index.md)
|
||||
- **LXC containers** → [knowledge/wiki/containers/index.md](knowledge/wiki/containers/index.md)
|
||||
- **Cross-cutting infrastructure** (DNS, ingress, mesh, backups, monitoring, auto-deploy, VPS) → [knowledge/wiki/infrastructure/index.md](knowledge/wiki/infrastructure/index.md)
|
||||
|
||||
### Knowledge & References
|
||||
|
||||
- **Glossary** — [GLOSSARY.md](knowledge/GLOSSARY.md)
|
||||
- **Incidents & investigations** — [knowledge/sources/investigations/index.md](knowledge/sources/investigations/index.md) (active + [archive](knowledge/sources/investigations/archive/))
|
||||
- **Plans & design docs** — [plans/index.md](plans/index.md)
|
||||
- **Hermes agent** (for Hermes-enrolled clients) — [HERMES.md](.agents/HERMES.md)
|
||||
|
||||
---
|
||||
### Operations
|
||||
- [Command cheatsheet](operations/commands.md)
|
||||
- [Agent enrollment](operations/agent-enrollment.md) — bootstrap a new client (workstation, LXC, VM) into the homelab context system
|
||||
|
||||
## Conventions
|
||||
|
||||
All pages follow:
|
||||
- **Each node page** ends with a `## Changelog` section. Reverse-chronological. Entry format:
|
||||
```
|
||||
### YYYY-MM-DD — short title
|
||||
one or two lines on what changed and why.
|
||||
```
|
||||
- **Cross-linking is mandatory.** If a page references another node or system, link to it. Treat orphans as a bug.
|
||||
- **Live state wins.** When something here disagrees with `pct config` / `docker inspect` / running config, fix the wiki *and* note the change in the relevant changelog.
|
||||
- **Tracked configs.** A node whose config lives in a Gitea repo (Caddy, Gitea customizations, Artifacto, mule-image) is auto-deployed via webhook — see [auto-deploy](infrastructure/auto-deploy.md). Edits there must be pushed, not left local.
|
||||
- **No secrets.** This is a private repo on `git.hubris.network`, but still: paths to secret files are fine, secret values are not.
|
||||
|
||||
- **File naming.** Foundational docs (entry-points, agent instruction, references) are ALL-CAPS (`AGENTS.md`, `OIKOS.md`, `GLOSSARY.md`); containers use `<id>-<name>.md`; infrastructure pages use lowercase-with-dashes; plans and incidents use `YYYY-MM-DD-slug.md`; skills are `<name>/SKILL.md`. See [page-templates.md](.agents/shared/page-templates.md#file-naming) for the full rules.
|
||||
- **Voice & vocabulary.** Concise, technical, sysadmin-to-sysadmin. No marketing prose, no puffers (seamless, robust, leverage, etc.). Full rules in [writing-style.md](.agents/shared/writing-style.md).
|
||||
- **Cross-linking is mandatory.** If a page references a node or system, link to it. Treat orphans as a bug.
|
||||
- **Live state wins.** When something here disagrees with `pct config` / `docker inspect` / running state, fix the wiki *and* add a changelog entry *in the same session*.
|
||||
- **Tracked configs.** Pages for configs living in git repos (Caddy, Gitea, Artifacto, mule-image) must note the repo. Edits go through commit+push, never local changes. See [auto-deploy](knowledge/wiki/infrastructure/auto-deploy.md).
|
||||
- **No secrets.** This is a private repo, but still: reference secret *paths*, never secret *values*.
|
||||
## Maintaining this wiki
|
||||
|
||||
**For agents:** Read [caveman.md](.agents/shared/caveman.md) (terse communication standard). Use templates at [page-templates.md](.agents/shared/page-templates.md) when creating pages.
|
||||
When you change a node:
|
||||
1. Update the relevant page (config snapshot, ports, mounts).
|
||||
2. Add a changelog entry at the bottom of that page.
|
||||
3. If the change touches a cross-cutting system (DNS, Caddy, Authentik, mesh), update *that* page too and link it from the changelog entry.
|
||||
4. If it's an incident, add an entry to [`investigations/`](investigations/index.md).
|
||||
|
||||
---
|
||||
## See also
|
||||
|
||||
## Updating the Wiki
|
||||
|
||||
### When You Change Infrastructure
|
||||
|
||||
1. Update the relevant page (config snapshot, ports, mounts, IP address).
|
||||
2. Add a `### YYYY-MM-DD — title` entry to the page's `## Changelog` section (reverse chronological order).
|
||||
3. If the change touches a cross-cutting system (DNS, Caddy, Authentik, mesh), update *that* page too and link from the changelog.
|
||||
4. If it's an incident, add a record to [`knowledge/sources/investigations/`](knowledge/sources/investigations/index.md).
|
||||
|
||||
### When You Restructure the Wiki
|
||||
|
||||
1. Update the relevant `index.md` / `README.md` in that section.
|
||||
2. Add a single-line entry to [`knowledge/log.md`](knowledge/log.md): `## [YYYY-MM-DD] <operation> | <summary>` (e.g., `## [2026-07-06] restructure | split infrastructure/dns into dns.md + dns-advanced.md`).
|
||||
|
||||
### The Same-Session Update Rule
|
||||
|
||||
**Any meaningful state change made in this session requires a wiki update before the session closes.** A change that touches a container page must also update:
|
||||
- The `containers/index.md` table (IPs, host, mounts, status)
|
||||
- The root `README.md` table (if affected)
|
||||
- The Caddy page site list (if affects `*.hubris.network` routing)
|
||||
- The DNS / ingress infrastructure pages (if affects routing)
|
||||
- The `hosts/hubris.md` or `hosts/strong.md` page (if container count changes)
|
||||
- The `inventory.yaml` host entry (source of truth for `hosts/*.yaml` generation)
|
||||
- The `knowledge/wiki/infrastructure/topology.md` (regenerate if needed)
|
||||
|
||||
Not updating all linked places is a bug. See [page-templates.md — same-session update rule](.agents/shared/page-templates.md#same-session-update-rule).
|
||||
|
||||
---
|
||||
|
||||
## More Information
|
||||
|
||||
- **For Hermes agents** → [HERMES.md](.agents/HERMES.md) (persona, source-of-truth hierarchy, token efficiency)
|
||||
- **For manual workflows** → [.agents/operations/](.agents/operations/) (commands cheatsheet, agent enrollment, Hermes guide)
|
||||
- **For skills/runbooks** → [.agents/skills/](.agents/skills/) (load the matching SKILL.md before acting; includes risk class + verification)
|
||||
- **MCP tools** → [AGENTS.md §3](AGENTS.md#3-the-mcp-server) (available tools, when to use MCP vs files)
|
||||
- **Page templates & voice** → [.agents/shared/](.agents/shared/) (page-templates.md, writing-style.md, caveman.md, llm-wiki.md)
|
||||
- **Machine-readable substrate** → `inventory.yaml`, `oikos/policy.yaml`, `oikos/ontology.yaml` (not part of the wiki; see [llm-wiki.md](.agents/shared/llm-wiki.md#rules))
|
||||
- [`CONTRIBUTING.md`](CONTRIBUTING.md) — page templates and tone
|
||||
|
||||
@@ -1,8 +0,0 @@
|
||||
# oapi-codegen config — `make generate` regenerates internal/httpapi/gen.
|
||||
package: gen
|
||||
output: internal/httpapi/gen/api.gen.go
|
||||
generate:
|
||||
models: true
|
||||
chi-server: true
|
||||
strict-server: true
|
||||
embedded-spec: true
|
||||
2891
api/openapi.yaml
2891
api/openapi.yaml
File diff suppressed because it is too large
Load Diff
@@ -1,7 +0,0 @@
|
||||
# Redocly lint config for api/openapi.yaml (CI runs: redocly lint api/openapi.yaml)
|
||||
extends:
|
||||
- recommended
|
||||
rules:
|
||||
# Every operation declares `default` → RFC 9457 problem+json instead of
|
||||
# enumerating each 4XX (plan R3-3); oapi-codegen handles `default` fine.
|
||||
operation-4xx-response: off
|
||||
388
bin/homelab
388
bin/homelab
@@ -37,21 +37,6 @@ INVENTORY = CONTEXT / "inventory.yaml"
|
||||
HOSTS_DIR = CONTEXT / "hosts"
|
||||
AGE_KEY = Path(os.environ.get("SOPS_AGE_KEY_FILE", "/etc/age/key.txt"))
|
||||
|
||||
# Oikos kernel modules (policy classification, ontology relations, change
|
||||
# ledger). Optional at import time so a stale/partial checkout degrades to
|
||||
# "feature unavailable" instead of crashing every subcommand.
|
||||
sys.path.insert(0, str(CONTEXT))
|
||||
try:
|
||||
from oikos import approve as oikos_approve
|
||||
from oikos import decide as oikos_decide
|
||||
from oikos import ledger as oikos_ledger
|
||||
from oikos import policy as oikos_policy
|
||||
from oikos import relations as oikos_relations
|
||||
from oikos import signal as oikos_signal
|
||||
except ImportError:
|
||||
oikos_approve = oikos_decide = oikos_ledger = oikos_policy = None
|
||||
oikos_relations = oikos_signal = None
|
||||
|
||||
|
||||
# ---------- helpers ----------
|
||||
|
||||
@@ -187,25 +172,6 @@ def service_backend_host(name: str) -> str:
|
||||
return service(name)["backend"]
|
||||
|
||||
|
||||
def _record_change(entity: str, action: str, risk: str, *,
|
||||
result: str | None = None, verification: str | None = None,
|
||||
approval_ref: str | None = None) -> None:
|
||||
"""Append a ledger entry and push it standalone (used by mutations that
|
||||
don't already go through push_inventory, e.g. restart)."""
|
||||
if oikos_ledger is None:
|
||||
return
|
||||
oikos_ledger.append(entity, action, risk, result=result, verification=verification,
|
||||
approval_ref=approval_ref)
|
||||
try:
|
||||
subprocess.run(["git", "add", "ledger/"], check=True, cwd=CONTEXT)
|
||||
if subprocess.run(["git", "diff", "--cached", "--quiet"], cwd=CONTEXT).returncode != 0:
|
||||
subprocess.run(["git", "commit", "-m", f"ledger: {entity} {action} ({risk})"],
|
||||
check=True, cwd=CONTEXT)
|
||||
subprocess.run(["git", "push"], check=True, cwd=CONTEXT)
|
||||
except subprocess.CalledProcessError as e:
|
||||
print(f"warning: could not commit/push ledger entry: {e}", file=sys.stderr)
|
||||
|
||||
|
||||
def push_inventory(message: str, extra_paths: list[str] | None = None) -> None:
|
||||
"""Stage + commit + push inventory + regenerated hosts/ (+ any extras)."""
|
||||
subprocess.run(["python3", str(CONTEXT / "mcp" / "build_host_files.py")],
|
||||
@@ -604,31 +570,11 @@ def cmd_restart(args: argparse.Namespace) -> int:
|
||||
svc = args.service
|
||||
host_name = service_backend_host(svc)
|
||||
unit = service(svc).get("systemd_unit", svc)
|
||||
risk = (oikos_policy.classify_action("service-restart", svc)
|
||||
if oikos_policy else "reversible_low") or "reversible_low"
|
||||
approval = oikos_policy.approval_for(risk) if oikos_policy else "none"
|
||||
|
||||
# Mechanical gate: config_mutation/destructive risk classes require a
|
||||
# live grant regardless of -y/interactivity — an agent (or a human
|
||||
# bypassing the confirm() prompt with -y) cannot mutate a gated service
|
||||
# without a real oikos/approve.py approval. See oikos/policy.yaml.
|
||||
if approval != "none":
|
||||
if not args.approval_id:
|
||||
die(f"restarting '{svc}' is risk class '{risk}' (approval: {approval}) — "
|
||||
f"pass --approval-id <id> from an approved 'homelab approval request'")
|
||||
ok, reason = oikos_approve.check_grant(args.approval_id, f"service:{svc}", "service-restart")
|
||||
if not ok:
|
||||
die(f"approval {args.approval_id} not valid for this action: {reason}")
|
||||
|
||||
if not args.yes:
|
||||
if not confirm(f"restart systemd unit '{unit}' on {host_name}?"):
|
||||
return 1
|
||||
base = ssh_base(host_name)
|
||||
rc = subprocess.call(base + ["--", "systemctl", "restart", unit])
|
||||
_record_change(f"service:{svc}", "restart", risk,
|
||||
result=("ok" if rc == 0 else f"failed rc={rc}"),
|
||||
approval_ref=args.approval_id)
|
||||
return rc
|
||||
return subprocess.call(base + ["--", "systemctl", "restart", unit])
|
||||
|
||||
|
||||
def cmd_open(args: argparse.Namespace) -> int:
|
||||
@@ -1085,19 +1031,14 @@ def cmd_refresh_creds(args: argparse.Namespace) -> int:
|
||||
|
||||
|
||||
def cmd_sync(args: argparse.Namespace) -> int:
|
||||
# Unlike every other mutating command here, this one had no os.geteuid()
|
||||
# guard — always shelled out to sudo. Fails outright with "No such file
|
||||
# or directory: 'sudo'" on minimal root-only Linux images (no sudo
|
||||
# binary installed at all) reached via `ssh root@host`, e.g. strong.
|
||||
needs_sudo = os.geteuid() != 0
|
||||
if sys.platform == "darwin":
|
||||
cmd = ["launchctl", "kickstart", "-k",
|
||||
"system/network.hubris.homelab-context-sync"]
|
||||
else:
|
||||
cmd = ["systemctl", "start", "homelab-context-sync.service"]
|
||||
if needs_sudo:
|
||||
cmd = ["sudo"] + cmd
|
||||
return subprocess.call(cmd)
|
||||
return subprocess.call(
|
||||
["sudo", "launchctl", "kickstart", "-k",
|
||||
"system/network.hubris.homelab-context-sync"]
|
||||
)
|
||||
return subprocess.call(
|
||||
["sudo", "systemctl", "start", "homelab-context-sync.service"]
|
||||
)
|
||||
|
||||
|
||||
def cmd_mcp(args: argparse.Namespace) -> int:
|
||||
@@ -1157,11 +1098,9 @@ def cmd_client_add(args: argparse.Namespace) -> int:
|
||||
print("granting hermes-only secrets...")
|
||||
_grant_shared_secrets(pubkey, HERMES_SECRETS)
|
||||
commit_subject = f"client-add: {name} (finalize age_pubkey + grant shared + hermes secrets)"
|
||||
if oikos_ledger is not None:
|
||||
oikos_ledger.append(f"host:{name}", "client-add-finalize", "config_mutation", result="ok")
|
||||
push_inventory(
|
||||
commit_subject,
|
||||
extra_paths=[".sops.yaml", "secrets/", "ledger/"],
|
||||
extra_paths=[".sops.yaml", "secrets/"],
|
||||
)
|
||||
print(f"finalized {name}.")
|
||||
return 0
|
||||
@@ -1223,11 +1162,9 @@ def cmd_client_remove(args: argparse.Namespace) -> int:
|
||||
print(f" issuance revoke failed: {e}")
|
||||
|
||||
# 4. Commit + push (extras: .sops.yaml + secrets/ may also have changed).
|
||||
if oikos_ledger is not None:
|
||||
oikos_ledger.append(f"host:{name}", "client-remove", "destructive", result="ok")
|
||||
push_inventory(
|
||||
f"client-remove: {name}",
|
||||
extra_paths=[".sops.yaml", "secrets/", "ledger/"],
|
||||
extra_paths=[".sops.yaml", "secrets/"],
|
||||
)
|
||||
|
||||
print()
|
||||
@@ -1240,211 +1177,6 @@ def cmd_client_remove(args: argparse.Namespace) -> int:
|
||||
return 0
|
||||
|
||||
|
||||
def _require_oikos() -> None:
|
||||
if oikos_policy is None or oikos_relations is None or oikos_ledger is None:
|
||||
die("oikos/ kernel modules not importable — is this checkout up to date?")
|
||||
|
||||
|
||||
def cmd_service(args: argparse.Namespace) -> int:
|
||||
"""Service Console v0 — explain/health/docs/log/actions/history for one service."""
|
||||
_require_oikos()
|
||||
name = args.name
|
||||
svc = service(name) # dies with a clear message if unknown
|
||||
|
||||
if args.action == "explain":
|
||||
card = CONTEXT / "oikos" / "cards" / f"service-{name}.md"
|
||||
if not card.exists():
|
||||
die(f"no context card for {name} — run: python3 oikos/gen-topology.py")
|
||||
print(card.read_text())
|
||||
return 0
|
||||
|
||||
if args.action == "health":
|
||||
url = svc.get("url") or svc.get("endpoint")
|
||||
if not url:
|
||||
die(f"service {name} has no url/endpoint in inventory")
|
||||
if not args.live:
|
||||
try:
|
||||
from oikos import scheduler as oikos_scheduler
|
||||
cached = oikos_scheduler.cached_service_health(name)
|
||||
except ImportError:
|
||||
cached = None
|
||||
if cached is not None and cached.get("checked"):
|
||||
status = "ok" if cached.get("ok") else "unhealthy"
|
||||
print(f"{name}: {cached.get('checked_url', url)} -> "
|
||||
f"{cached.get('http_code') or 'no response'} ({status}, "
|
||||
f"as of {cached['as_of']} — pass --live to force a fresh probe)")
|
||||
return 0 if cached.get("ok") else 1
|
||||
proc = subprocess.run(
|
||||
["curl", "-sS", "-o", "/dev/null", "-w", "%{http_code}", "--max-time", "5", url],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
code = proc.stdout.strip() or "no response"
|
||||
print(f"{name}: {url} -> {code} (live probe)")
|
||||
return 0 if code.startswith(("2", "3")) else 1
|
||||
|
||||
if args.action == "docs":
|
||||
doc = svc.get("doc_page")
|
||||
if not doc:
|
||||
die(f"no doc_page recorded for {name} in inventory.yaml")
|
||||
path = CONTEXT / doc
|
||||
if not path.exists():
|
||||
die(f"doc_page {doc} does not exist")
|
||||
print(path.read_text())
|
||||
return 0
|
||||
|
||||
if args.action == "log":
|
||||
return cmd_logs(argparse.Namespace(service=name, lines=args.lines, follow=False))
|
||||
|
||||
if args.action == "actions":
|
||||
for a in oikos_policy.safe_actions_for_service(name, svc):
|
||||
print(f"{a['action']:<24} {a['risk']:<16} approval={a['approval']}")
|
||||
return 0
|
||||
|
||||
if args.action == "history":
|
||||
entries = oikos_ledger.history(f"service:{name}", limit=args.limit)
|
||||
if not entries:
|
||||
print(f"(no ledger entries for service:{name} yet)")
|
||||
for e in entries:
|
||||
print(json.dumps(e))
|
||||
return 0
|
||||
|
||||
die(f"unknown service action: {args.action}")
|
||||
|
||||
|
||||
def cmd_change_preflight(args: argparse.Namespace) -> int:
|
||||
"""Dry-run report before mutating a service: health, risk class, approval
|
||||
requirement, and the verification command to run after."""
|
||||
_require_oikos()
|
||||
name = args.service
|
||||
svc = service(name)
|
||||
risk = (oikos_policy.classify_action("tracked-config-edit", name)
|
||||
if svc.get("config_repo")
|
||||
else oikos_policy.classify_action("service-restart", name)) or "config_mutation"
|
||||
approval = oikos_policy.approval_for(risk)
|
||||
|
||||
print(f"Preflight: {name}")
|
||||
print(f" risk class: {risk} (approval: {approval})")
|
||||
|
||||
url = svc.get("url") or svc.get("endpoint")
|
||||
if url:
|
||||
proc = subprocess.run(
|
||||
["curl", "-sS", "-o", "/dev/null", "-w", "%{http_code}", "--max-time", "3", url],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
print(f" current health: {url} -> {proc.stdout.strip() or 'no response'}")
|
||||
if svc.get("config_repo"):
|
||||
print(f" config repo: {svc['config_repo']} "
|
||||
f"(verify the backend's working tree is clean before editing)")
|
||||
if svc.get("risk_notes"):
|
||||
print(f" risk notes: {svc['risk_notes']}")
|
||||
print(f" verification after change: "
|
||||
+ (f"curl -sf {url}" if url else f"homelab logs {name}"))
|
||||
if approval != "none":
|
||||
print(f" requires operator approval before mutating ({approval})")
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_node_relations(args: argparse.Namespace) -> int:
|
||||
"""Walk the ontology graph both directions for a host or service name."""
|
||||
_require_oikos()
|
||||
results = oikos_relations.relations_for_name(args.name)
|
||||
if not results:
|
||||
die(f"unknown entity: {args.name}")
|
||||
for r in results:
|
||||
print(f"entity: {r['entity']}")
|
||||
print(f" impacts: {', '.join(r['impacts']) or '(none)'}")
|
||||
print(f" affected by: {', '.join(r['affected_by']) or '(none)'}")
|
||||
print(f" full blast radius: {', '.join(r['blast_radius']) or '(none)'}")
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_decide(args: argparse.Namespace) -> int:
|
||||
"""Route a proposed action: auto-act or escalate. See oikos/decide.py."""
|
||||
_require_oikos()
|
||||
result = oikos_decide.classify(args.action, args.entity, service_name=args.service_name,
|
||||
record=not args.no_record)
|
||||
print(json.dumps(result, indent=2))
|
||||
return 0 if result["route"] == "auto-act" else 1
|
||||
|
||||
|
||||
def cmd_approval_request(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
entry = oikos_approve.request(
|
||||
args.entity, args.action, args.risk, args.evidence,
|
||||
verification=args.verification, requires_phrase=args.requires_phrase,
|
||||
ttl_hours=args.ttl_hours,
|
||||
)
|
||||
print(json.dumps({k: v for k, v in entry.items() if k != "matrix_message"}, indent=2))
|
||||
print()
|
||||
print("--- post this to Matrix ---")
|
||||
print(entry["matrix_message"])
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_approval_list(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
for e in oikos_approve.list_approvals(state=args.state):
|
||||
print(json.dumps(e))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_approval_reply(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
try:
|
||||
entry = oikos_approve.reply(args.id, args.decision, phrase=args.phrase,
|
||||
decided_by=args.decided_by)
|
||||
except (ValueError, RuntimeError) as e:
|
||||
die(str(e))
|
||||
print(json.dumps(entry, indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_approval_check(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
ok, reason = oikos_approve.check_grant(args.id, args.entity, args.action)
|
||||
print(f"{'GRANTED' if ok else 'DENIED'}: {reason}")
|
||||
return 0 if ok else 1
|
||||
|
||||
|
||||
def cmd_signal_raise(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
action = None
|
||||
if args.action_runbook or args.action_risk:
|
||||
action = {"runbook": args.action_runbook, "risk": args.action_risk}
|
||||
entry = oikos_signal.raise_signal(args.kind, args.severity, args.entity, args.evidence,
|
||||
likely_cause=args.likely_cause,
|
||||
recommended_action=action,
|
||||
verification=args.verification)
|
||||
print(json.dumps(entry, indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_signal_list(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
for e in oikos_signal.list_signals(state=args.state, entity=args.entity,
|
||||
severity=args.severity, kind=args.kind):
|
||||
print(json.dumps(e))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_signal_ack(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
print(json.dumps(oikos_signal.acknowledge(args.id, args.note), indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_signal_resolve(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
print(json.dumps(oikos_signal.resolve(args.id, args.note), indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_signal_mute(args: argparse.Namespace) -> int:
|
||||
_require_oikos()
|
||||
print(json.dumps(oikos_signal.mute(args.id, args.ttl_hours, args.note), indent=2))
|
||||
return 0
|
||||
|
||||
|
||||
def cmd_nuke(args: argparse.Namespace) -> int:
|
||||
name = args.name
|
||||
if not args.yes:
|
||||
@@ -1762,9 +1494,6 @@ def main() -> int:
|
||||
sp = sub.add_parser("restart", help="restart a service")
|
||||
sp.add_argument("service")
|
||||
sp.add_argument("--yes", "-y", action="store_true")
|
||||
sp.add_argument("--approval-id", default=None,
|
||||
help="required if the service's risk class needs approval "
|
||||
"(see 'homelab approval request')")
|
||||
sp.set_defaults(func=cmd_restart)
|
||||
|
||||
sp = sub.add_parser("open", help="open a service's URL in browser")
|
||||
@@ -1820,101 +1549,6 @@ def main() -> int:
|
||||
help="skip the pre-flight dpkg-audit gate AND proceed past snapshot failures")
|
||||
sp.set_defaults(func=cmd_apt_upgrade)
|
||||
|
||||
sp = sub.add_parser("service", help="Service Console v0 — explain/health/docs/log/actions/history")
|
||||
sp.add_argument("name")
|
||||
sp.add_argument("action", choices=["explain", "health", "docs", "log", "actions", "history"])
|
||||
sp.add_argument("--lines", "-n", type=int, default=200, help="for 'log'")
|
||||
sp.add_argument("--limit", type=int, default=20, help="for 'history'")
|
||||
sp.add_argument("--live", action="store_true",
|
||||
help="for 'health': force a fresh probe instead of the scheduler's cache")
|
||||
sp.set_defaults(func=cmd_service)
|
||||
|
||||
change = sub.add_parser("change", help="change/mutation workflow")
|
||||
chsub = change.add_subparsers(dest="action", required=True)
|
||||
ch_preflight = chsub.add_parser("preflight")
|
||||
ch_preflight.add_argument("service")
|
||||
ch_preflight.set_defaults(func=cmd_change_preflight)
|
||||
|
||||
sp = sub.add_parser("node", help="ontology queries on a host/service")
|
||||
sp.add_argument("name")
|
||||
sp.add_argument("action", choices=["relations"])
|
||||
sp.set_defaults(func=cmd_node_relations)
|
||||
|
||||
sp = sub.add_parser("decide", help="classify a proposed action: auto-act or escalate")
|
||||
sp.add_argument("action")
|
||||
sp.add_argument("entity")
|
||||
sp.add_argument("--service-name", default=None)
|
||||
sp.add_argument("--no-record", action="store_true",
|
||||
help="skip writing this classification to the change ledger")
|
||||
sp.set_defaults(func=cmd_decide)
|
||||
|
||||
approval = sub.add_parser("approval", help="approval-engine requests (escalate route)")
|
||||
apsub = approval.add_subparsers(dest="action", required=True)
|
||||
|
||||
ap_req = apsub.add_parser("request")
|
||||
ap_req.add_argument("entity")
|
||||
ap_req.add_argument("action")
|
||||
ap_req.add_argument("risk")
|
||||
ap_req.add_argument("evidence")
|
||||
ap_req.add_argument("--verification")
|
||||
ap_req.add_argument("--requires-phrase", action="store_true")
|
||||
ap_req.add_argument("--ttl-hours", type=int, default=24)
|
||||
ap_req.set_defaults(func=cmd_approval_request)
|
||||
|
||||
ap_list = apsub.add_parser("list")
|
||||
ap_list.add_argument("--state", choices=["pending", "approved", "denied", "expired", "executed"])
|
||||
ap_list.set_defaults(func=cmd_approval_list)
|
||||
|
||||
ap_reply = apsub.add_parser("reply")
|
||||
ap_reply.add_argument("id")
|
||||
ap_reply.add_argument("decision", choices=["approve", "deny"])
|
||||
ap_reply.add_argument("--phrase")
|
||||
ap_reply.add_argument("--decided-by")
|
||||
ap_reply.set_defaults(func=cmd_approval_reply)
|
||||
|
||||
ap_check = apsub.add_parser("check")
|
||||
ap_check.add_argument("id")
|
||||
ap_check.add_argument("entity")
|
||||
ap_check.add_argument("action")
|
||||
ap_check.set_defaults(func=cmd_approval_check)
|
||||
|
||||
signal = sub.add_parser("signal", help="the attention layer (oikos/signal.py)")
|
||||
sigsub = signal.add_subparsers(dest="action", required=True)
|
||||
|
||||
sig_raise = sigsub.add_parser("raise")
|
||||
sig_raise.add_argument("kind")
|
||||
sig_raise.add_argument("severity", choices=["info", "warning", "critical"])
|
||||
sig_raise.add_argument("entity")
|
||||
sig_raise.add_argument("evidence")
|
||||
sig_raise.add_argument("--likely-cause")
|
||||
sig_raise.add_argument("--action-runbook")
|
||||
sig_raise.add_argument("--action-risk")
|
||||
sig_raise.add_argument("--verification")
|
||||
sig_raise.set_defaults(func=cmd_signal_raise)
|
||||
|
||||
sig_list = sigsub.add_parser("list")
|
||||
sig_list.add_argument("--state", choices=["raised", "acknowledged", "acting", "resolved", "muted"])
|
||||
sig_list.add_argument("--entity")
|
||||
sig_list.add_argument("--severity", choices=["info", "warning", "critical"])
|
||||
sig_list.add_argument("--kind")
|
||||
sig_list.set_defaults(func=cmd_signal_list)
|
||||
|
||||
sig_ack = sigsub.add_parser("ack")
|
||||
sig_ack.add_argument("id")
|
||||
sig_ack.add_argument("--note")
|
||||
sig_ack.set_defaults(func=cmd_signal_ack)
|
||||
|
||||
sig_resolve = sigsub.add_parser("resolve")
|
||||
sig_resolve.add_argument("id")
|
||||
sig_resolve.add_argument("--note")
|
||||
sig_resolve.set_defaults(func=cmd_signal_resolve)
|
||||
|
||||
sig_mute = sigsub.add_parser("mute")
|
||||
sig_mute.add_argument("id")
|
||||
sig_mute.add_argument("--ttl-hours", type=int, default=24)
|
||||
sig_mute.add_argument("--note")
|
||||
sig_mute.set_defaults(func=cmd_signal_mute)
|
||||
|
||||
sp = sub.add_parser("nuke", help="shred /etc/age/key.txt + /opt/homelab-context on a host")
|
||||
sp.add_argument("name")
|
||||
sp.add_argument("--yes", "-y", action="store_true")
|
||||
@@ -1929,7 +1563,7 @@ def main() -> int:
|
||||
csub_add.add_argument("--with-hermes", action="store_true",
|
||||
help="also grant secrets/openrouter-api-key.yaml so this "
|
||||
"host can run the Hermes agent (see "
|
||||
".agents/operations/hermes-agent.md). Combine with --finalize-pubkey.")
|
||||
"operations/hermes-agent.md). Combine with --finalize-pubkey.")
|
||||
csub_add.set_defaults(func=cmd_client_add)
|
||||
csub_rm = csub.add_parser("remove")
|
||||
csub_rm.add_argument("name")
|
||||
|
||||
85
bootstrap.sh
85
bootstrap.sh
@@ -7,11 +7,7 @@
|
||||
# curl ... | sudo bash -s -- --with-mcp # also wire Claude's .mcp.json
|
||||
# curl ... | sudo bash -s -- --with-hermes # also install Goose + Hermes wrapper
|
||||
# curl ... | sudo bash -s -- --dry-run # show what would happen
|
||||
# curl ... | sudo bash -s -- --no-secrets # skip age-key issuance entirely
|
||||
# curl ... | sudo bash -s -- --no-mesh # get secrets over LAN only, skip
|
||||
# # installing/connecting Netbird
|
||||
# # (host must be on 192.168.8.0/24
|
||||
# # or otherwise reach secrets.hubris.network)
|
||||
# curl ... | sudo bash -s -- --no-secrets # skip age-key issuance
|
||||
#
|
||||
# Prerequisites the script verifies:
|
||||
# - running as root
|
||||
@@ -35,7 +31,6 @@ WITH_MCP=0
|
||||
WITH_HERMES=0
|
||||
DRY_RUN=0
|
||||
NO_SECRETS=0
|
||||
NO_MESH=0
|
||||
GITEA_TOKEN="${HOMELAB_GITEA_TOKEN:-}"
|
||||
GITEA_USER="${HOMELAB_GITEA_USER:-dtoro}"
|
||||
|
||||
@@ -46,7 +41,6 @@ while [ $# -gt 0 ]; do
|
||||
--with-hermes) WITH_HERMES=1; shift ;;
|
||||
--dry-run) DRY_RUN=1; shift ;;
|
||||
--no-secrets) NO_SECRETS=1; shift ;;
|
||||
--no-mesh) NO_MESH=1; shift ;;
|
||||
--gitea-token) GITEA_TOKEN="$2"; shift 2 ;;
|
||||
--gitea-user) GITEA_USER="$2"; shift 2 ;;
|
||||
--help|-h)
|
||||
@@ -89,21 +83,6 @@ run() {
|
||||
fi
|
||||
}
|
||||
|
||||
# Run a command as the enrolling human user when one exists (i.e. this
|
||||
# script was invoked via `sudo bash bootstrap.sh` from a real login), and
|
||||
# directly otherwise. Minimal Linux images (bare Proxmox/Debian installs
|
||||
# reached via `ssh root@host`) often don't even have a `sudo` binary
|
||||
# installed — calling `sudo -u root ...` on those unconditionally fails
|
||||
# with "sudo: command not found" even though we're already root and don't
|
||||
# need to switch users at all.
|
||||
run_as() {
|
||||
if [ -n "${SUDO_USER:-}" ] && [ "$SUDO_USER" != "root" ]; then
|
||||
sudo -u "$SUDO_USER" -- "$@"
|
||||
else
|
||||
"$@"
|
||||
fi
|
||||
}
|
||||
|
||||
# -------- preflight --------
|
||||
if [ "$(id -u)" -ne 0 ]; then
|
||||
echo "bootstrap.sh must run as root (use sudo)." >&2
|
||||
@@ -139,22 +118,6 @@ fi
|
||||
if [ "$NO_SECRETS" -eq 0 ]; then
|
||||
for cmd in age sops; do command -v "$cmd" >/dev/null || missing+=("$cmd"); done
|
||||
fi
|
||||
# sops isn't a real Debian/Fedora package (there is no apt/dnf "sops"), so it
|
||||
# always needs the direct-binary-download path, on both distros. Only Darwin
|
||||
# (brew) can install it via a package manager.
|
||||
install_sops_binary() {
|
||||
local sops_version=v3.9.4
|
||||
local arch
|
||||
arch="$(uname -m)"
|
||||
case "$arch" in
|
||||
x86_64|amd64) arch=amd64 ;;
|
||||
aarch64|arm64) arch=arm64 ;;
|
||||
*) echo "[bootstrap] unsupported arch for sops binary download: $arch" >&2; return 1 ;;
|
||||
esac
|
||||
curl -fsSL "https://github.com/getsops/sops/releases/download/${sops_version}/sops-${sops_version}.linux.${arch}" \
|
||||
-o /usr/local/bin/sops && chmod +x /usr/local/bin/sops
|
||||
}
|
||||
|
||||
if [ "${#missing[@]}" -gt 0 ]; then
|
||||
if [ "$DRY_RUN" -eq 1 ]; then
|
||||
echo "+ would install missing tools: ${missing[*]}"
|
||||
@@ -175,23 +138,13 @@ if [ "${#missing[@]}" -gt 0 ]; then
|
||||
for m in "${missing[@]}"; do
|
||||
case "$m" in
|
||||
python3-yaml) dnf_list+=("python3-pyyaml") ;;
|
||||
sops) install_sops_binary ;;
|
||||
*) dnf_list+=("$m") ;;
|
||||
esac
|
||||
done
|
||||
[ "${#dnf_list[@]}" -gt 0 ] && dnf install -y "${dnf_list[@]}"
|
||||
dnf install -y "${dnf_list[@]}"
|
||||
elif command -v apt-get >/dev/null 2>&1; then
|
||||
apt_list=()
|
||||
for m in "${missing[@]}"; do
|
||||
case "$m" in
|
||||
sops) install_sops_binary ;;
|
||||
*) apt_list+=("$m") ;;
|
||||
esac
|
||||
done
|
||||
if [ "${#apt_list[@]}" -gt 0 ]; then
|
||||
DEBIAN_FRONTEND=noninteractive apt-get update
|
||||
DEBIAN_FRONTEND=noninteractive apt-get install -y "${apt_list[@]}"
|
||||
fi
|
||||
DEBIAN_FRONTEND=noninteractive apt-get update
|
||||
DEBIAN_FRONTEND=noninteractive apt-get install -y "${missing[@]}"
|
||||
else
|
||||
echo "[bootstrap] no supported package manager for: ${missing[*]}" >&2
|
||||
echo "[bootstrap] install with your package manager + re-run" >&2
|
||||
@@ -212,14 +165,10 @@ if [ "${#missing[@]}" -gt 0 ]; then
|
||||
fi
|
||||
|
||||
# -------- ensure netbird is installed + connected (workstation/VM hosts) --------
|
||||
# Skipped on --no-secrets (LXCs that route via the LAN already), --no-mesh
|
||||
# (explicit opt-out — secrets issuance still works if the mesh check below
|
||||
# falls back to LAN reachability), and --dry-run. Installs netbird if
|
||||
# missing, then drives `netbird up` against the homelab management server.
|
||||
# The operator clicks the printed device-code URL once — this blocks
|
||||
# indefinitely if nobody approves it, so don't skip --no-mesh on a host
|
||||
# nobody's watching interactively.
|
||||
if [ "$NO_SECRETS" -eq 0 ] && [ "$DRY_RUN" -eq 0 ] && [ "$NO_MESH" -eq 0 ]; then
|
||||
# Skipped on --no-secrets (LXCs that route via the LAN already) and --dry-run.
|
||||
# Installs netbird if missing, then drives `netbird up` against the homelab
|
||||
# management server. The operator clicks the printed device-code URL once.
|
||||
if [ "$NO_SECRETS" -eq 0 ] && [ "$DRY_RUN" -eq 0 ]; then
|
||||
if ! command -v netbird >/dev/null 2>&1 && ! command -v tailscale >/dev/null 2>&1; then
|
||||
echo "[bootstrap] no mesh CLI found; installing netbird..."
|
||||
if [ "$OS" = "Darwin" ]; then
|
||||
@@ -481,9 +430,9 @@ if [ "$WITH_HERMES" -eq 1 ]; then
|
||||
echo "+ would run upstream goose installer and symlink to /usr/local/bin/goose"
|
||||
else
|
||||
# Upstream installer drops the binary at ~/.local/bin/goose for the
|
||||
# invoking user. We run it as $H_USER (via run_as) then symlink
|
||||
# system-wide.
|
||||
run_as env CONFIGURE=false \
|
||||
# invoking user. We run it as $H_USER then symlink system-wide.
|
||||
sudo -u "$H_USER" \
|
||||
env CONFIGURE=false \
|
||||
bash -c 'curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash'
|
||||
if [ -x "$H_HOME/.local/bin/goose" ]; then
|
||||
ln -sfn "$H_HOME/.local/bin/goose" /usr/local/bin/goose
|
||||
@@ -507,7 +456,7 @@ if [ "$WITH_HERMES" -eq 1 ]; then
|
||||
Linux) HERMES_LINK=/root/HERMES.md ;;
|
||||
Darwin) HERMES_LINK=/etc/HERMES.md ;;
|
||||
esac
|
||||
run "ln -sfn '$CLONE_DIR/.agents/HERMES.md' '$HERMES_LINK'"
|
||||
run "ln -sfn '$CLONE_DIR/HERMES.md' '$HERMES_LINK'"
|
||||
echo "[bootstrap] linked HERMES.md → $HERMES_LINK"
|
||||
|
||||
# 4. Drop the Goose config. Idempotent YAML merge — preserves any keys the
|
||||
@@ -563,7 +512,7 @@ PYEOF
|
||||
|
||||
# 5. Symlink HERMES.md as the global .goosehints — Goose injects it into
|
||||
# the system prompt on every session start.
|
||||
run "ln -sfn '$CLONE_DIR/.agents/HERMES.md' '$GOOSEHINTS'"
|
||||
run "ln -sfn '$CLONE_DIR/HERMES.md' '$GOOSEHINTS'"
|
||||
if [ "$DRY_RUN" -eq 0 ]; then
|
||||
chown -h "$H_USER" "$GOOSEHINTS" 2>/dev/null || true
|
||||
fi
|
||||
@@ -646,7 +595,7 @@ if [ "$HKIND" != "lxc" ] && [ "$HKIND" != "vm" ]; then
|
||||
# Make sure pipx is available; OS-specific install.
|
||||
if ! command -v pipx >/dev/null 2>&1; then
|
||||
if [ "$OS" = "Darwin" ] && command -v brew >/dev/null 2>&1; then
|
||||
run_as brew install pipx 2>&1 | tail -2 || true
|
||||
sudo -u "${SUDO_USER:-$USER}" brew install pipx 2>&1 | tail -2 || true
|
||||
elif command -v dnf >/dev/null 2>&1; then
|
||||
dnf install -y pipx 2>&1 | tail -2 || true
|
||||
elif command -v apt-get >/dev/null 2>&1; then
|
||||
@@ -654,9 +603,9 @@ if [ "$HKIND" != "lxc" ] && [ "$HKIND" != "vm" ]; then
|
||||
fi
|
||||
fi
|
||||
if command -v pipx >/dev/null 2>&1; then
|
||||
INVOKING_USER="${SUDO_USER:-root}"
|
||||
run_as bash -lc "pipx install 'mcp[cli]'" 2>&1 | tail -3 || true
|
||||
run_as bash -lc "pipx ensurepath" >/dev/null 2>&1 || true
|
||||
INVOKING_USER="${SUDO_USER:-$USER}"
|
||||
sudo -u "$INVOKING_USER" -- bash -lc "pipx install 'mcp[cli]'" 2>&1 | tail -3 || true
|
||||
sudo -u "$INVOKING_USER" -- bash -lc "pipx ensurepath" >/dev/null 2>&1 || true
|
||||
echo "[bootstrap] mcp CLI installed for $INVOKING_USER via pipx"
|
||||
else
|
||||
echo "[bootstrap] WARNING: pipx unavailable; install manually: pipx install 'mcp[cli]'" >&2
|
||||
|
||||
@@ -1,254 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"os"
|
||||
"os/signal"
|
||||
"syscall"
|
||||
|
||||
"net/http"
|
||||
|
||||
"github.com/dtoro/oikos/internal/config"
|
||||
"github.com/dtoro/oikos/internal/db"
|
||||
"github.com/dtoro/oikos/internal/httpapi"
|
||||
"github.com/dtoro/oikos/internal/observability"
|
||||
"github.com/jackc/pgx/v5"
|
||||
)
|
||||
|
||||
// SchedulerRunner is set by the scheduler init() to avoid circular imports.
|
||||
var SchedulerRunner func(context.Context, *db.Pool, config.Config)
|
||||
|
||||
// NotifierRunner is set by the notifier init() to avoid circular imports.
|
||||
var NotifierRunner func(context.Context, *db.Pool, config.Config)
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 2 {
|
||||
usage()
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
role := os.Args[1]
|
||||
cfg := config.FromEnv()
|
||||
|
||||
// Structured logging (slog)
|
||||
logger := observability.NewLogger(cfg.Debug)
|
||||
slog.SetDefault(logger)
|
||||
|
||||
slog.Info("starting oikos", "role", role, "config", cfg)
|
||||
|
||||
ctx, cancel := signal.NotifyContext(context.Background(),
|
||||
syscall.SIGTERM, syscall.SIGINT)
|
||||
defer cancel()
|
||||
|
||||
switch role {
|
||||
case "migrate":
|
||||
if err := runMigrate(ctx, cfg); err != nil {
|
||||
slog.Error("migrate failed", "error", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
case "seed":
|
||||
if err := runSeed(ctx, cfg); err != nil {
|
||||
slog.Error("seed failed", "error", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
case "export":
|
||||
if err := runExport(ctx, cfg); err != nil {
|
||||
slog.Error("export failed", "error", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
case "api":
|
||||
if err := runAPI(ctx, cfg); err != nil {
|
||||
slog.Error("api failed", "error", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
case "scheduler":
|
||||
if SchedulerRunner != nil {
|
||||
SchedulerRunner(ctx, nil, cfg)
|
||||
} else {
|
||||
slog.Error("scheduler not compiled in (import internal/scheduler)")
|
||||
os.Exit(1)
|
||||
}
|
||||
case "notifier":
|
||||
if NotifierRunner != nil {
|
||||
NotifierRunner(ctx, nil, cfg)
|
||||
} else {
|
||||
slog.Error("notifier not compiled in (import internal/notifier)")
|
||||
os.Exit(1)
|
||||
}
|
||||
case "all":
|
||||
slog.Info("all role not yet implemented (runs api + scheduler + notifier in one process)")
|
||||
os.Exit(1)
|
||||
case "version":
|
||||
fmt.Println("oikos dev (Phase 1)")
|
||||
case "help", "--help", "-h":
|
||||
usage()
|
||||
default:
|
||||
fmt.Fprintf(os.Stderr, "unknown role: %s\n", role)
|
||||
usage()
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
func usage() {
|
||||
fmt.Println(`oikos — the homelab OS
|
||||
|
||||
Usage: oikos <role> [flags]
|
||||
|
||||
Roles:
|
||||
migrate Run database migrations (forward-only, idempotent)
|
||||
seed Ingest seed YAML files into the database
|
||||
export Export DB state back to seed YAMLs (DR / version control)
|
||||
api Run the REST + MCP API server (Phase 2)
|
||||
scheduler Run the observe + act loop (Phase 3)
|
||||
notifier Run the notification service (Phase 3)
|
||||
all Run all roles in one process (dev mode)
|
||||
version Print version info
|
||||
|
||||
Environment:
|
||||
OIKOS_DATABASE_URL Postgres connection string
|
||||
OIKOS_API_LISTEN API listen address (default :8090)
|
||||
OIKOS_ENV Environment (dev, prod)
|
||||
OIKOS_DEBUG Enable verbose logging (true/1)
|
||||
OIKOS_SEEDS_DIR Path to seeds directory (default: seeds)
|
||||
OIKOS_MCP_BEARER_TOKEN Shared secret for MCP auth`)
|
||||
}
|
||||
|
||||
func runMigrate(ctx context.Context, cfg config.Config) error {
|
||||
pool, err := db.New(ctx, cfg.DatabaseURL)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer pool.Close()
|
||||
|
||||
slog.Info("running migrations")
|
||||
if err := pool.Migrate(ctx); err != nil {
|
||||
return err
|
||||
}
|
||||
slog.Info("migrations complete")
|
||||
return nil
|
||||
}
|
||||
|
||||
func runSeed(ctx context.Context, cfg config.Config) error {
|
||||
pool, err := db.New(ctx, cfg.DatabaseURL)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer pool.Close()
|
||||
|
||||
// Ensure migrations are applied first
|
||||
if err := pool.Migrate(ctx); err != nil {
|
||||
return fmt.Errorf("migrations: %w", err)
|
||||
}
|
||||
|
||||
seedsDir := cfg.SeedsDir
|
||||
if seedsDir == "" {
|
||||
seedsDir = "seeds"
|
||||
}
|
||||
|
||||
// Ingest ontology seed
|
||||
ontoContent, err := os.ReadFile(seedsDir + "/ontology.yaml")
|
||||
if err != nil {
|
||||
return fmt.Errorf("read ontology seed: %w", err)
|
||||
}
|
||||
err = pool.SeedIngest(ctx, "ontology.yaml", ontoContent,
|
||||
func(ctx context.Context, tx pgx.Tx, data map[string]any) error {
|
||||
r, err := db.IngestOntologySeed(ctx, tx, data)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
slog.Info("ontology ingested",
|
||||
"lifecycles", r.Lifecycles,
|
||||
"entity_types", r.EntityTypes,
|
||||
"relationship_types", r.RelationshipTypes)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
// Ingest inventory seed
|
||||
invContent, err := os.ReadFile(seedsDir + "/inventory.yaml")
|
||||
if err != nil {
|
||||
return fmt.Errorf("read inventory seed: %w", err)
|
||||
}
|
||||
err = pool.SeedIngest(ctx, "inventory.yaml", invContent,
|
||||
func(ctx context.Context, tx pgx.Tx, data map[string]any) error {
|
||||
r, err := db.IngestInventorySeed(ctx, tx, data)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
slog.Info("inventory ingested",
|
||||
"entities", r.Entities,
|
||||
"relationships", r.Relationships)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
// Ingest policy seed
|
||||
polContent, err := os.ReadFile(seedsDir + "/policy.yaml")
|
||||
if err != nil {
|
||||
return fmt.Errorf("read policy seed: %w", err)
|
||||
}
|
||||
err = pool.SeedIngest(ctx, "policy.yaml", polContent,
|
||||
func(ctx context.Context, tx pgx.Tx, data map[string]any) error {
|
||||
r, err := db.IngestPolicySeed(ctx, tx, data)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
slog.Info("policy ingested",
|
||||
"risk_classes", r.RiskClasses,
|
||||
"approval_rules", r.ApprovalRules,
|
||||
"autonomy_settings", r.AutonomySettings)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
slog.Info("seed ingest complete")
|
||||
return nil
|
||||
}
|
||||
|
||||
func runAPI(ctx context.Context, cfg config.Config) error {
|
||||
pool, err := db.New(ctx, cfg.DatabaseURL)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer pool.Close()
|
||||
|
||||
if err := pool.Migrate(ctx); err != nil {
|
||||
return fmt.Errorf("migrations: %w", err)
|
||||
}
|
||||
|
||||
err = httpapi.ListenAndServe(ctx, pool, cfg)
|
||||
if err == http.ErrServerClosed {
|
||||
return nil
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
func runExport(ctx context.Context, cfg config.Config) error {
|
||||
pool, err := db.New(ctx, cfg.DatabaseURL)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer pool.Close()
|
||||
|
||||
exports, err := db.ExportToYAML(ctx, pool)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
for name, content := range exports {
|
||||
path := cfg.SeedsDir + "/" + name
|
||||
if err := os.WriteFile(path, content, 0644); err != nil {
|
||||
return fmt.Errorf("write %s: %w", path, err)
|
||||
}
|
||||
slog.Info("exported", "file", path, "bytes", len(content))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -1,21 +0,0 @@
|
||||
# Multi-stage Dockerfile for Oikos (ADR 0001: single binary)
|
||||
FROM golang:1.26-alpine AS builder
|
||||
|
||||
RUN apk add --no-cache git ca-certificates
|
||||
|
||||
WORKDIR /build
|
||||
COPY go.mod go.sum ./
|
||||
RUN go mod download
|
||||
|
||||
COPY . .
|
||||
|
||||
RUN CGO_ENABLED=0 go build -o /oikos -tags timetzdata -ldflags="-s -w" ./cmd/oikos
|
||||
|
||||
# --- Runtime: distroless static ---
|
||||
FROM gcr.io/distroless/static:nonroot
|
||||
|
||||
COPY --from=builder /oikos /oikos
|
||||
COPY --from=builder /build/seeds /seeds
|
||||
COPY --from=builder /build/migrations /migrations
|
||||
|
||||
ENTRYPOINT ["/oikos"]
|
||||
34
containers/101-jellyfin.md
Normal file
34
containers/101-jellyfin.md
Normal file
@@ -0,0 +1,34 @@
|
||||
# 101 — `jellyfin`
|
||||
|
||||
Media server: serves the movies / TV / anime / music / audiobooks / podcasts libraries from `/mnt/library` to LAN clients.
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `jellyfin`
|
||||
- **IP:** `192.168.8.206`
|
||||
- **Privilege:** **unprivileged** + idmap (so it can write to the `media` group on `/mnt/library`)
|
||||
- **Resources:** 2 cores / 4 GiB RAM / 16 GiB rootfs
|
||||
- **Mounts:** `/mnt/library` ↔ `/mnt/library`
|
||||
- **Public hostname:** [`media.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md) → `:8096`
|
||||
|
||||
## Service / port map
|
||||
|
||||
| Service | Listen | Notes |
|
||||
| -------- | ------ | ----- |
|
||||
| jellyfin | `:8096` | HTTP (caddy terminates TLS) |
|
||||
|
||||
## Permissions
|
||||
Member of the [media GID 10000](../infrastructure/media-permissions.md) standard. Service user `jellyfin` is in the `media` group inside the container; idmap block in `/etc/pve/lxc/101.conf` maps in-container GID 10000 to host GID 10000.
|
||||
|
||||
## Related
|
||||
- [Caddy reverse proxy](121-caddy.md)
|
||||
- [Media permissions](../infrastructure/media-permissions.md)
|
||||
- [arriman](122-arriman.md) — \*arr stack writes the libraries jellyfin reads
|
||||
- [DNS split-horizon](../infrastructure/dns.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-04-28 — wiki entry created
|
||||
Initial documentation. No config changes.
|
||||
|
||||
### 2026-04-20 — joined the `media` GID 10000 standard
|
||||
Idmap block applied; in-container `media` group at GID 10000 mapped to host GID 10000. See [media permissions](../infrastructure/media-permissions.md). Config backup: `/root/101.conf.bak.*`.
|
||||
@@ -54,7 +54,7 @@ We considered three options before building this:
|
||||
|
||||
| Option | Outcome |
|
||||
|---|---|
|
||||
| **NFS on hubris bare-metal host** | Best performance, but adds long-lived NFS/RPC daemons to a host with a recent crash episode ([hubris crash 2026-04-21/22](../../sources/investigations/index.md)). Rejected. |
|
||||
| **NFS on hubris bare-metal host** | Best performance, but adds long-lived NFS/RPC daemons to a host with a recent crash episode ([hubris crash 2026-04-21/22](../investigations/index.md)). Rejected. |
|
||||
| **SMB on host** | Same host-blast-radius problem, plus 30–50% lower throughput than NFS on Linux↔Linux. Rejected. |
|
||||
| **NFS in a dedicated LXC** ← this | Within ~2% of host performance (LXC is namespace isolation; IO path is unchanged), zero new daemons on hubris, matches the existing fleet pattern. Selected. |
|
||||
|
||||
@@ -18,7 +18,7 @@ Paperless-ngx for document management. Ingests scans / PDFs from `/mnt/library/d
|
||||
| paperless-task-queue, paperless-scheduler, paperless-consumer | — | systemd workers |
|
||||
|
||||
## Auth
|
||||
Behind [Authentik forward-auth](106-auth-outpost.md). API path `/api/*` bypasses forward-auth (mobile clients can't follow the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates matching users on first SSO login.
|
||||
Behind [Authentik forward-auth](124-authentik.md). API path `/api/*` bypasses forward-auth (mobile clients can't follow the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates matching users on first SSO login.
|
||||
|
||||
## Storage
|
||||
- Documents at `/mnt/library/documents` (owner `www-data:www-data`, mode 750 — *not* on the `media` group, by design).
|
||||
@@ -27,7 +27,7 @@ Behind [Authentik forward-auth](106-auth-outpost.md). API path `/api/*` bypasses
|
||||
- ~~Disk usage was 86.9% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
|
||||
|
||||
## Related
|
||||
- [Authentik](106-auth-outpost.md)
|
||||
- [Authentik](124-authentik.md)
|
||||
- [Caddy](121-caddy.md)
|
||||
- [DNS](../infrastructure/dns.md)
|
||||
- [Monitoring](../infrastructure/monitoring.md)
|
||||
@@ -55,7 +55,7 @@ Initial documentation.
|
||||
Added `192.168.8.205`. See [Artifacto auto-deploy on apps (105)](105-apps.md).
|
||||
|
||||
### 2026-04-21 — `/etc/hosts` override for `auth.hubris.network` added
|
||||
For OIDC integration with [authentik (124)](106-auth-outpost.md). Outside the PVE markers, with a hubris-hosts-override.service for idempotency.
|
||||
For OIDC integration with [authentik (124)](124-authentik.md). Outside the PVE markers, with a hubris-hosts-override.service for idempotency.
|
||||
|
||||
### 2026-04-20 — gitea customizations + auto-deploy pipeline shipped
|
||||
`dtoro/gitea-customizations` repo created; webhook receiver at loopback `:9797` validates HMAC and runs `deploy.sh`. CAD and PlantUML loaders live in `footer.tmpl`.
|
||||
@@ -100,7 +100,7 @@ Native OIDC via `[oauth.generic]` in `config/config.ini`. `host = https://auth.h
|
||||
## Related
|
||||
- [Gitea (104)](104-gitea.md) — uses the PlantUML server
|
||||
- [Caddy (121)](121-caddy.md)
|
||||
- [Authentik (124)](106-auth-outpost.md)
|
||||
- [Authentik (124)](124-authentik.md)
|
||||
- [DNS](../infrastructure/dns.md)
|
||||
- [Auto-deploy](../infrastructure/auto-deploy.md)
|
||||
- [Public ingress (Artifacto + blog)](../infrastructure/ingress.md)
|
||||
@@ -116,7 +116,7 @@ Two new services from the [homelab-context distribution plan](../infrastructure/
|
||||
`secrets-issuance.service` on `:9820` (per-client age-key provisioning).
|
||||
Caddy fronts both with Let's Encrypt; new vhosts on
|
||||
[caddy](121-caddy.md), split-horizon DNS entries on
|
||||
[authentik (124)](106-auth-outpost.md). Gitea webhook ids 10 + 11 wire
|
||||
[authentik (124)](124-authentik.md). Gitea webhook ids 10 + 11 wire
|
||||
auto-deploy. LXC is itself an enrolled context client
|
||||
(`/opt/homelab-context/`).
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 106 — `auth-outpost`
|
||||
|
||||
Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that connects outbound to the [VPS Authentik core](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md) and serves forward-auth locally, so [Caddy (121)](121-caddy.md) never hairpins auth through VPS Traefik.
|
||||
Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that connects outbound to the [VPS Authentik core](../investigations/2026-05-31-authentik-vps-migration.md) and serves forward-auth locally, so [Caddy (121)](121-caddy.md) never hairpins auth through VPS Traefik.
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `auth-outpost`
|
||||
@@ -8,11 +8,11 @@ Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that co
|
||||
- **Privilege:** privileged (Docker-in-LXC, `features: nesting=1`)
|
||||
- **Resources:** 1 core / 512 MiB / 4 GiB rootfs
|
||||
- **Mounts:** none
|
||||
- **Created:** 2026-06-01, Debian 13, replacing the embedded outpost on [124](106-auth-outpost.md)
|
||||
- **Created:** 2026-06-01, Debian 13, replacing the embedded outpost on [124](124-authentik.md)
|
||||
|
||||
## Role
|
||||
|
||||
Runs one container — `ghcr.io/goauthentik/proxy` — that opens an outbound websocket to `https://auth.hubris.network` (the VPS core), pulls its proxy-provider config, and answers Caddy's `forward_auth` subrequests on `192.168.8.6:9000` (LAN-only bind). Because the call path is **Caddy → outpost (LAN)**, with no Traefik in between, `X-Forwarded-Host` is preserved — the failure that 404s when Caddy is pointed at `https://auth.hubris.network` directly (Traefik rewrites the header). See the [migration investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md).
|
||||
Runs one container — `ghcr.io/goauthentik/proxy` — that opens an outbound websocket to `https://auth.hubris.network` (the VPS core), pulls its proxy-provider config, and answers Caddy's `forward_auth` subrequests on `192.168.8.6:9000` (LAN-only bind). Because the call path is **Caddy → outpost (LAN)**, with no Traefik in between, `X-Forwarded-Host` is preserved — the failure that 404s when Caddy is pointed at `https://auth.hubris.network` directly (Traefik rewrites the header). See the [migration investigation](../investigations/2026-05-31-authentik-vps-migration.md).
|
||||
|
||||
## Service / port map
|
||||
| Service | Listen | Notes |
|
||||
@@ -42,15 +42,15 @@ Fix: the LAN outpost gets its **own** domain.
|
||||
**Lesson:** when the IdP core and the forward-auth outpost live on different hosts, the outpost needs a dedicated domain distinct from the core's — and proxy-provider `redirect_uris` must be regenerated, not just `external_host`.
|
||||
|
||||
## Related
|
||||
- [124 — authentik](106-auth-outpost.md) — old embedded-outpost host (now DNS-only)
|
||||
- [124 — authentik](124-authentik.md) — old embedded-outpost host (now DNS-only)
|
||||
- [Caddy (121)](121-caddy.md) — forward-auth consumer
|
||||
- [Ingress (VPS traefik)](../infrastructure/ingress.md)
|
||||
- [Authentik VPS migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)
|
||||
- [Authentik VPS migration](../investigations/2026-05-31-authentik-vps-migration.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-06-06 — Authentik session lifetime extended to 30 days
|
||||
VPS Authentik core `user_login` stage updated: `session_duration` changed from `seconds=0` (session cookie, cleared on browser close) to `days=30` (persistent 30-day cookie). Also set `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30` in `/opt/authentik.env` on the VPS. See [investigation](../../sources/investigations/2026-06-06-authentik-session-lifetime.md).
|
||||
VPS Authentik core `user_login` stage updated: `session_duration` changed from `seconds=0` (session cookie, cleared on browser close) to `days=30` (persistent 30-day cookie). Also set `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30` in `/opt/authentik.env` on the VPS. See [investigation](../investigations/2026-06-06-authentik-session-lifetime.md).
|
||||
|
||||
### 2026-06-01 — created; forward-auth cut over from LXC 124
|
||||
New dedicated LXC for the LAN forward-auth outpost (Phase 1 of the [architecture migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)). Deployed `goauthentik/proxy:2026.5.2` pointed at the VPS core; repointed Caddy `(authentik)` from `192.168.8.180:9000` → `192.168.8.6:9000`. Verified Paperless/qBittorrent/Artifacto return the SSO redirect with **124-Authentik stopped**, confirming the frozen instance is out of the path. dnsmasq stays on 124 until [DNS is relocated](106-auth-outpost.md).
|
||||
New dedicated LXC for the LAN forward-auth outpost (Phase 1 of the [architecture migration](../investigations/2026-05-31-authentik-vps-migration.md)). Deployed `goauthentik/proxy:2026.5.2` pointed at the VPS core; repointed Caddy `(authentik)` from `192.168.8.180:9000` → `192.168.8.6:9000`. Verified Paperless/qBittorrent/Artifacto return the SSO redirect with **124-Authentik stopped**, confirming the frozen instance is out of the path. dnsmasq stays on 124 until [DNS is relocated](124-authentik.md).
|
||||
@@ -1,6 +1,6 @@
|
||||
# 107 — `dns`
|
||||
|
||||
Homelab DNS server (Technitium). Replaces the dnsmasq that lived on [124 — authentik](106-auth-outpost.md); single-purpose, one job.
|
||||
Homelab DNS server (Technitium). Replaces the dnsmasq that lived on [124 — authentik](124-authentik.md); single-purpose, one job.
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `dns`
|
||||
@@ -29,7 +29,7 @@ Authoritative split-horizon DNS for `hubris.network` on the LAN/mesh, plus recur
|
||||
- **Plain LAN clients (`192.168.178.x`):** Fritz!Box DHCP still hands out Fritz!Box itself (`192.168.178.1`) as DNS — no split-horizon for non-mesh clients. Changing this requires a secondary DNS fallback, which Fritz!OS 8.x doesn't expose in a single DHCP field.
|
||||
|
||||
## dns-sync (Technitium = authoring source)
|
||||
`/opt/dns-sync/sync.py` (cron `*/10`, logs `/var/log/dns-sync.log`) reconciles this zone's named A-records → the NetBird managed DNS zone via the NetBird API (`/api/dns/zones/{id}/records`). Token at `/opt/dns-sync/netbird-token` (mode 600; source of truth in sops `secrets/netbird-pat.yaml`). **Edit DNS only here**; the sync propagates to the mesh. It deletes NetBird records absent from Technitium. Tracked: [scripts/dns-sync.py](../../../scripts/dns-sync.py). *Why this exists:* NetBird won't forward to Technitium for mesh peers (self-IP / nameserver-group quirks), so we sync into the managed zone instead — see [dns.md](../infrastructure/dns.md).
|
||||
`/opt/dns-sync/sync.py` (cron `*/10`, logs `/var/log/dns-sync.log`) reconciles this zone's named A-records → the NetBird managed DNS zone via the NetBird API (`/api/dns/zones/{id}/records`). Token at `/opt/dns-sync/netbird-token` (mode 600; source of truth in sops `secrets/netbird-pat.yaml`). **Edit DNS only here**; the sync propagates to the mesh. It deletes NetBird records absent from Technitium. Tracked: [scripts/dns-sync.py](../scripts/dns-sync.py). *Why this exists:* NetBird won't forward to Technitium for mesh peers (self-IP / nameserver-group quirks), so we sync into the managed zone instead — see [dns.md](../infrastructure/dns.md).
|
||||
|
||||
## DHCP
|
||||
|
||||
@@ -42,7 +42,7 @@ Technitium also runs a DHCP server for the homelab subnet (enabled 2026-06-02):
|
||||
Replaces the DHCP that was previously served by the Slate AX router. Static-IP LXCs (`.101–.239`) are excluded from the pool. Pool narrowed from `.100–.240` to `.241–.254` on 2026-06-03 to eliminate IP conflict risk.
|
||||
|
||||
## Related
|
||||
- [124 — authentik](106-auth-outpost.md) — retired host of the old dnsmasq
|
||||
- [124 — authentik](124-authentik.md) — retired host of the old dnsmasq
|
||||
- [DNS split-horizon](../infrastructure/dns.md)
|
||||
- [Mesh](../infrastructure/mesh.md)
|
||||
|
||||
@@ -55,7 +55,7 @@ Added for [trmnl (128)](128-trmnl.md) (LAN path via [Caddy (121)](121-caddy.md))
|
||||
Although the 2026-06-03 changelog claimed "cron */10", **no crontab was actually configured** on the LXC. The sync was running only via ad-hoc manual invocations during incident debugging. Fixed by adding `/etc/cron.d/dns-sync`.
|
||||
|
||||
### 2026-06-03 — DHCP pool narrowed to `.241–.254`
|
||||
Previous pool `.100–.240` overlapped with all static LXCs/VMs (`.101–.239`). Shrunk via API (`/api/dhcp/scopes/set`). 11 stale DHCP leases in `.101–.110` remain until natural expiry (2026-06-04). See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).
|
||||
Previous pool `.100–.240` overlapped with all static LXCs/VMs (`.101–.239`). Shrunk via API (`/api/dhcp/scopes/set`). 11 stale DHCP leases in `.101–.110` remain until natural expiry (2026-06-04). See [plan](../plans/2026-06-03-dhcp-pool-exclude-static-ips.md).
|
||||
|
||||
### 2026-06-03 — dns-sync added (Technitium → NetBird managed zone)
|
||||
This Technitium became the single DNS authoring source; `/opt/dns-sync/sync.py` (cron */10) reconciles named A-records into the NetBird managed zone via the API. Fixed previously-broken mesh names (`sso`, `nfs-export`, `mcp`, `secrets`) by adding them to the managed zone; reaped obsolete `files`/`photos-new`. See [dns.md](../infrastructure/dns.md).
|
||||
@@ -64,4 +64,4 @@ This Technitium became the single DNS authoring source; `/opt/dns-sync/sync.py`
|
||||
Enabled Technitium's built-in DHCP server for `192.168.8.0/24` (scope `homelab`, range `.100–.240`, gateway `192.168.8.1`, DNS self). Previously the Slate AX sub-router served DHCP for the homelab subnet. With the Slate AX retired and Proxmox now the subnet router, Technitium takes over DHCP. Configured via the Technitium API (`/api/dhcp/scopes/set`). DHCP LXCs kept their Slate AX leases until expiry, then renewed from Technitium.
|
||||
|
||||
### 2026-06-01 — created; replaced dnsmasq on 124
|
||||
Stood up Technitium at `192.168.8.2`, imported the split-horizon zone (specific A + wildcard + MX/SPF/CAA), made it the primary nameserver in the NetBird `home-lab-dns` group. Verified all names resolve with dnsmasq/124 stopped; [LXC 124 retired](106-auth-outpost.md).
|
||||
Stood up Technitium at `192.168.8.2`, imported the split-horizon zone (specific A + wildcard + MX/SPF/CAA), made it the primary nameserver in the NetBird `home-lab-dns` group. Verified all names resolve with dnsmasq/124 stopped; [LXC 124 retired](124-authentik.md).
|
||||
@@ -11,7 +11,7 @@ Personal cloud / file collaboration. Source-of-truth for the photo libraries sur
|
||||
- **Public hostname:** [`cloud.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md)
|
||||
|
||||
## Auth
|
||||
Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin` via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider <name> --mapping-uid=nc_uid` and `--scope="openid profile email <app>-uid"`. See [Authentik](106-auth-outpost.md#per-app-username-override-pattern-authentik) for the full pattern.
|
||||
Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin` via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider <name> --mapping-uid=nc_uid` and `--scope="openid profile email <app>-uid"`. See [Authentik](124-authentik.md#per-app-username-override-pattern-authentik) for the full pattern.
|
||||
|
||||
Redirect URI: `/index.php/apps/user_oidc/code` (NOT `/apps/...` — pretty URLs aren't on).
|
||||
|
||||
@@ -74,7 +74,7 @@ Apply with `systemctl restart mariadb` (not reload — `innodb_log_file_size` ne
|
||||
|
||||
## Related
|
||||
- [mulita (120)](120-mule-images.md) — reads NC user trees + writes back via WebDAV
|
||||
- [Authentik (124)](106-auth-outpost.md)
|
||||
- [Authentik (124)](124-authentik.md)
|
||||
- [Caddy (121)](121-caddy.md)
|
||||
- [DNS](../infrastructure/dns.md)
|
||||
- [Mesh migration (DNS overrides explained)](../infrastructure/mesh.md)
|
||||
@@ -4,8 +4,7 @@ Matrix homeserver (Synapse). Backs `@dtoro:avispero`.
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `elementsynapse`
|
||||
- **IP:** `192.168.8.242`
|
||||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||||
- **IP:** `192.168.8.239`
|
||||
- **Privilege:** **unprivileged**
|
||||
- **Resources:** 1 core / 2 GiB RAM / **16 GiB rootfs** (grown from 8 GiB on 2026-05-15 after disk-full incident)
|
||||
- **Mounts:** none from `/mnt/library`
|
||||
@@ -37,7 +36,7 @@ All five bridges run as plain `docker compose` stacks under `/root/mautrix-<name
|
||||
- ~~Disk usage was 86.8% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
|
||||
|
||||
## Related
|
||||
- ~~[claudio-bot (123)](archive/123-claudio-bot.md)~~ — decommissioned 2026-06-04, replaced by Hermes Agent
|
||||
- ~~[claudio-bot (123)](123-claudio-bot.md)~~ — decommissioned 2026-06-04, replaced by Hermes Agent
|
||||
- [Caddy](121-caddy.md)
|
||||
- [DNS](../infrastructure/dns.md)
|
||||
- [Monitoring](../infrastructure/monitoring.md)
|
||||
@@ -60,7 +60,7 @@ For pushes from inside the LXC, gitea creds at `/etc/mule-deploy/git-credentials
|
||||
|
||||
## Related
|
||||
- [Nextcloud (114)](114-nextcloud.md) — source of truth for photo libraries
|
||||
- [Authentik (124)](106-auth-outpost.md)
|
||||
- [Authentik (124)](124-authentik.md)
|
||||
- [Caddy (121)](121-caddy.md)
|
||||
- [DNS](../infrastructure/dns.md)
|
||||
- [Auto-deploy](../infrastructure/auto-deploy.md)
|
||||
@@ -11,16 +11,16 @@ The reverse proxy. Terminates TLS for every `*.hubris.network` hostname on the L
|
||||
- **Config:** `/etc/caddy/Caddyfile` is a [git checkout of `dtoro/caddy-conf`](#auto-deploy)
|
||||
- **Cert source:** Let's Encrypt **DNS-01** via IONOS API (`IONOS_AUTH_API_TOKEN`).
|
||||
|
||||
## Sites currently served (live as of 2026-07-06)
|
||||
## Sites currently served (live as of 2026-04-28)
|
||||
|
||||
- `artifacto.hubris.network` → [apps (105)](105-apps.md) `:3100`
|
||||
- `auth.hubris.network` → [authentik (124)](124-authentik.md) `:9000`
|
||||
- `blog.hubris.network` → [apps (105)](105-apps.md) `:8080`
|
||||
- `books.hubris.network` → [grimmory (130)](130-grimmory.md) `:6060`
|
||||
- `books.hubris.network` → [apps (105)](105-apps.md) `:6060`
|
||||
- `cloud.hubris.network` → [nextcloud (114)](114-nextcloud.md) `:443`
|
||||
- `docker.hubris.network` → [apps (105)](105-apps.md) `:9443`
|
||||
- `git.hubris.network` → [gitea (104)](104-gitea.md) `:3000` (+ `handle_path /_plantuml/*` → apps `:8079`)
|
||||
- `home.hubris.network` → [haos VM (108)](../vms/108-haos.md) `192.168.8.101:8123`
|
||||
- `house.hubris.network` → [house (129)](129-house.md) `:3000`
|
||||
- `jellyseerr.hubris.network` → [arriman (122)](122-arriman.md) `:5056`
|
||||
- `matrix.hubris.network` → [elementsynapse (118)](118-elementsynapse.md) `:8008`
|
||||
- `media.hubris.network` → [jellyfin (101)](101-jellyfin.md) `:8096`
|
||||
@@ -28,18 +28,15 @@ The reverse proxy. Terminates TLS for every `*.hubris.network` hostname on the L
|
||||
- `photos.hubris.network` → [mule-images (120)](120-mule-images.md) `:3000`
|
||||
- `proxmox.hubris.network` → [hubris host](../hosts/hubris.md) `:8006`
|
||||
- `qbit.hubris.network` → [arriman (122)](122-arriman.md) `:8080`
|
||||
- `roms.hubris.network` → [romm (134)](134-romm.md) `:80`
|
||||
- `sab.hubris.network` → [arriman (122)](122-arriman.md) `:8082` (Authentik forward-auth)
|
||||
- `teddy.hubris.network` → LXC 131 `192.168.8.150:8443`
|
||||
- `trmnl.hubris.network` → [trmnl (128)](128-trmnl.md) `:9851`
|
||||
|
||||
> **Reminder:** Caddy alone isn't enough to make a new subdomain reachable on the LAN. Each one needs an entry in [DNS split-horizon](../infrastructure/dns.md) too.
|
||||
|
||||
## Snippet: `(authentik)` forward-auth
|
||||
|
||||
A snippet at the top of the Caddyfile (used as `import authentik` in any site block) wires forward-auth to the embedded Authentik outpost. It points at `http://192.168.8.180:9000` directly (NOT `https://auth.hubris.network`) to avoid Caddy-to-self round-tripping that strips `X-Forwarded-Host`. The forward-auth block must explicitly set `header_up X-Forwarded-Host {host}`. See [Authentik](106-auth-outpost.md#forward-auth-domain-level-setup).
|
||||
A snippet at the top of the Caddyfile (used as `import authentik` in any site block) wires forward-auth to the embedded Authentik outpost. It points at `http://192.168.8.180:9000` directly (NOT `https://auth.hubris.network`) to avoid Caddy-to-self round-tripping that strips `X-Forwarded-Host`. The forward-auth block must explicitly set `header_up X-Forwarded-Host {host}`. See [Authentik](124-authentik.md#forward-auth-domain-level-setup).
|
||||
|
||||
For apps with mobile clients, `/api/*` (or equivalent) bypasses forward-auth — see the per-app gotchas in [Authentik](106-auth-outpost.md).
|
||||
For apps with mobile clients, `/api/*` (or equivalent) bypasses forward-auth — see the per-app gotchas in [Authentik](124-authentik.md).
|
||||
|
||||
## Caddy environment
|
||||
|
||||
@@ -60,7 +57,7 @@ Gitea webhook id 2 on `dtoro/caddy-conf`. Receiver, deploy script, install scrip
|
||||
|
||||
## Related
|
||||
- [DNS split-horizon](../infrastructure/dns.md) — must add entry for every new subdomain
|
||||
- [Authentik (124)](106-auth-outpost.md) — forward-auth + IdP
|
||||
- [Authentik (124)](124-authentik.md) — forward-auth + IdP
|
||||
- [Auto-deploy](../infrastructure/auto-deploy.md)
|
||||
- [Public ingress (VPS traefik)](../infrastructure/ingress.md) — mirrors Caddy's certs to the VPS for public exposure
|
||||
- [Gitea (104)](104-gitea.md) — webhook source
|
||||
@@ -85,7 +82,7 @@ Gitea webhook id 2 on `dtoro/caddy-conf`. Receiver, deploy script, install scrip
|
||||
- **Dirty-tree auto-stash:** stashes local changes before `git pull --ff-only` so the webhook doesn't fail on local edits
|
||||
- **Auto-backup:** saves `Caddyfile.bak.<timestamp>` before any modifications, keeps last 5
|
||||
|
||||
Also: [elementsynapse LXC 118](118-elementsynapse.md) found to have DHCP-overridden static IP (actual `.244` vs config `.239`) during incident investigation — fixed.
|
||||
Also: [elementsynapse LXC 118](../containers/118-elementsynapse.md) found to have DHCP-overridden static IP (actual `.244` vs config `.239`) during incident investigation — fixed.
|
||||
|
||||
### 2026-06-02 — caddy.service unit missing; recreated
|
||||
After the Slate AX → SODOLA network migration, Caddy was not listening (ports 80/443 dead). Root cause: the custom hubris1 Debian package (`caddy_1:2.11.3-hubris1_amd64`) does not ship a systemd service unit file. The unit had previously existed but was lost (likely on a package reinstall). Recreated at `/lib/systemd/system/caddy.service` with standard Caddy service config + `EnvironmentFile=/etc/caddy/caddy.env` (already present in `caddy.service.d/override.conf`). **Risk:** the unit will be lost again if the package is reinstalled without the file being tracked. Fix: add the service unit to the `caddy-conf` repo or rebuild the hubris1 package to include it.
|
||||
@@ -4,11 +4,10 @@ Docker host running the \*arr stack via [`ezarr`](https://github.com/ezarr/ezarr
|
||||
|
||||
## At a glance
|
||||
- **Hostname:** `arriman`
|
||||
- **IP:** `192.168.8.245`
|
||||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||||
- **IP:** `192.168.8.132`
|
||||
- **Privilege:** privileged
|
||||
- **Resources:** 4 cores / 8 GiB RAM / 24 GiB rootfs
|
||||
- **Mounts:** `/mnt/media_local` ↔ `/mnt/library`
|
||||
- **Mounts:** `/mnt/library` ↔ `/mnt/library`
|
||||
- **Public hostnames:** `jellyseerr` / `qbit` / `sab` (see below)
|
||||
|
||||
## Compose
|
||||
@@ -115,7 +114,7 @@ Member of [media GID 10000](../infrastructure/media-permissions.md). The LXC has
|
||||
|
||||
## Related
|
||||
- [Caddy (121)](121-caddy.md)
|
||||
- [Authentik (124)](106-auth-outpost.md) — forward-auth wiring + per-app `/api/*` bypass
|
||||
- [Authentik (124)](124-authentik.md) — forward-auth wiring + per-app `/api/*` bypass
|
||||
- [DNS](../infrastructure/dns.md)
|
||||
- [Media permissions](../infrastructure/media-permissions.md)
|
||||
- [Hubris host](../hosts/hubris.md)
|
||||
@@ -3,7 +3,7 @@
|
||||
> **This LXC was destroyed on 2026-06-04.** Replaced by Hermes Agent on mac-mini.
|
||||
> Monitoring migrated to `homelab-hardware-health` skill + 15-min Hermes cronjob.
|
||||
> Repos `dtoro/claudio-bot` and `dtoro/claudio-monitor` archived (read-only) on Gitea.
|
||||
> See [deprecation plan](../../../../plans/done/2026-06-04_130000-deprecate-claudio-bot.md) for full details.
|
||||
> See [deprecation plan](../plans/2026-06-04_130000-deprecate-claudio-bot.md) for full details.
|
||||
|
||||
Matrix-resident control plane. Bot account `@claudio:avispero` joined to a private room; accepts slash commands and natural language; relays infra notifications.
|
||||
|
||||
@@ -17,7 +17,7 @@ Matrix-resident control plane. Bot account `@claudio:avispero` joined to a priva
|
||||
|
||||
## Stack
|
||||
|
||||
Repo `dtoro/claudio-bot`, checkout at `/opt/claudio-bot`, systemd unit `claudio-bot.service`. Connects to Matrix at `http://192.168.8.239:8008` (direct LAN to [synapse (118)](../118-elementsynapse.md), avoids hairpin-NAT TLS issue on `matrix.hubris.network`).
|
||||
Repo `dtoro/claudio-bot`, checkout at `/opt/claudio-bot`, systemd unit `claudio-bot.service`. Connects to Matrix at `http://192.168.8.239:8008` (direct LAN to [synapse (118)](118-elementsynapse.md), avoids hairpin-NAT TLS issue on `matrix.hubris.network`).
|
||||
|
||||
Room: `!dEUVJArVKPorxHHJZK:avispero` (invite-only, `@dtoro:avispero` allowed).
|
||||
|
||||
@@ -43,8 +43,8 @@ Currently set to `lmstudio` → `google/gemma-4-e4b` on the Mac mini at `192.168
|
||||
## IPC
|
||||
|
||||
`http://192.168.8.230:9090/{notify,propose,status}`. Header `X-Bot-Token` must match `/etc/claudio-bot/ipc.token`. Used by:
|
||||
- [claudio-monitor on hubris](../../infrastructure/monitoring.md) for edge-triggered alerts (token in `/etc/claudio-monitor/bot.token`)
|
||||
- The (currently disabled) [restic backup wrapper](../../infrastructure/backups.md) (token in `/etc/restic/bot.token`)
|
||||
- [claudio-monitor on hubris](../infrastructure/monitoring.md) for edge-triggered alerts (token in `/etc/claudio-monitor/bot.token`)
|
||||
- The (currently disabled) [restic backup wrapper](../infrastructure/backups.md) (token in `/etc/restic/bot.token`)
|
||||
|
||||
> Token files at the source side **must hold the same value as `/etc/claudio-bot/ipc.token`** — rotate together.
|
||||
|
||||
@@ -61,13 +61,13 @@ Active plugins:
|
||||
|
||||
Push to `dtoro/claudio-bot` → gitea webhook → `http://192.168.8.230:9797/deploy` → pull + `pip install` + restart. Same shape as caddy-conf.
|
||||
|
||||
`app.ini` `ALLOWED_HOST_LIST` on [gitea](../104-gitea.md) includes `192.168.8.230`.
|
||||
`app.ini` `ALLOWED_HOST_LIST` on [gitea](104-gitea.md) includes `192.168.8.230`.
|
||||
|
||||
## Related
|
||||
- [elementsynapse (118)](../118-elementsynapse.md)
|
||||
- [Monitoring (claudio-monitor)](../../infrastructure/monitoring.md)
|
||||
- [Backups (disabled)](../../infrastructure/backups.md)
|
||||
- [Auto-deploy](../../infrastructure/auto-deploy.md)
|
||||
- [elementsynapse (118)](118-elementsynapse.md)
|
||||
- [Monitoring (claudio-monitor)](../infrastructure/monitoring.md)
|
||||
- [Backups (disabled)](../infrastructure/backups.md)
|
||||
- [Auto-deploy](../infrastructure/auto-deploy.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
@@ -84,7 +84,7 @@ Initial documentation.
|
||||
`backend: lmstudio` → `google/gemma-4-e4b` on the Mac mini. Anthropic key still present so the swap is reversible by flipping the config field.
|
||||
|
||||
### 2026-04-21 — `monitor` plugin added
|
||||
Receives events from [claudio-monitor](../../infrastructure/monitoring.md). Slash commands + tools registered. See `dtoro/claudio-bot` commit `e56da25`.
|
||||
Receives events from [claudio-monitor](../infrastructure/monitoring.md). Slash commands + tools registered. See `dtoro/claudio-bot` commit `e56da25`.
|
||||
|
||||
### 2026-04-20 — claudio-bot deployed
|
||||
LXC 123 provisioned. Repo, systemd unit, Matrix wiring, `system` + `backup` plugins, IPC server.
|
||||
@@ -1,7 +1,7 @@
|
||||
# 127 — `mule-photos-new`
|
||||
|
||||
Side-by-side **PhotoPrism M0 test** of the `dtoro/mule-image` `new` branch
|
||||
at `photos-new.hubris.network`. Production [LXC 120](../120-mule-images.md) keeps
|
||||
at `photos-new.hubris.network`. Production [LXC 120](120-mule-images.md) keeps
|
||||
running on the legacy stack at `photos.hubris.network` until M5 cutover.
|
||||
|
||||
## At a glance
|
||||
@@ -11,7 +11,7 @@ running on the legacy stack at `photos.hubris.network` until M5 cutover.
|
||||
- **Resources:** 6 cores / 8 GiB RAM / 40 GiB rootfs / 1 GiB swap
|
||||
- **Features:** `nesting=1,fuse=1,keyctl=1`
|
||||
- **Mounts:** *(none — see scratch copy below)*
|
||||
- **Public hostname:** [`photos-new.hubris.network`](../../infrastructure/dns.md) → [caddy (121)](../121-caddy.md) → split (PhotoPrism `:2342`, sidecar `:8000`, Vite `:5173`)
|
||||
- **Public hostname:** [`photos-new.hubris.network`](../infrastructure/dns.md) → [caddy (121)](121-caddy.md) → split (PhotoPrism `:2342`, sidecar `:8000`, Vite `:5173`)
|
||||
|
||||
## Stack (`/opt/mule-image`)
|
||||
|
||||
@@ -65,7 +65,7 @@ unprivileged LXCs can't see through.
|
||||
|
||||
## Auth — Authentik OIDC
|
||||
|
||||
PhotoPrism's "Sign in with OIDC" button delegates to [Authentik (124)](../106-auth-outpost.md).
|
||||
PhotoPrism's "Sign in with OIDC" button delegates to [Authentik (124)](124-authentik.md).
|
||||
|
||||
- **Provider/Application slug:** `mule-photos-new`
|
||||
- **Issuer:** `https://auth.hubris.network/application/o/mule-photos-new/`
|
||||
@@ -111,7 +111,7 @@ Mirrors the LXC 120 pattern.
|
||||
Push to the `new` branch on [git.hubris.network/dtoro/mule-image](http://git.hubris.network/dtoro/mule-image) → webhook fires → rebuild. The legacy LXC 120 watches `main` and is unaffected.
|
||||
|
||||
**Gitea gotcha:** the receiver IP must be in `[webhook] ALLOWED_HOST_LIST`
|
||||
in `/etc/gitea/app.ini` on [LXC 104](../104-gitea.md). LXC 127's
|
||||
in `/etc/gitea/app.ini` on [LXC 104](104-gitea.md). LXC 127's
|
||||
`192.168.8.181` was missing on first bring-up; every push delivered
|
||||
status 0 with the message `webhook can only call allowed HTTP servers`.
|
||||
Adding the IP and `systemctl restart gitea` is enough — same list is
|
||||
@@ -145,10 +145,10 @@ curl -sk --resolve photos-new.hubris.network:443:192.168.8.175 \
|
||||
```
|
||||
|
||||
> **Decommissioned 2026-05-22.** The PhotoPrism + sidecar + SvelteKit stack
|
||||
> validated here was promoted into production on [LXC 120](../120-mule-images.md)
|
||||
> validated here was promoted into production on [LXC 120](120-mule-images.md)
|
||||
> via the `Mulimage 2.0` merge (`dtoro/mule-image` commit `70dc1b6`). This
|
||||
> page is retained for archaeology; everything below is historic. See the
|
||||
> 2026-05-22 entry in [120-mule-images.md](../120-mule-images.md#changelog) for
|
||||
> 2026-05-22 entry in [120-mule-images.md](120-mule-images.md#changelog) for
|
||||
> the cutover detail.
|
||||
|
||||
## Changelog
|
||||
@@ -35,7 +35,7 @@ Not yet SOPS-enrolled. The poll token is set directly in `/etc/trmnl-plugins/env
|
||||
- [VPS ingress](../infrastructure/ingress.md) — public edge (cert mirror + traefik router)
|
||||
- [DNS (107)](107-dns.md) — Technitium A record `trmnl → 192.168.8.175` (LAN path via Caddy)
|
||||
- [Gitea (104)](104-gitea.md) — source repo `dtoro/terminalito`
|
||||
- [Plan: 2026-06-24 TRMNL plugins LXC](../../../plans/2026-06-24-trmnl-plugins-lxc.md)
|
||||
- [Plan: 2026-06-24 TRMNL plugins LXC](../plans/2026-06-24-trmnl-plugins-lxc.md)
|
||||
|
||||
## Changelog
|
||||
### 2026-06-24 — auto-deploy + LAN DNS wired
|
||||
@@ -5,8 +5,7 @@ Yuvomi family planner (formerly Oikos). Self-hosted family planner with 14 modul
|
||||
## At a glance
|
||||
|
||||
- **Hostname:** `house`
|
||||
- **IP:** `192.168.8.244`
|
||||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||||
- **IP:** `192.168.8.212` (static)
|
||||
- **Privilege:** unprivileged
|
||||
- **Resources:** 1 core / 1344 MiB RAM / 8 GiB rootfs (Debian 13)
|
||||
- **Mounts:** none
|
||||
@@ -40,7 +39,7 @@ Yuvomi family planner (formerly Oikos). Self-hosted family planner with 14 modul
|
||||
- [DNS (107)](107-dns.md) — Technitium A record `house → 192.168.8.175` (LAN path via Caddy)
|
||||
- [Paperless (103)](103-paperless.md) — native DMS connector (API at `:8000`)
|
||||
- [TRMNL (128)](128-trmnl.md) — Google Calendar tokens source
|
||||
- [Deployment plan](../../../plans/done/2026-06-25-yuvomi-deployment.md)
|
||||
- [Deployment plan](../plans/2026-06-25-yuvomi-deployment.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
@@ -5,18 +5,17 @@ Self-hosted digital library (eBooks, comics, audiobooks). Community fork/success
|
||||
## At a glance
|
||||
|
||||
- **Hostname:** `grimmory`
|
||||
- **IP:** `192.168.8.247`
|
||||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||||
- **IP:** `192.168.8.213` (static, set in PVE `net0` config — same pattern as all other LXCs)
|
||||
- **Privilege:** privileged (UID = host UID for `/mnt/library` media GID)
|
||||
- **Resources:** 1 core / 2 GiB RAM / 16 GiB rootfs (Debian 13)
|
||||
- **Mounts:** `/mnt/media_local` ↔ `/mnt/library`
|
||||
- **Mounts:** `/mnt/library`
|
||||
- **Public hostname:** `books.hubris.network`
|
||||
|
||||
## Service / port map
|
||||
|
||||
| Service | Listen | Notes |
|
||||
|---------|--------|-------|
|
||||
| Grimmory | `192.168.8.247:6060` | Docker Compose at `/opt/grimmory/` |
|
||||
| Grimmory | `192.168.8.213:6060` | Docker Compose at `/opt/grimmory/` |
|
||||
| MariaDB | internal only | Sidecar in the same compose stack |
|
||||
|
||||
## Compose
|
||||
@@ -44,7 +43,7 @@ Uses Confidential client (client secret stored in Grimmory's DB — migrated fro
|
||||
- **Client type:** Confidential (client secret in `oidc_provider_details` in MariaDB `app_settings`)
|
||||
- **Redirect URI:** `https://books.hubris.network/oauth2-callback`
|
||||
- **Scopes:** openid, profile, email, offline_access
|
||||
- **Back-channel logout:** `http://192.168.8.247:6060/api/v1/auth/oidc/backchannel-logout`
|
||||
- **Back-channel logout:** `http://192.168.8.213:6060/api/v1/auth/oidc/backchannel-logout`
|
||||
- **Application slug:** `booklore` → Issuer URI: `https://auth.hubris.network/application/o/booklore/`
|
||||
|
||||
## Media permissions
|
||||
@@ -54,8 +53,8 @@ LXC is privileged → in-container UID = host UID. Docker container gets media G
|
||||
## Related
|
||||
|
||||
- [apps (105)](105-apps.md) — previous host (Booklore)
|
||||
- [Caddy (121)](121-caddy.md) — `books.hubris.network → 192.168.8.247:6060`
|
||||
- [Authentik (124)](106-auth-outpost.md) — OIDC provider `Grimmory`
|
||||
- [Caddy (121)](121-caddy.md) — `books.hubris.network → 192.168.8.213:6060`
|
||||
- [Authentik (124)](124-authentik.md) — OIDC provider `Grimmory`
|
||||
- [DNS (107)](107-dns.md) — `books.hubris.network → 192.168.8.175` (unchanged from Booklore)
|
||||
- [Media permissions](../infrastructure/media-permissions.md)
|
||||
|
||||
112
containers/131-teddycloud.md
Normal file
112
containers/131-teddycloud.md
Normal file
@@ -0,0 +1,112 @@
|
||||
# 131 — `teddycloud`
|
||||
|
||||
Open-source replacement server for Toniebox smart audio devices (Tonieboxes). Serves device content and API on port 443 and exposes a management web UI at `teddy.hubris.network`.
|
||||
|
||||
## At a glance
|
||||
|
||||
- **Hostname:** `teddycloud`
|
||||
- **IP:** `192.168.8.243` (DHCP reservation; MAC `bc:24:11:11:7a:df`)
|
||||
- **Privilege:** privileged
|
||||
- **Resources:** 1 core / 1 GiB RAM / 16 GiB rootfs (Debian 12)
|
||||
- **Mounts:** `/mnt/library` (`mp0`) — TeddyCloud content at `/mnt/library/cloud/leon`
|
||||
- **Public hostname:** none (LAN-only)
|
||||
|
||||
## Role
|
||||
|
||||
Replaces the Boxine cloud (`prod.de.bb-online.com`) as the backend for Leon's Toniebox. Tonieboxes connect on port 443 using a custom CA cert issued by TeddyCloud. Content (Tonies) is stored on the NAS at `/mnt/library/cloud/leon` and is accessible from the management UI.
|
||||
|
||||
## Service / port map
|
||||
|
||||
| Service | Listen | Notes |
|
||||
|---------|--------|-------|
|
||||
| TeddyCloud device API | `0.0.0.0:443` | HTTPS, TeddyCloud self-signed CA, Toniebox connects here |
|
||||
| TeddyCloud HTTP | `0.0.0.0:80` | Redirects to 443 |
|
||||
| TeddyCloud web UI | `0.0.0.0:8443` | HTTPS management UI — fronted by Caddy at `teddy.hubris.network` (backend uses `tls_insecure_skip_verify` for self-signed cert on LAN hop) |
|
||||
|
||||
## Docker Compose
|
||||
|
||||
`/opt/teddycloud/docker-compose.yml`:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
teddycloud:
|
||||
image: ghcr.io/toniebox-reverse-engineering/teddycloud:latest
|
||||
ports:
|
||||
- "80:80"
|
||||
- "443:443"
|
||||
- "8443:8443"
|
||||
volumes:
|
||||
- certs:/teddycloud/certs
|
||||
- config:/teddycloud/config
|
||||
- /mnt/library/cloud/leon:/teddycloud/content
|
||||
- /mnt/library/cloud/leon:/teddycloud/library
|
||||
restart: unless-stopped
|
||||
|
||||
volumes:
|
||||
certs:
|
||||
config:
|
||||
```
|
||||
|
||||
`certs` and `config` are Docker named volumes (runtime state). `content` and `library` are bind-mounted from `/mnt/library/cloud/leon` so audio content persists across container rebuilds and is browsable from the host.
|
||||
|
||||
## Storage / config paths
|
||||
|
||||
- `/opt/teddycloud/docker-compose.yml` — compose file
|
||||
- Docker volume `teddycloud_certs` — TeddyCloud CA + server certs (generated on first boot)
|
||||
- Docker volume `teddycloud_config` — TeddyCloud config
|
||||
- `/mnt/library/cloud/leon/` — Tonie content + library (NAS bind mount)
|
||||
|
||||
## Networking
|
||||
|
||||
Two separate traffic paths — different IPs, no port 443 conflict:
|
||||
|
||||
**Management UI (browser):**
|
||||
```
|
||||
teddy.hubris.network → Technitium → 192.168.8.175 (Caddy) → 192.168.8.243:8443
|
||||
```
|
||||
|
||||
**Toniebox device traffic:**
|
||||
```
|
||||
prod.de.bb-online.com → Technitium override → 192.168.8.243:443 (TeddyCloud direct)
|
||||
```
|
||||
|
||||
Caddy terminates TLS for the management UI (IONOS DNS-01 wildcard cert). TeddyCloud terminates TLS for device traffic with its own self-signed CA — the Toniebox must have this CA installed.
|
||||
|
||||
### DNS overrides in Technitium
|
||||
|
||||
| Record | Type | Value | Purpose |
|
||||
|--------|------|-------|---------|
|
||||
| `teddy.hubris.network` | A | `192.168.8.175` | Management UI → Caddy (standard pattern) |
|
||||
| `prod.de.bb-online.com` | A | `192.168.8.243` | Toniebox device traffic → TeddyCloud direct |
|
||||
|
||||
The `prod.de.bb-online.com` override is Technitium-only — it intercepts Toniebox DNS locally without touching public DNS. The `dns-sync.py` cron on LXC 107 skips non-`hubris.network` records, so it stays local.
|
||||
|
||||
## Config notes
|
||||
|
||||
- `core.boxCertAuth=false` — client cert validation disabled. The box connects without presenting its unique client cert. Set in `/var/lib/docker/volumes/teddycloud_config/_data/config.ini` (TeddyCloud hot-reloads on change).
|
||||
- If you ever want per-box auth, flip to `true` and supply `certs/client/ca.der`, `client.der`, `private.der` extracted from the box flash.
|
||||
|
||||
## Toniebox onboarding — ESP32 SD card method
|
||||
|
||||
Leon's box is ESP32 generation. No hardware mod required.
|
||||
|
||||
1. Download the TeddyCloud CA cert from the web UI: **Security → CA Certificate → Download CA** (`ca.der`).
|
||||
2. Power off the Toniebox, remove the SD card.
|
||||
3. On the SD card, create folder `cert/` at the root.
|
||||
4. Copy the downloaded `ca.der` into `cert/ca.der` on the SD card.
|
||||
5. Reinsert SD card, power on the box.
|
||||
6. The box patches itself to trust TeddyCloud's CA, then resolves `prod.de.bb-online.com` via Technitium's override (`192.168.8.243`) and connects on port 443.
|
||||
|
||||
Reference: [upstream wiki — ESP32 SD card method](https://github.com/toniebox-reverse-engineering/teddycloud/wiki).
|
||||
|
||||
## Related
|
||||
|
||||
- [Caddy (121)](121-caddy.md) — LAN reverse proxy (`teddy.hubris.network → 192.168.8.243:8443`)
|
||||
- [DNS (107)](../infrastructure/dns.md) — Technitium A records for `teddy.hubris.network` and `prod.de.bb-online.com`
|
||||
- [Media permissions](../infrastructure/media-permissions.md) — NAS `/mnt/library` mount pattern
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-06-29 — provisioned
|
||||
|
||||
LXC 131 created (Debian 12, privileged, nesting=1). Docker installed. TeddyCloud running via Docker Compose at `/opt/teddycloud/`. Content bind-mounted from `/mnt/library/cloud/leon`. Caddy block added at `teddy.hubris.network → :8443`. Technitium A records: `teddy.hubris.network → 192.168.8.175` (Caddy), `prod.de.bb-online.com → 192.168.8.243` (device traffic direct).
|
||||
54
containers/index.md
Normal file
54
containers/index.md
Normal file
@@ -0,0 +1,54 @@
|
||||
# LXC containers — index
|
||||
|
||||
All containers live on [`hubris`](../hosts/hubris.md). Each row links to the per-container page.
|
||||
|
||||
| ID | Name | IP | Priv | Cores | RAM | Disk | Mounts | Public hostname | Status |
|
||||
| --- | ---------------- | --------------- | ---- | ----- | ----- | ----- | --------------------- | ------------------------------------- | -------- |
|
||||
| 101 | [jellyfin](101-jellyfin.md) | 192.168.8.206 | unpriv (idmap) | 2 | 4 GiB | 16 GiB | `/mnt/library` | `media.hubris.network` | running |
|
||||
| 103 | [paperless](103-paperless.md) | 192.168.8.130 | priv | 2 | 3 GiB | 8 GiB | `/mnt/library` | `paperless.hubris.network` | running |
|
||||
| 104 | [gitea](104-gitea.md) | 192.168.8.121 | priv | 1 | 1 GiB | 8 GiB | `/mnt/library` | `git.hubris.network` | running |
|
||||
| 105 | [apps](105-apps.md) | 192.168.8.205 | priv | 2 | 4 GiB | 30 GiB | `/mnt/library` | `docker` / `artifacto` / `blog` | running |
|
||||
| 114 | [nextcloud](114-nextcloud.md) | 192.168.8.224 | priv | 4 | 6 GiB | 25 GiB | `/mnt/library` | `cloud.hubris.network` | running |
|
||||
| 118 | [elementsynapse](118-elementsynapse.md) | 192.168.8.239 | unpriv | 1 | 2 GiB | 8 GiB | — | `matrix.hubris.network` | running |
|
||||
| 119 | [sophia](119-sophia.md) | 192.168.8.157 | priv | 2 | 1 GiB | 10 GiB | `/mnt/library` | — | running |
|
||||
| 120 | [mule-images](120-mule-images.md) | 192.168.8.136 | priv | 6 | 12 GiB | 60 GiB | `/mnt/library` + `/dev/dri` (iGPU passthrough) | `photos.hubris.network` | running |
|
||||
| 121 | [caddy](121-caddy.md) | 192.168.8.175 | unpriv | 1 | 512 MiB | 6 GiB | — | (terminates all `*.hubris.network`) | running |
|
||||
| 122 | [arriman](122-arriman.md) | 192.168.8.132 | priv | 4 | 8 GiB | 24 GiB | `/mnt/library` | `jellyseerr` / `qbit` / `sab` | running |
|
||||
| 124 | [authentik](124-authentik.md) | 192.168.8.180 | priv | 2 | 4 GiB | 20 GiB | — | `auth.hubris.network` | running |
|
||||
| 128 | [trmnl](128-trmnl.md) | 192.168.8.211 | unpriv | 1 | 768 MiB | 8 GiB | — | `trmnl.hubris.network` | running |
|
||||
| 129 | [house](129-house.md) | 192.168.8.212 | unpriv | 1 | 1344 MiB | 8 GiB | — | `house.hubris.network` | running |
|
||||
| 130 | [grimmory](130-grimmory.md) | 192.168.8.213 | priv | 1 | 2 GiB | 16 GiB | `/mnt/library` | `books.hubris.network` | running |
|
||||
| 131 | [teddycloud](131-teddycloud.md) | 192.168.8.243 | priv | 1 | 1 GiB | 16 GiB | `/mnt/library` | `teddy.hubris.network` (LAN only) | running |
|
||||
|
||||
## Recently destroyed (kept for archaeology)
|
||||
|
||||
| ID | Name | Destroyed | Reason |
|
||||
| --- | ---------------- | --------------- | --------------------------------------------- |
|
||||
| 127 | mule-photos-new | 2026-05-22 | PhotoPrism + sidecar + SvelteKit stack promoted to LXC 120 via Mulimage 2.0 merge (`70dc1b6`); M0 test LXC retired. Caddy + dnsmasq + gitea webhook + NC webhook listeners all cleaned up in the same cutover. |
|
||||
| 100 | arr (yunohost) | ~2026-04-28 | Migrated to docker stack on [arriman](122-arriman.md); planned retention window expired |
|
||||
| 106 | flaresolverr | ~2026-04-28 | Folded into the arriman docker compose |
|
||||
| 116 | heaper | 2026-05-14 | Decommissioned by user; data subtree at `/mnt/library/heaper` (224 MiB) retained |
|
||||
| 126 | plato | 2026-06-28 | Notes/discovery workspace decommissioned; data at `/mnt/library/documents/plato` retained for archaeology |
|
||||
| 123 | claudio-bot | 2026-06-04 | Replaced by Hermes Agent on mac-mini; monitoring migrated to `homelab-health-watchdog` cron. See [deprecation plan](../plans/2026-06-04_130000-deprecate-claudio-bot.md) |
|
||||
| 109 | syncthing | 2026-05-14 | Decommissioned by user; `/mnt/library/syncthing` was already empty |
|
||||
| 125 | seafile | 2026-05-13 | Seafile Pro evaluation, user disliked the product; teardown also removed `files.hubris.network` from caddy + dnsmasq |
|
||||
| 107 | marimo | between 2026-04-21 and 2026-04-28 | Decommissioned |
|
||||
| 110 | photoprism | between 2026-04-21 and 2026-04-28 | Replaced by [mulita](120-mule-images.md) |
|
||||
| 111 | karakeep | between 2026-04-21 and 2026-04-28 | Decommissioned |
|
||||
| 112 | immich | between 2026-04-21 and 2026-04-28 | Replaced by [mulita](120-mule-images.md) |
|
||||
| 115 | reticulum | between 2026-04-21 and 2026-04-28 | Decommissioned |
|
||||
|
||||
> Several `.conf.bak` files survive under `/etc/pve/lxc/` if you need to recover any of the configs.
|
||||
|
||||
## Conventions
|
||||
|
||||
- All net0 are `bridge=vmbr0`, `ip=dhcp` except [124 (authentik)](124-authentik.md) which is statically `192.168.8.180/24`. IPs are stable via the LAN router's DHCP reservations.
|
||||
- `onboot=1` on every container — the host brings them up after `pve-guests.service`.
|
||||
- Bind mounts are declared as `mp0: /mnt/library,mp=/mnt/library`. Containers that don't mount `/mnt/library` don't need it.
|
||||
- Most containers are privileged. Unprivileged ones (`101`, `118`, `121`) require an idmap block in their conf to participate in the [media GID 10000](../infrastructure/media-permissions.md) standard.
|
||||
|
||||
## Related
|
||||
- [Hubris host](../hosts/hubris.md)
|
||||
- [Media permissions](../infrastructure/media-permissions.md)
|
||||
- [Caddy](121-caddy.md) — terminates every public hostname
|
||||
- [DNS](../infrastructure/dns.md) — split-horizon entries for each subdomain
|
||||
@@ -1,70 +0,0 @@
|
||||
# Docker Compose for Oikos development
|
||||
# Usage: docker compose up -d postgres (just the DB)
|
||||
# make dev (full dev stack)
|
||||
|
||||
services:
|
||||
postgres:
|
||||
image: timescale/timescaledb:2.17.2-pg16
|
||||
environment:
|
||||
POSTGRES_DB: oikos
|
||||
POSTGRES_USER: oikos
|
||||
POSTGRES_PASSWORD: ${OIKOS_DB_PASSWORD:-oikos_dev}
|
||||
ports:
|
||||
- "5432:5432"
|
||||
volumes:
|
||||
- pg-data:/var/lib/postgresql/data
|
||||
healthcheck:
|
||||
test: ["CMD", "pg_isready", "-U", "oikos"]
|
||||
interval: 5s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
# One-shot: run migrations then exit
|
||||
migrate:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: compose/oikos/Dockerfile
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
environment:
|
||||
OIKOS_DATABASE_URL: postgres://oikos:${OIKOS_DB_PASSWORD:-oikos_dev}@postgres:5432/oikos?sslmode=disable
|
||||
command: ["migrate"]
|
||||
restart: "no"
|
||||
|
||||
# One-shot: ingest seeds then exit
|
||||
seed:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: compose/oikos/Dockerfile
|
||||
depends_on:
|
||||
migrate:
|
||||
condition: service_completed_successfully
|
||||
environment:
|
||||
OIKOS_DATABASE_URL: postgres://oikos:${OIKOS_DB_PASSWORD:-oikos_dev}@postgres:5432/oikos?sslmode=disable
|
||||
OIKOS_SEEDS_DIR: /app/seeds
|
||||
command: ["seed"]
|
||||
restart: "no"
|
||||
|
||||
# API server (Phase 2)
|
||||
api:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: compose/oikos/Dockerfile
|
||||
profiles: ["dev", "full"]
|
||||
depends_on:
|
||||
seed:
|
||||
condition: service_completed_successfully
|
||||
environment:
|
||||
OIKOS_DATABASE_URL: postgres://oikos:${OIKOS_DB_PASSWORD:-oikos_dev}@postgres:5432/oikos?sslmode=disable
|
||||
OIKOS_API_LISTEN: ":8090"
|
||||
OIKOS_ENV: dev
|
||||
OIKOS_DEBUG: "true"
|
||||
ports:
|
||||
- "8090:8090"
|
||||
command: ["api"]
|
||||
stop_signal: SIGTERM
|
||||
stop_grace_period: 30s
|
||||
|
||||
volumes:
|
||||
pg-data:
|
||||
@@ -1,23 +0,0 @@
|
||||
# ADR 0001 — Go with single-binary role packaging
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, R3-4
|
||||
|
||||
## Context
|
||||
The OS has three long-running roles (api, scheduler+actuator+learning,
|
||||
notifier) plus one-shot jobs (migrate, seed, export). Rev 2 planned three
|
||||
binaries with three Dockerfiles.
|
||||
|
||||
## Decision
|
||||
One Go binary `oikos` with role subcommands (`oikos api | scheduler |
|
||||
notifier | all | migrate | seed | export`), one multi-stage Dockerfile, one
|
||||
image tagged `oikos:<git-sha>`. Compose runs the image N times with
|
||||
different commands (Loki/Temporal pattern). Go over Python for static
|
||||
typing, small static binaries (CGO_ENABLED=0, distroless), and goroutines
|
||||
for concurrent probes.
|
||||
|
||||
## Consequences
|
||||
- One build, guaranteed version consistency across roles, trivial local dev
|
||||
(`oikos all`), simpler rollback (retag one image).
|
||||
- Full rewrite of ~4,400 Python lines (logic carries over per plan reuse map).
|
||||
- All roles share a dependency set; image is slightly larger than per-role
|
||||
minimal images — accepted.
|
||||
@@ -1,25 +0,0 @@
|
||||
# ADR 0002 — PostgreSQL + TimescaleDB as the only datastore
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3
|
||||
|
||||
## Context
|
||||
The OS needs a graph (entities/relationships), operational tables
|
||||
(signals/executions/approvals), a learning corpus, time-series metrics,
|
||||
audit and event logs. Alternatives: dedicated graph DB (Neo4j), dedicated
|
||||
TSDB (Prometheus/VictoriaMetrics), or one Postgres.
|
||||
|
||||
## Decision
|
||||
One PostgreSQL 16 instance with the TimescaleDB extension
|
||||
(timescale/timescaledb:2-pg16). Graph traversal via recursive CTEs
|
||||
(cycle-safe blast_radius); time-series via hypertables + continuous
|
||||
aggregates + retention policies; events via table + LISTEN/NOTIFY.
|
||||
|
||||
## Consequences
|
||||
- One backup/restore/DR story, one connection pool, transactional
|
||||
consistency between graph and operational writes (e.g. event emission in
|
||||
the same transaction as state change).
|
||||
- Postgres is the accepted SPOF — mitigated by daily pg_dump + WAL PITR +
|
||||
off-host copies + monthly restore drills; streaming replication is the
|
||||
future path if needed.
|
||||
- Homelab graph scale (hundreds of nodes) is far below where a dedicated
|
||||
graph DB pays for itself.
|
||||
@@ -1,23 +0,0 @@
|
||||
# ADR 0003 — DB-native ontology with YAML seed manifests
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, R3-1
|
||||
|
||||
## Context
|
||||
Rev 1 kept inventory/ontology/policy as YAML files parsed at runtime.
|
||||
Agents need graph queries (blast radius), transactional mutations with
|
||||
audit, and a future UI needs to edit the model without file round-trips.
|
||||
|
||||
## Decision
|
||||
The DB is the runtime source of truth. entity_types form an is-a hierarchy
|
||||
(parent_type, is_abstract); relationship endpoint constraints may name
|
||||
abstract types and validation walks the hierarchy. YAML files under seeds/
|
||||
bootstrap the DB (idempotent, content-hashed via seed_versions) and serve
|
||||
DR; `GET /api/v1/export` regenerates them for version control (round-trip
|
||||
byte-stable, tested in CI).
|
||||
|
||||
## Consequences
|
||||
- Ontology changes are API calls (policy-gated), not redeploys.
|
||||
- Seeds can drift from DB between exports — export is part of the routine
|
||||
(commit after meaningful model edits).
|
||||
- Abstract types let policy rules and relationships bind once at the right
|
||||
altitude (e.g. `compute-entity provides service`).
|
||||
@@ -1,21 +0,0 @@
|
||||
# ADR 0004 — Contract-first OpenAPI API
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, R3-2/R3-3
|
||||
|
||||
## Context
|
||||
Future UIs, a CLI client, and an MCP surface must stay in sync with the
|
||||
API. Code-first (Gin + generated docs) drifts.
|
||||
|
||||
## Decision
|
||||
api/openapi.yaml (OpenAPI 3.1) is the source of truth. Server stubs via
|
||||
oapi-codegen (strict server, chi router); clients generated for Go (CLI)
|
||||
and TypeScript (future UI). Conventions: RFC 9457 problem+json errors,
|
||||
{items, next_cursor} envelopes, cursor pagination, Idempotency-Key on
|
||||
unsafe POSTs, ETag/If-Match optimistic concurrency, scopes
|
||||
(operator/viewer/agent) annotated per operation. CI fails on spec/handler
|
||||
drift. MCP tools wrap the same service layer.
|
||||
|
||||
## Consequences
|
||||
- UI development needs only the running API (spec served at /openapi.yaml).
|
||||
- Handler changes require spec changes first — deliberate friction.
|
||||
- Breaking changes ship as /api/v2 side by side; v1 is additive-only.
|
||||
@@ -1,18 +0,0 @@
|
||||
# ADR 0005 — UUIDv7 + slug entity identity
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, R3-5 (resolves audit D1)
|
||||
|
||||
## Context
|
||||
Rev 2 used TEXT primary keys ('host:hubris') — renames break FKs, and
|
||||
date-string signal IDs are race-prone.
|
||||
|
||||
## Decision
|
||||
Primary keys are UUIDv7 (time-ordered, generated in Go). Every entity also
|
||||
carries a unique human slug ('host:hubris'); (type, name) is unique too.
|
||||
The API accepts UUID or slug everywhere; slugs may change (rename), UUIDs
|
||||
never do.
|
||||
|
||||
## Consequences
|
||||
- Renames are metadata updates; history and edges survive.
|
||||
- UUIDv7's time-ordering keeps B-tree inserts append-mostly.
|
||||
- Seeds and exports use slugs (human-diffable); ingest resolves to UUIDs.
|
||||
@@ -1,23 +0,0 @@
|
||||
# ADR 0006 — Learning is proposal-only (no self-authorization)
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3 (resolves audit S3/S4/SA2)
|
||||
|
||||
## Context
|
||||
The learning loop (feedback → patterns → skills) informs the classifier
|
||||
that decides auto-act vs escalate. If learning could expand its own
|
||||
autonomy, poisoned feedback (flapping services, biased probes) could
|
||||
unlock destructive auto-act.
|
||||
|
||||
## Decision
|
||||
The learning engine cannot write to governance (policy/autonomy) tables —
|
||||
enforced structurally: its DB role has no grants on them. Pattern
|
||||
activation (validated → active) and any autonomy expansion require operator
|
||||
approval. Confidence is the Wilson lower bound capped by evidence_count/5;
|
||||
anomalous feedback bursts quarantine the pattern; no skill ever
|
||||
auto-promotes an action into destructive autonomy (hard-coded). Lowering
|
||||
autonomy (kill-switch) is always immediate, never gated.
|
||||
|
||||
## Consequences
|
||||
- Cold start is slow by design — the agent escalates until trust is earned.
|
||||
- The operator is the only path to more autonomy; the audit trail shows
|
||||
every grant.
|
||||
@@ -1,27 +0,0 @@
|
||||
# ADR 0007 — Threat model and trust zones
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, Security model section
|
||||
|
||||
## Context
|
||||
The control plane can restart services and (eventually) mutate config
|
||||
fleet-wide. Compromise of any one container must not equal compromise of
|
||||
the fleet.
|
||||
|
||||
## Decision
|
||||
Trust zones as Docker networks: net-front (Caddy→api only), net-data
|
||||
(Postgres), net-ops (SSH egress, actuator only). Hermes holds no SSH keys;
|
||||
the actuator uses a restricted key (command=/from= in authorized_keys)
|
||||
until the /executions gateway fully brokers actions. Caddy is an explicit
|
||||
trust root but the API independently validates OIDC JWTs — network origin
|
||||
is defense-in-depth, never the auth (this enables the LAN break-glass API
|
||||
binding; the Hermes gateway remains mesh-only). Policy changes are
|
||||
dual-controlled with before/after hash auditing and a startup
|
||||
hash-vs-known-good check. Approval tokens are single-use HMAC, hashed at
|
||||
rest, TTL-bound.
|
||||
|
||||
## Consequences
|
||||
- Documented residual risks: plaintext LAN break-glass hop (emergency use),
|
||||
Postgres as shared dependency of all roles, macOS host itself unmanaged
|
||||
by the OS.
|
||||
- Rotation cadences: actuator SSH key 6mo, machine tokens 90d, webhook
|
||||
HMAC 1y — scheduler raises expiry signals 2 weeks ahead.
|
||||
@@ -1,20 +0,0 @@
|
||||
# ADR 0008 — Forward-only migrations
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3 (resolves audit D5/O1)
|
||||
|
||||
## Context
|
||||
Down-migrations are rarely tested and lie about reversibility once data
|
||||
has flowed. Rollback needs a strategy that works with real data.
|
||||
|
||||
## Decision
|
||||
golang-migrate, embedded (//go:embed), up-only. Migrations run in a
|
||||
one-shot init container with a DDL-only DB user before app roles start.
|
||||
Within one deploy window migrations are additive-only (new columns
|
||||
nullable, new tables optional) so previous-SHA images tolerate the new
|
||||
schema. Rollback = redeploy previous image tag; if the migration itself is
|
||||
the problem, pg_restore the automatic pre-deploy dump. Mistakes roll
|
||||
forward via compensating migrations.
|
||||
|
||||
## Consequences
|
||||
- No down.sql to write or test; the pre-deploy dump is the real safety net.
|
||||
- Destructive schema changes (drop/rename) take two deploys by design.
|
||||
@@ -1,19 +0,0 @@
|
||||
# ADR 0009 — SSE over WebSocket for the event stream
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, R3-14
|
||||
|
||||
## Context
|
||||
Live updates (signals, executions, approvals) push server→client only.
|
||||
Rev 2 specified WebSocket.
|
||||
|
||||
## Decision
|
||||
Server-Sent Events at GET /api/v1/events/stream: plain HTTP (proxies
|
||||
through Caddy without upgrade handling), native browser EventSource with
|
||||
auto-reconnect, Last-Event-ID resume backed by the events table. Bounded
|
||||
per-subscriber buffers with drop-oldest; heartbeat comments every 15s.
|
||||
Delivery is best-effort — GET /events backfills. Transactional emission +
|
||||
post-commit LISTEN/NOTIFY feed the stream.
|
||||
|
||||
## Consequences
|
||||
- No bidirectional channel; if one is ever needed (interactive terminals),
|
||||
add WebSocket alongside — this ADR covers the event feed only.
|
||||
@@ -1,20 +0,0 @@
|
||||
# ADR 0010 — Infisical secrets with SOPS DR fallback
|
||||
|
||||
Status: accepted (2026-07-07) · Plan: rev 3, Phase 5 (resolves audit S9)
|
||||
|
||||
## Context
|
||||
SOPS+age is file-based: no runtime API, no machine identities, no
|
||||
rotation tracking, and every consumer needs the age key.
|
||||
|
||||
## Decision
|
||||
Infisical in the Docker stack; services fetch via machine identities;
|
||||
secrets never in env files or plain config (config hierarchy: defaults →
|
||||
file → env → Infisical, secrets only). Bootstrap root of trust: Infisical
|
||||
master key in the mac-mini Keychain, backed up offline. One age key is
|
||||
retained and all secrets are exported to a SOPS-encrypted fallback file
|
||||
until an Infisical restore drill has passed; the fallback is refreshed on
|
||||
rotation.
|
||||
|
||||
## Consequences
|
||||
- Chicken-and-egg is explicit: the Keychain + offline copy are the root.
|
||||
- SOPS retirement is gated on a passed restore drill, not on the calendar.
|
||||
@@ -1,18 +0,0 @@
|
||||
# Architecture Decision Records
|
||||
|
||||
MADR-style records for Oikos. One decision per file, numbered, never edited
|
||||
after acceptance — superseding decisions get a new ADR that links back.
|
||||
Statuses: proposed | accepted | superseded-by-NNNN.
|
||||
|
||||
| ADR | Title |
|
||||
|---|---|
|
||||
| [0001](0001-go-single-binary.md) | Go with single-binary role packaging |
|
||||
| [0002](0002-postgres-timescale-only-datastore.md) | PostgreSQL + TimescaleDB as the only datastore |
|
||||
| [0003](0003-db-native-ontology-yaml-seeds.md) | DB-native ontology with YAML seed manifests |
|
||||
| [0004](0004-openapi-first.md) | Contract-first OpenAPI API |
|
||||
| [0005](0005-uuidv7-plus-slug-identity.md) | UUIDv7 + slug entity identity |
|
||||
| [0006](0006-learning-proposal-only.md) | Learning is proposal-only (no self-authorization) |
|
||||
| [0007](0007-threat-model.md) | Threat model and trust zones |
|
||||
| [0008](0008-forward-only-migrations.md) | Forward-only migrations |
|
||||
| [0009](0009-sse-over-websocket.md) | SSE over WebSocket for the event stream |
|
||||
| [0010](0010-infisical-with-sops-fallback.md) | Infisical secrets with SOPS DR fallback |
|
||||
36
go.mod
36
go.mod
@@ -1,36 +0,0 @@
|
||||
module github.com/dtoro/oikos
|
||||
|
||||
go 1.26.3
|
||||
|
||||
require (
|
||||
github.com/getkin/kin-openapi v0.140.0
|
||||
github.com/go-chi/chi/v5 v5.3.1
|
||||
github.com/golang-jwt/jwt/v5 v5.3.1
|
||||
github.com/google/jsonschema-go v0.4.3
|
||||
github.com/google/uuid v1.6.0
|
||||
github.com/jackc/pgx/v5 v5.10.0
|
||||
github.com/modelcontextprotocol/go-sdk v1.6.1
|
||||
github.com/oapi-codegen/runtime v1.4.2
|
||||
golang.org/x/crypto v0.53.0
|
||||
golang.org/x/sync v0.21.0
|
||||
gopkg.in/yaml.v3 v3.0.1
|
||||
)
|
||||
|
||||
require (
|
||||
github.com/apapsch/go-jsonmerge/v2 v2.0.0 // indirect
|
||||
github.com/go-openapi/jsonpointer v0.22.5 // indirect
|
||||
github.com/go-openapi/swag/jsonname v0.25.5 // indirect
|
||||
github.com/jackc/pgpassfile v1.0.0 // indirect
|
||||
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
|
||||
github.com/jackc/puddle/v2 v2.2.2 // indirect
|
||||
github.com/oasdiff/yaml v0.1.0 // indirect
|
||||
github.com/oasdiff/yaml3 v0.0.13 // indirect
|
||||
github.com/rogpeppe/go-internal v1.15.0 // indirect
|
||||
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 // indirect
|
||||
github.com/segmentio/asm v1.1.3 // indirect
|
||||
github.com/segmentio/encoding v0.5.4 // indirect
|
||||
github.com/yosida95/uritemplate/v3 v3.0.2 // indirect
|
||||
golang.org/x/oauth2 v0.35.0 // indirect
|
||||
golang.org/x/sys v0.46.0 // indirect
|
||||
golang.org/x/text v0.38.0 // indirect
|
||||
)
|
||||
88
go.sum
88
go.sum
@@ -1,88 +0,0 @@
|
||||
github.com/RaveNoX/go-jsoncommentstrip v1.0.0/go.mod h1:78ihd09MekBnJnxpICcwzCMzGrKSKYe4AqU6PDYYpjk=
|
||||
github.com/apapsch/go-jsonmerge/v2 v2.0.0 h1:axGnT1gRIfimI7gJifB699GoE/oq+F2MU7Dml6nw9rQ=
|
||||
github.com/apapsch/go-jsonmerge/v2 v2.0.0/go.mod h1:lvDnEdqiQrp0O42VQGgmlKpxL1AP2+08jFMw88y4klk=
|
||||
github.com/bmatcuk/doublestar v1.1.1/go.mod h1:UD6OnuiIn0yFxxA2le/rnRU1G4RaI4UvFv1sNto9p6w=
|
||||
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
|
||||
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||
github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI=
|
||||
github.com/dlclark/regexp2 v1.11.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8=
|
||||
github.com/getkin/kin-openapi v0.140.0 h1:JFn675aXRFjyiZKa/BFWploGldQlI0gobp4J5k0EZ2g=
|
||||
github.com/getkin/kin-openapi v0.140.0/go.mod h1:lISrB64F0CPcuDJ3LdtPTMJBY8VENjR9wJBdrcT6J3g=
|
||||
github.com/go-chi/chi/v5 v5.3.1 h1:3j4HZLGZQ3JpMCrPJF/Jl3mYJfWLKBfNJ6quurUGCf8=
|
||||
github.com/go-chi/chi/v5 v5.3.1/go.mod h1:R+tYY2hNuVUUjxoPtqUdgBqevM9s9njzkTLutVsOCto=
|
||||
github.com/go-openapi/jsonpointer v0.22.5 h1:8on/0Yp4uTb9f4XvTrM2+1CPrV05QPZXu+rvu2o9jcA=
|
||||
github.com/go-openapi/jsonpointer v0.22.5/go.mod h1:gyUR3sCvGSWchA2sUBJGluYMbe1zazrYWIkWPjjMUY0=
|
||||
github.com/go-openapi/swag/jsonname v0.25.5 h1:8p150i44rv/Drip4vWI3kGi9+4W9TdI3US3uUYSFhSo=
|
||||
github.com/go-openapi/swag/jsonname v0.25.5/go.mod h1:jNqqikyiAK56uS7n8sLkdaNY/uq6+D2m2LANat09pKU=
|
||||
github.com/go-openapi/testify/v2 v2.4.0 h1:8nsPrHVCWkQ4p8h1EsRVymA2XABB4OT40gcvAu+voFM=
|
||||
github.com/go-openapi/testify/v2 v2.4.0/go.mod h1:HCPmvFFnheKK2BuwSA0TbbdxJ3I16pjwMkYkP4Ywn54=
|
||||
github.com/golang-jwt/jwt/v5 v5.3.1 h1:kYf81DTWFe7t+1VvL7eS+jKFVWaUnK9cB1qbwn63YCY=
|
||||
github.com/golang-jwt/jwt/v5 v5.3.1/go.mod h1:fxCRLWMO43lRc8nhHWY6LGqRcf+1gQWArsqaEUEa5bE=
|
||||
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
|
||||
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
|
||||
github.com/google/jsonschema-go v0.4.3 h1:/DBOLZTfDow7pe2GmaJNhltueGTtDKICi8V8p+DQPd0=
|
||||
github.com/google/jsonschema-go v0.4.3/go.mod h1:r5quNTdLOYEz95Ru18zA0ydNbBuYoo9tgaYcxEYhJVE=
|
||||
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
|
||||
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
|
||||
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
|
||||
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
|
||||
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
|
||||
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
|
||||
github.com/jackc/pgx/v5 v5.10.0 h1:VhSvgU2jSli8o3AqIEOTJr7rZwAEUVo4E4XhR94Zfr0=
|
||||
github.com/jackc/pgx/v5 v5.10.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
|
||||
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
|
||||
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
|
||||
github.com/juju/gnuflag v0.0.0-20171113085948-2ce1bb71843d/go.mod h1:2PavIy+JPciBPrBUjwbNvtwB6RQlve+hkpll6QSNmOE=
|
||||
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0=
|
||||
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk=
|
||||
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
|
||||
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
|
||||
github.com/modelcontextprotocol/go-sdk v1.6.1 h1:0zOSupjKUxPKSocPT1Wtago+mUHU2/uZ4xSOY0FGReU=
|
||||
github.com/modelcontextprotocol/go-sdk v1.6.1/go.mod h1:kzm3kzFL1/+AziGOE0nUs3gvPoNxMCvkxokMkuFapXQ=
|
||||
github.com/oapi-codegen/nullable v1.1.0 h1:eAh8JVc5430VtYVnq00Hrbpag9PFRGWLjxR1/3KntMs=
|
||||
github.com/oapi-codegen/nullable v1.1.0/go.mod h1:KUZ3vUzkmEKY90ksAmit2+5juDIhIZhfDl+0PwOQlFY=
|
||||
github.com/oapi-codegen/runtime v1.4.2 h1:GMxFVYLzoYLua+/KvzgSphkyK1lLTReQI9Vf4hvATKE=
|
||||
github.com/oapi-codegen/runtime v1.4.2/go.mod h1:GwV7hC2hviaMzj+ITfHVRESK5J2W/GefVwIND/bMGvU=
|
||||
github.com/oasdiff/yaml v0.1.0 h1:0bqZjfKc/8S9urj4JuwepX41WX9EoA6ifhU3SV06cXg=
|
||||
github.com/oasdiff/yaml v0.1.0/go.mod h1:kOlRmMdL2X3vucLCEQO5u61SU22RysnfXvcttrZA1O0=
|
||||
github.com/oasdiff/yaml3 v0.0.13 h1:06svmvOHOVBqF81+sY2EUScvUI/iS/vl2VIeUUxZQwg=
|
||||
github.com/oasdiff/yaml3 v0.0.13/go.mod h1:y5+oSEHCPT/DGrS++Wc/479ERge0zTFxaF8PbGKcg2o=
|
||||
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
|
||||
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
||||
github.com/rogpeppe/go-internal v1.15.0 h1:D0RCU5rMAp+SpgkiNdrjfJ+LX4J1M32V2NeCY7EJ6hc=
|
||||
github.com/rogpeppe/go-internal v1.15.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
|
||||
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 h1:KRzFb2m7YtdldCEkzs6KqmJw4nqEVZGK7IN2kJkjTuQ=
|
||||
github.com/santhosh-tekuri/jsonschema/v6 v6.0.2/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU=
|
||||
github.com/segmentio/asm v1.1.3 h1:WM03sfUOENvvKexOLp+pCqgb/WDjsi7EK8gIsICtzhc=
|
||||
github.com/segmentio/asm v1.1.3/go.mod h1:Ld3L4ZXGNcSLRg4JBsZ3//1+f/TjYl0Mzen/DQy1EJg=
|
||||
github.com/segmentio/encoding v0.5.4 h1:OW1VRern8Nw6ITAtwSZ7Idrl3MXCFwXHPgqESYfvNt0=
|
||||
github.com/segmentio/encoding v0.5.4/go.mod h1:HS1ZKa3kSN32ZHVZ7ZLPLXWvOVIiZtyJnO1gPH1sKt0=
|
||||
github.com/spkg/bom v0.0.0-20160624110644-59b7046e48ad/go.mod h1:qLr4V1qq6nMqFKkMo8ZTx3f+BZEkzsRUY10Xsm2mwU0=
|
||||
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
|
||||
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
|
||||
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
|
||||
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
|
||||
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
|
||||
github.com/yosida95/uritemplate/v3 v3.0.2 h1:Ed3Oyj9yrmi9087+NczuL5BwkIc4wvTb5zIM+UJPGz4=
|
||||
github.com/yosida95/uritemplate/v3 v3.0.2/go.mod h1:ILOh0sOhIJR3+L/8afwt/kE++YT040gmv5BQTMR2HP4=
|
||||
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
|
||||
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
|
||||
golang.org/x/oauth2 v0.35.0 h1:Mv2mzuHuZuY2+bkyWXIHMfhNdJAdwW3FuWeCPYN5GVQ=
|
||||
golang.org/x/oauth2 v0.35.0/go.mod h1:lzm5WQJQwKZ3nwavOZ3IS5Aulzxi68dUSgRHujetwEA=
|
||||
golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
|
||||
golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
|
||||
golang.org/x/sys v0.46.0 h1:noSf2Fq6F8DBgS+LysIkx7rIExoNHJsxOAtPp4rthXw=
|
||||
golang.org/x/sys v0.46.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
|
||||
golang.org/x/term v0.44.0 h1:0rLvDRCtNj0gZkyIXhCyOb2OAzEhLVqc4B+hrsBhrmc=
|
||||
golang.org/x/term v0.44.0/go.mod h1:7ze4MdzUzLXpSAoFP1H0bOI9aXDqveSvatT5vKcFh2Y=
|
||||
golang.org/x/text v0.38.0 h1:sXmwo9DwP3OK9EZ7PqAdaooSGozfl/3a6/xJcbzPRhE=
|
||||
golang.org/x/text v0.38.0/go.mod h1:YXZt3QhHUKYT53r2lLKFIVi6Ao1jdzrTR/KQ09qyxF4=
|
||||
golang.org/x/tools v0.45.0 h1:18qN3FAooORvApf5XjCXgsuayZOEtXf6JK18I3+ONa8=
|
||||
golang.org/x/tools v0.45.0/go.mod h1:LuUGqqaXcXMEFEruIVJVm5mgDD8vww/z/SR1gQ4uE/0=
|
||||
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
|
||||
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
|
||||
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
|
||||
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
||||
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
|
||||
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
||||
@@ -5,7 +5,6 @@ name: apps
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: docker-apps
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 105
|
||||
lan_ip: 192.168.8.205
|
||||
@@ -34,29 +33,23 @@ services_hosted:
|
||||
- name: artifacto
|
||||
backend: apps
|
||||
url: https://artifacto.hubris.network
|
||||
doc_page: knowledge/wiki/containers/105-apps.md
|
||||
config_repo: dtoro/Artifacto
|
||||
- name: homelab_mcp
|
||||
backend: apps
|
||||
port: 9810
|
||||
systemd_unit: homelab-mcp
|
||||
public_host: mcp.hubris.network
|
||||
endpoint: https://mcp.hubris.network/mcp
|
||||
doc_page: knowledge/wiki/infrastructure/homelab-context.md
|
||||
config_repo: dtoro/Homelab-Docs
|
||||
note: MCP server. Read-only context + management. Reachable on the LAN via Caddy and from off-LAN via
|
||||
Netbird (192.168.8.0/24 is a network resource routed through hubris).
|
||||
risk_notes: "agents' primary read surface \u2014 outage degrades every agent to grepping the clone"
|
||||
- name: secrets_issuance
|
||||
backend: apps
|
||||
port: 9820
|
||||
systemd_unit: secrets-issuance
|
||||
public_host: secrets.hubris.network
|
||||
endpoint: https://secrets.hubris.network/issue
|
||||
doc_page: .agents/operations/agent-enrollment.md
|
||||
config_repo: dtoro/Homelab-Docs
|
||||
note: Issues per-client age private keys. Gated at source-IP layer (mesh + LAN subnets in MESH_SUBNETS).
|
||||
risk_notes: "identity issuance \u2014 any change is security-sensitive; key operations are destructive-class"
|
||||
age_pubkey: age1duyl8mkpgu80uv934dy8q7enqjms6yvdz264hme8uryuxmvvqesq6rusq0
|
||||
see_also:
|
||||
- containers/105-apps.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,10 +5,9 @@ name: arriman
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: arr-stack
|
||||
state: active
|
||||
host: strong
|
||||
host: hubris
|
||||
pve_id: 122
|
||||
lan_ip: 192.168.8.245
|
||||
lan_ip: 192.168.8.132
|
||||
mesh:
|
||||
tailscale:
|
||||
fqdn: arr
|
||||
@@ -18,7 +17,7 @@ mesh_globals:
|
||||
- netbird
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/media_local
|
||||
- /mnt/library
|
||||
public_hosts:
|
||||
- jellyseerr.hubris.network
|
||||
- qbit.hubris.network
|
||||
@@ -29,10 +28,7 @@ services_hosted:
|
||||
- name: arr_stack
|
||||
backend: arriman
|
||||
note: jellyseerr / qbit / sab on docker compose
|
||||
doc_page: knowledge/wiki/containers/122-arriman.md
|
||||
notes:
|
||||
- Migrated from hubris to strong 2026-07-05 (Phase 2). Library on ludo-lvm.
|
||||
- Contains homarr, radarr, sonarr, lidarr, sabnzbd, qbittorrent, bazarr, flaresolverr, prowlarr, jellyseerr
|
||||
- qBittorrent auth subnet whitelist expanded to 192.168.8.0/24 (for seanime + Caddy access)
|
||||
see_also:
|
||||
- containers/122-arriman.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: auth-outpost
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: authentik-gateway
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 106
|
||||
lan_ip: 192.168.8.6
|
||||
@@ -16,5 +15,7 @@ mesh_globals:
|
||||
- tailscale
|
||||
notes:
|
||||
- Runs Authentik outpost (reverse-proxy/SSO enforcement) for protected services
|
||||
see_also:
|
||||
- containers/106-auth-outpost.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: caddy
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: reverse-proxy
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 121
|
||||
lan_ip: 192.168.8.175
|
||||
@@ -24,12 +23,10 @@ services_hosted:
|
||||
backend: caddy
|
||||
role: reverse-proxy
|
||||
note: terminates all *.hubris.network
|
||||
doc_page: knowledge/wiki/containers/121-caddy.md
|
||||
config_repo: dtoro/caddy-conf
|
||||
risk_notes: "wide blast radius \u2014 every *.hubris.network route rides on it (see oikos/policy.yaml\
|
||||
\ service_overrides)"
|
||||
notes:
|
||||
- Terminates all *.hubris.network
|
||||
- /etc/caddy is a git checkout of dtoro/caddy-conf
|
||||
see_also:
|
||||
- containers/121-caddy.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: dns
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: dns-server
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 107
|
||||
lan_ip: 192.168.8.2
|
||||
@@ -15,15 +14,17 @@ mesh_globals:
|
||||
- netbird
|
||||
- tailscale
|
||||
runs:
|
||||
- dns
|
||||
- authentik
|
||||
services_hosted:
|
||||
- name: dns
|
||||
- name: authentik
|
||||
url: https://auth.hubris.network
|
||||
backend: dns
|
||||
dns: null
|
||||
note: Technitium DNS, split-horizon zone
|
||||
doc_page: knowledge/wiki/containers/107-dns.md
|
||||
risk_notes: "LAN-wide resolver \u2014 misconfig breaks name resolution for every client"
|
||||
notes:
|
||||
- Technitium DNS, split-horizon zone for *.hubris.network
|
||||
- Primary DNS for 192.168.8.0/24 LAN (inventory.services.dns references this)
|
||||
see_also:
|
||||
- containers/107-dns.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,12 +5,12 @@ name: elementsynapse
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: matrix-server
|
||||
state: active
|
||||
host: strong
|
||||
host: hubris
|
||||
pve_id: 118
|
||||
lan_ip: 192.168.8.242
|
||||
lan_ip: 192.168.8.239
|
||||
mesh:
|
||||
tailscale: {}
|
||||
tailscale:
|
||||
fqdn: elementsynapse
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
@@ -23,9 +23,7 @@ services_hosted:
|
||||
- name: matrix
|
||||
url: https://matrix.hubris.network
|
||||
backend: elementsynapse
|
||||
doc_page: knowledge/wiki/containers/118-elementsynapse.md
|
||||
risk_notes: "alert/approval channel for Oikos \u2014 outage silences agent escalation"
|
||||
notes:
|
||||
- Migrated from hubris to strong 2026-07-05 (Phase 1 of strong migration plan).
|
||||
see_also:
|
||||
- containers/118-elementsynapse.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: gitea
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: git-server
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 104
|
||||
lan_ip: 192.168.8.121
|
||||
@@ -27,10 +26,9 @@ services_hosted:
|
||||
url: https://git.hubris.network
|
||||
backend: gitea
|
||||
backend_url: http://192.168.8.121:3000
|
||||
doc_page: knowledge/wiki/containers/104-gitea.md
|
||||
config_repo: dtoro/gitea-customizations
|
||||
risk_notes: hosts all config repos + deploy webhooks; outage blocks auto-deploy and sync
|
||||
notes:
|
||||
- Bare repos live at /mnt/library/repos/dtoro/*.git
|
||||
see_also:
|
||||
- containers/104-gitea.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -1,25 +0,0 @@
|
||||
# Generated by mcp/build_host_files.py from inventory.yaml.
|
||||
# Do NOT edit by hand — your changes will be overwritten.
|
||||
# Source of truth: ../inventory.yaml
|
||||
name: grimmory
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: book-library
|
||||
state: active
|
||||
host: strong
|
||||
pve_id: 130
|
||||
lan_ip: 192.168.8.247
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/media_local
|
||||
public_host: books.hubris.network
|
||||
notes:
|
||||
- Docker host for Grimmory (community fork of Booklore). Created 2026-06-29.
|
||||
- Migrated from hubris to strong 2026-07-05 (Phase 2d). Books on ludo-lvm.
|
||||
age_pubkey: age1uellsemnjrzgfg9fxw4jefpy05laxzggwnwhh6ny3wl7alyp6v8q0muxet
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
@@ -5,7 +5,6 @@ name: haos
|
||||
kind: vm
|
||||
os: linux
|
||||
role: home-automation
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 108
|
||||
lan_ip: 192.168.8.101
|
||||
@@ -22,6 +21,7 @@ runs:
|
||||
services_hosted:
|
||||
- name: haos
|
||||
backend: haos
|
||||
doc_page: knowledge/wiki/vms/108-haos.md
|
||||
see_also:
|
||||
- vms/108-haos.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,10 +5,9 @@ name: house
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: family-planner
|
||||
state: active
|
||||
host: strong
|
||||
host: hubris
|
||||
pve_id: 129
|
||||
lan_ip: 192.168.8.244
|
||||
lan_ip: 192.168.8.212
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
@@ -17,10 +16,9 @@ mesh_globals:
|
||||
public_host: house.hubris.network
|
||||
notes:
|
||||
- Docker host for Yuvomi (family planner). Created 2026-06-26.
|
||||
- Migrated from hubris to strong 2026-07-05 (Phase 1 of strong migration plan).
|
||||
- Runs Yuvomi container + WebDAV doc bridge to paperless
|
||||
- 192.168.8.212 was the hubris IP before migration (briefly picked up by teddycloud via DHCP; teddycloud
|
||||
has since been given a static IP, see hosts.teddycloud)
|
||||
age_pubkey: age1s07zs83ehtlg8jtwvr75ltc3c4cdlemfwjuxrwjtwkqxkl9tpggsyrzn2h
|
||||
see_also:
|
||||
- containers/129-house.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -1,14 +1,11 @@
|
||||
# `hubris` — Proxmox host
|
||||
|
||||
Proxmox VE host running 1 VM and 13 LXC containers — the whole homelab's
|
||||
workloads still live here. As of 2026-07-01, hubris is node 1 of the 2-node
|
||||
`Homelab` cluster (see [Cluster](#cluster)); the second node is
|
||||
[strong](strong.md), which hosts nothing yet.
|
||||
Single-node Proxmox VE running 1 VM and 13 LXC containers. The whole homelab.
|
||||
|
||||
## At a glance
|
||||
- **Role:** Proxmox VE 9.1.2 hypervisor (kernel `6.14.11-4-pve`)
|
||||
- **Hardware:** GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS (Phoenix APU), 12 vCPU / ~28 GiB RAM, 2× Samsung 990 EVO Plus NVMe (one SSD primary, one for `library` LVM). 2× Realtek RTL8125 NICs (`r8169`).
|
||||
- **BIOS:** 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||||
- **BIOS:** 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../investigations/2026-04-21-hubris-crash-loop.md).
|
||||
- **Uplink:** `vmbr1` (slave: `eno1`) → SODOLA switch → Fritz!Box 7590. DHCP-reserved `192.168.178.10/24`, gateway `192.168.178.1`.
|
||||
- **Homelab bridge:** `vmbr0` — portless internal bridge, `192.168.8.77/24` + `192.168.8.1/24` alias (LXC default gateway). All 16 LXCs and the HAOS VM are on `vmbr0`. Proxmox routes between `vmbr0` and `vmbr1`; Fritz!Box has a static route `192.168.8.0/24 → 192.168.178.10`.
|
||||
- **WiFi:** disabled 2026-06-02 — `wlp3s0` removed from `/etc/network/interfaces`, wpa config deleted. Was used as a failover to the now-retired Slate AX AP.
|
||||
@@ -25,37 +22,13 @@ workloads still live here. As of 2026-07-01, hubris is node 1 of the 2-node
|
||||
|
||||
`/mnt/library` holds the shared media + data pool: `anime`, `audiobooks`, `books`, `comics`, `documents`, `downloads`, `heaper`, `homecloud`, `images`, `marimo`, `movies`, `music`, `notes`, `podcasts`, `repos`, `roms`, `sophia`. Bind-mounted into every container that needs it. Permissions standard: [media GID 10000](../infrastructure/media-permissions.md).
|
||||
|
||||
## Cluster
|
||||
|
||||
Member of `Homelab`, a 2-node Proxmox cluster with [strong](strong.md)
|
||||
(cluster/OS hostname `strong`), formed 2026-07-01.
|
||||
|
||||
- **Corosync ring0:** hubris's internal `192.168.8.77` (the `vmbr0` address).
|
||||
strong reaches it via the existing Fritz!Box static route
|
||||
(`192.168.8.0/24 → 192.168.178.10`) — no dedicated corosync link, just the
|
||||
household LAN. Fine for a home cluster; not latency-isolated.
|
||||
- **Quorum:** 2 nodes, 1 vote each, no QDevice tiebreaker. Quorum needs both
|
||||
votes — if either node is down (reboot, maintenance, network hiccup), the
|
||||
survivor's running guests keep working but `/etc/pve` goes read-only:
|
||||
no start/stop/create/edit until quorum returns. Decided to skip a QDevice
|
||||
for now; revisit if hubris's periodic reboots (BIOS/thermal work, see
|
||||
Quirks below) make this painful in practice.
|
||||
- **Storage:** `local` / `local-lvm` are the standard per-node default IDs
|
||||
(every node has its own, not actually shared). The `library` lvmthin pool
|
||||
is explicitly restricted to `nodes hubris` in `/etc/pve/storage.cfg` since
|
||||
it's a physical thinpool that only exists on this host's hardware.
|
||||
- strong currently hosts no LXCs/VMs — it exists solely as a cluster
|
||||
member so far. See [strong.md](strong.md) and the [library-SSD
|
||||
migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)
|
||||
for what comes next (physical drive move, service migration — not started).
|
||||
|
||||
## Tenants
|
||||
|
||||
### VMs
|
||||
- [108 — `haos-16.3`](../vms/108-haos.md) — Home Assistant OS, 4 GiB / 32 GiB
|
||||
|
||||
### LXC containers
|
||||
See [containers/index](../containers/index.md). 10 active on hubris (101, 118, 122, 129, 130 migrated to [strong](strong.md) 2026-07-05).
|
||||
See [containers/index](../containers/index.md). 13 active (109 syncthing destroyed 2026-05-14).
|
||||
|
||||
## Boot-time tuning (load-bearing)
|
||||
|
||||
@@ -89,7 +62,7 @@ See [monitoring](../infrastructure/monitoring.md), [backups](../infrastructure/b
|
||||
|
||||
## Quirks
|
||||
|
||||
- `/etc/pve` is fuse — the Proxmox cluster filesystem, now genuinely cluster-synced (2-node) rather than the single-node-but-still-fuse case this note used to describe.
|
||||
- `/etc/pve` is fuse — normal for the Proxmox cluster filesystem, even on a single-node install.
|
||||
- ZFS is **not** in use; storage is LVM-thin + ext4.
|
||||
- Two Realtek 8125 NICs use the in-tree `r8169` driver, not the OOT `r8125`.
|
||||
- Hardware is thermally marginal. NVMe sensors live near warn temp under load. Thermal pads installed on the SSDs 2026-04-23; host relocated to a better-ventilated spot 2026-04-29.
|
||||
@@ -100,8 +73,6 @@ See [monitoring](../infrastructure/monitoring.md), [backups](../infrastructure/b
|
||||
|
||||
- `root@hubris` (self, RSA) — local
|
||||
- `d.toro.v@pm.me` (ed25519) — user's iMac, added 2026-04-22
|
||||
- `root@strong` (RSA) — strong's cluster-join key, added 2026-07-01 so
|
||||
`pvecm add` could authenticate without a password prompt
|
||||
|
||||
OpenSSH on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:22022` and bypasses `authorized_keys` (OIDC/browser). See [SSH access](../infrastructure/ssh-access.md) for the dual-server gotcha.
|
||||
|
||||
@@ -113,20 +84,16 @@ OpenSSH on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:220
|
||||
- [Media permissions](../infrastructure/media-permissions.md)
|
||||
- [Monitoring](../infrastructure/monitoring.md)
|
||||
- [Backups (disabled)](../infrastructure/backups.md)
|
||||
- [Operations cheatsheet](../../../.agents/operations/commands.md)
|
||||
- [Investigation: 2026-04-21 crash loop](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md)
|
||||
- [strong — Proxmox host](strong.md)
|
||||
- [Operations cheatsheet](../operations/commands.md)
|
||||
- [Investigation: 2026-04-21 crash loop](../investigations/2026-04-21-hubris-crash-loop.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-07-01 — strong joined as a 2nd cluster node ("Homelab")
|
||||
User reformatted `strong` (formerly a Linux dev workstation, `192.168.178.181`) to Proxmox VE 9.2.3. Cluster/OS hostname on that box is `strong` (left as-is from install). Bootstrapped root SSH on strong from a one-time console password (installed hubris's existing trusted key set: `root@hubris`, `d.toro.v@pm.me`), then generated a keypair on strong and pre-authorized it here (`root@strong`) so `pvecm add 192.168.8.77 --use_ssh 1` (run from strong) could join without an interactive password prompt. No cabling/routing changes needed — strong reaches hubris's corosync address (`192.168.8.77`) via the existing Fritz!Box static route. Cluster now 2 nodes, quorate, **no QDevice** (explicit choice — see [Cluster](#cluster) above for the quorum tradeoff this implies). strong hosts no guests yet; this is Phase 1 of the [library-SSD migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md), nothing further from that plan has been executed.
|
||||
|
||||
### 2026-06-02 — Slate AX retired; SODOLA switch added; network restructured
|
||||
Replaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed. Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md) and [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).
|
||||
Replaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed. Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md) and [migration plan](../plans/2026-06-01-slate-ax-to-sodola-migration.md).
|
||||
|
||||
### 2026-05-14 — LXC 109 (syncthing) decommissioned
|
||||
User destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy 109 --purge` cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing` was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up. Entry moved to the "recently destroyed" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology); references stripped from [README](../../../README.md), [media-permissions](../infrastructure/media-permissions.md), [vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md).
|
||||
User destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy 109 --purge` cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing` was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up. Entry moved to the "recently destroyed" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology); references stripped from [README](../README.md), [media-permissions](../infrastructure/media-permissions.md), [vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md).
|
||||
|
||||
### 2026-05-14 — network performance baseline captured
|
||||
First explicit speed snapshot: WAN ↓113.5 / ↑19.9 Mbit (24.6 ms), `eno1` 1 Gb full-duplex negotiated, intra-host `vmbr0` ~34.7 Gbit/s host↔LXC and ~34.8 Gbit/s LXC↔LXC (single TCP stream, zero retransmits). `iperf3` + `speedtest-cli` installed on host. Noted `eno1` `rx_errors` at 1.62 M (~1.7 % of 96 M RX packets in 14 d uptime) plus 10.9 k `align_errors` — flagged for follow-up; expect to recheck the trend in ~1 week, suspect patch cable / switch port first if still climbing. See new "Network performance baseline" section above.
|
||||
@@ -138,7 +105,7 @@ User destroyed the heaper LXC. No `116.conf.bak` left behind in `/etc/pve/lxc/`.
|
||||
`/etc/sysctl.d/99-bbr.conf` switches `net.ipv4.tcp_congestion_control` from `cubic` to `bbr` and `net.core.default_qdisc` from `fq_codel` to `fq`. Also bumps `rmem_max`/`wmem_max` to 64 MiB and widens `tcp_rmem`/`tcp_wmem`. `tcp_bbr` module pinned at boot via `/etc/modules-load.d/bbr.conf`. Triggered by Nextcloud client downloads from a WiFi laptop pulling ~2 MB/s despite a 152 Mbps link — server-side baseline through Caddy with BBR is ~400 MB/s single-stream loopback, so any client-perceived single-stream improvement is pure congestion-control win. Touches every LXC's outbound TCP since they all share this kernel.
|
||||
|
||||
### 2026-04-29 — relocated to better-ventilated spot
|
||||
User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C, ACPI zone 47–49 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 60–61 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation).
|
||||
User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C, ACPI zone 47–49 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 60–61 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation).
|
||||
|
||||
### 2026-04-28 — Phase 1 WiFi failover
|
||||
Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failover, metric 200) on the GL-AXT1800-714-5G AP. Installed `wpasupplicant`+`iw`; added `wlp3s0` stanza to `/etc/network/interfaces` with `wpa-conf`; ARP isolation sysctls in `post-up`. Built `wan-failover.service` to remove the vmbr0 default route on `eno1` carrier loss, since the bridge's carrier doesn't follow `eno1` (the LXC veths keep it `1`). LXC/VM guests are still LAN-only — Phase 2 will migrate them.
|
||||
@@ -147,10 +114,10 @@ Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failov
|
||||
This wiki created. Live state at this date: 14 LXCs running (109 syncthing stopped), 1 VM, kernel `6.14.11-4-pve`, uptime 3 d 0 h post drive-removal A/B test. Compared to memory snapshot from a week ago, **destroyed**: LXC 100 (yunohost arr), 106 (flaresolverr), 107 (marimo), 110 (photoprism), 111 (karakeep), 112 (immich), 115 (reticulum). 100 + 106 destroyed per the planned 2026-04-21 \*arr migration retention; the others removed since.
|
||||
|
||||
### 2026-04-23 — SSD cooling + thermal pads installed
|
||||
Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 60–61 °C, nvme1 38–40 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict).
|
||||
Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 60–61 °C, nvme1 38–40 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict).
|
||||
|
||||
### 2026-04-22 — drive removal A/B test
|
||||
Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||||
Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md).
|
||||
|
||||
### 2026-04-22 — `cpu-epp.service` ordering bug fixed
|
||||
Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `pve-guests.service`, so the hottest boot window (20+ guests starting on `performance`) preceded EPP application. Now `After=sysinit.target` + `Before=pve-guests.service`.
|
||||
@@ -159,4 +126,4 @@ Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `
|
||||
`60-crash-capture.conf`, softdog `soft_panic=1`, RuntimeWatchdog 15 s. `rasdaemon` installed and enabled. Pure silicon hangs still leave no trace; this catches everything else.
|
||||
|
||||
### 2026-04-21 — `cpu-epp.service` deployed
|
||||
Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the [crash-loop incident](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||||
Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the [crash-loop incident](../investigations/2026-04-21-hubris-crash-loop.md).
|
||||
@@ -5,7 +5,6 @@ name: hubris
|
||||
kind: proxmox-host
|
||||
os: linux
|
||||
role: hypervisor
|
||||
state: active
|
||||
lan_ip: 192.168.8.77
|
||||
mesh:
|
||||
netbird:
|
||||
@@ -29,8 +28,8 @@ services_hosted:
|
||||
url: https://proxmox.hubris.network
|
||||
backend: hubris
|
||||
port: 8006
|
||||
doc_page: knowledge/wiki/hosts/hubris.md
|
||||
risk_notes: "hypervisor UI \u2014 changes here affect every guest on the node"
|
||||
age_pubkey: age1xkklkvnk5z0fsnh6cfgv70hy9ksfy8rdprwerzw4yk3p4p7cxcqs2yvpz6
|
||||
see_also:
|
||||
- hosts/hubris.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,10 +5,9 @@ name: jellyfin
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: media-server
|
||||
state: active
|
||||
host: strong
|
||||
host: hubris
|
||||
pve_id: 101
|
||||
lan_ip: 192.168.8.246
|
||||
lan_ip: 192.168.8.206
|
||||
mesh:
|
||||
tailscale:
|
||||
fqdn: jellyfin
|
||||
@@ -18,7 +17,7 @@ mesh_globals:
|
||||
- netbird
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/media_local
|
||||
- /mnt/library
|
||||
public_host: media.hubris.network
|
||||
runs:
|
||||
- jellyfin
|
||||
@@ -26,14 +25,7 @@ services_hosted:
|
||||
- name: jellyfin
|
||||
url: https://media.hubris.network
|
||||
backend: jellyfin
|
||||
doc_page: knowledge/wiki/containers/101-jellyfin.md
|
||||
risk_notes: native Authentik OIDC via SSO-Auth plugin, no Caddy forward-auth gate; VAAPI transcode depends
|
||||
on GPU passthrough on strong
|
||||
notes:
|
||||
- Jellyfin 10.11.11 with VAAPI hardware acceleration (Radeon 680M iGPU on strong)
|
||||
- 4 cores / 8 GiB RAM / 1 GiB swap
|
||||
- SSO-Auth plugin v4.0.0.4 with Authentik OIDC (no Caddy forward-auth gate)
|
||||
- GPU passed via dev0+dev1: /dev/dri/renderD128 + card0
|
||||
- Migrated from hubris to strong 2026-07-05 (Phase 2). Library on ludo-lvm.
|
||||
see_also:
|
||||
- containers/101-jellyfin.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -1,19 +1,20 @@
|
||||
# Generated by mcp/build_host_files.py from inventory.yaml.
|
||||
# Do NOT edit by hand — your changes will be overwritten.
|
||||
# Source of truth: ../inventory.yaml
|
||||
name: rclone
|
||||
kind: lxc
|
||||
name: ludo-mini
|
||||
kind: workstation
|
||||
os: linux
|
||||
role: backup
|
||||
state: active
|
||||
role: dev
|
||||
lan_ip: 192.168.178.181
|
||||
mesh:
|
||||
netbird:
|
||||
fqdn: rclone.netbird.selfhosted
|
||||
fqdn: ludo-mini.netbird.selfhosted
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
age_pubkey: age1pwtdws2thdh7vzp2dzttl3zxgcs2tgpcsjsqgw3q04nyml4kvuqq467u4x
|
||||
ssh:
|
||||
user: dtoro
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
@@ -5,7 +5,6 @@ name: mac-mini
|
||||
kind: workstation
|
||||
os: macos
|
||||
role: dev
|
||||
state: active
|
||||
lan_ip: 192.168.178.182
|
||||
mesh:
|
||||
netbird:
|
||||
|
||||
@@ -5,7 +5,6 @@ name: mule-images
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: photo-management
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 120
|
||||
lan_ip: 192.168.8.136
|
||||
@@ -26,7 +25,7 @@ services_hosted:
|
||||
- name: photos
|
||||
url: https://photos.hubris.network
|
||||
backend: mule-images
|
||||
doc_page: knowledge/wiki/containers/120-mule-images.md
|
||||
config_repo: dtoro/mule-image
|
||||
see_also:
|
||||
- containers/120-mule-images.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: netbird-vps
|
||||
kind: external
|
||||
os: linux
|
||||
role: netbird-mgmt
|
||||
state: active
|
||||
mesh:
|
||||
netbird:
|
||||
ip: 100.122.165.149
|
||||
@@ -17,16 +16,6 @@ mesh_globals:
|
||||
- tailscale
|
||||
ssh:
|
||||
user: root
|
||||
runs:
|
||||
- authentik
|
||||
services_hosted:
|
||||
- name: authentik
|
||||
url: https://auth.hubris.network
|
||||
backend: netbird-vps
|
||||
doc_page: knowledge/wiki/containers/106-auth-outpost.md
|
||||
note: core runs on the VPS since 2026-05-31; LAN forward-auth outpost is auth-outpost (LXC 106) at 192.168.8.6:9000.
|
||||
Previous backend value "authentik" referenced the retired embedded-outpost host (LXC 124).
|
||||
risk_notes: "SSO provider \u2014 outage locks login to OIDC/forward-auth services"
|
||||
notes:
|
||||
- "Public IONOS VPS \u2014 hosts the vanilla netbird mgmt+signal+relay+dashboard stack + host coturn (see\
|
||||
\ infrastructure/vps-hardening.md + infrastructure/mesh.md changelog 2026-05-21)."
|
||||
|
||||
@@ -5,7 +5,6 @@ name: nextcloud
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: file-sync
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 114
|
||||
lan_ip: 192.168.8.224
|
||||
@@ -26,6 +25,7 @@ services_hosted:
|
||||
- name: nextcloud
|
||||
url: https://cloud.hubris.network
|
||||
backend: nextcloud
|
||||
doc_page: knowledge/wiki/containers/114-nextcloud.md
|
||||
see_also:
|
||||
- containers/114-nextcloud.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: nfs-export
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: storage-export
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 102
|
||||
lan_ip: 192.168.8.200
|
||||
@@ -14,5 +13,7 @@ mesh_globals:
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
see_also:
|
||||
- containers/102-nfs-export.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: paperless
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: document-archive
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 103
|
||||
lan_ip: 192.168.8.130
|
||||
@@ -26,7 +25,7 @@ services_hosted:
|
||||
- name: paperless
|
||||
url: https://paperless.hubris.network
|
||||
backend: paperless
|
||||
doc_page: knowledge/wiki/containers/103-paperless.md
|
||||
risk_notes: "document archive \u2014 treat data as irreplaceable; DB operations are destructive-class"
|
||||
see_also:
|
||||
- containers/103-paperless.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: republic-laptop
|
||||
kind: workstation
|
||||
os: linux
|
||||
role: primary-dev
|
||||
state: active
|
||||
mesh:
|
||||
netbird:
|
||||
fqdn: republic-laptop.netbird.selfhosted
|
||||
|
||||
@@ -1,26 +0,0 @@
|
||||
# Generated by mcp/build_host_files.py from inventory.yaml.
|
||||
# Do NOT edit by hand — your changes will be overwritten.
|
||||
# Source of truth: ../inventory.yaml
|
||||
name: romm
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: rom-manager
|
||||
state: active
|
||||
host: strong
|
||||
pve_id: 134
|
||||
lan_ip: 192.168.8.249
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/media_local
|
||||
public_host: roms.hubris.network
|
||||
notes:
|
||||
- Docker host for RomM (romm.app) self-hosted ROM manager. Created 2026-07-05.
|
||||
- MariaDB sidecar at /opt/romm/docker-compose.yml.
|
||||
- ROMs on ludo-lvm media volume at /mnt/media_local/roms.
|
||||
- 1 core / 2 GiB RAM / 16 GiB rootfs (ludo-lvm).
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
@@ -1,29 +0,0 @@
|
||||
# Generated by mcp/build_host_files.py from inventory.yaml.
|
||||
# Do NOT edit by hand — your changes will be overwritten.
|
||||
# Source of truth: ../inventory.yaml
|
||||
name: seanime
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: anime-media-server
|
||||
state: active
|
||||
host: strong
|
||||
pve_id: 133
|
||||
lan_ip: 192.168.8.248
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/media_local/anime
|
||||
public_host: seanime.hubris.network
|
||||
notes:
|
||||
- Seanime anime media server for online streaming + local library scanning
|
||||
- Created 2026-07-05. Binary at /opt/seanime/bin/seanime, systemd service.
|
||||
- Connected to qBittorrent on arriman (192.168.8.245:8080)
|
||||
- 8 online streaming extensions installed (HiAnime, AniWatch, KickAssAnime, etc.)
|
||||
- /anime mounted from strong ludo-lvm (/mnt/media_local/anime)
|
||||
- Caddy: "https://seanime.hubris.network \u2192 192.168.8.248:43211"
|
||||
- qBittorrent auth subnet whitelist expanded to 192.168.8.0/24 for seanime access
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
@@ -5,7 +5,6 @@ name: sophia
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: workshop
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 119
|
||||
lan_ip: 192.168.8.109
|
||||
@@ -19,5 +18,7 @@ mesh_globals:
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/library
|
||||
see_also:
|
||||
- containers/119-sophia.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -1,32 +0,0 @@
|
||||
# Generated by mcp/build_host_files.py from inventory.yaml.
|
||||
# Do NOT edit by hand — your changes will be overwritten.
|
||||
# Source of truth: ../inventory.yaml
|
||||
name: strong
|
||||
kind: proxmox-host
|
||||
os: linux
|
||||
role: hypervisor
|
||||
state: active
|
||||
lan_ip: 192.168.178.181
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
ssh:
|
||||
user: root
|
||||
notes:
|
||||
- Reformatted from Linux workstation ("ludo-mini" in this wiki, still the machine's nickname) to Proxmox
|
||||
VE 9.2.3 on 2026-07-01. Renamed the inventory/wiki identity from ludo-mini to strong on the same day
|
||||
so it matches the OS/cluster hostname everywhere (bootstrap looks up hosts/$(hostname).yaml, so a mismatch
|
||||
would break enrollment).
|
||||
- "Joined hubris's \"Homelab\" cluster same day. 2-node, no QDevice tiebreaker yet \u2014 see hosts/hubris.md\
|
||||
\ quorum note."
|
||||
- Netbird not yet installed (fresh OS wiped prior enrollment); reachable today only via the household
|
||||
LAN / existing Fritz static route to 192.168.8.0/24. Re-enroll in mesh as a follow-up if off-LAN access
|
||||
to this host itself (not just its future guests) is needed.
|
||||
- "First step of the planned library-SSD migration \u2014 see .hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md\
|
||||
\ (filename kept as-is, it's a historical planning doc). Only Phase 1 (Proxmox install + cluster join)\
|
||||
\ is done; no physical drive move, service migration, or GPU passthrough has happened yet."
|
||||
age_pubkey: age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
@@ -1,42 +0,0 @@
|
||||
# Generated by mcp/build_host_files.py from inventory.yaml.
|
||||
# Do NOT edit by hand — your changes will be overwritten.
|
||||
# Source of truth: ../inventory.yaml
|
||||
name: teddycloud
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: teddycloud
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 131
|
||||
lan_ip: 192.168.8.150
|
||||
mesh_globals:
|
||||
primary: netbird
|
||||
accepted:
|
||||
- netbird
|
||||
- tailscale
|
||||
mounts:
|
||||
- /mnt/library
|
||||
public_host: teddy.hubris.network
|
||||
runs:
|
||||
- teddycloud
|
||||
services_hosted:
|
||||
- name: teddycloud
|
||||
url: https://teddy.hubris.network
|
||||
backend: teddycloud
|
||||
doc_page: knowledge/wiki/containers/131-teddycloud.md
|
||||
note: self-hosted TeddyCloud (Toniebox cloud reimplementation), docker compose
|
||||
risk_notes: "no Caddy forward-auth gate (unlike sab.hubris.network on the same Caddyfile) \u2014 reachable\
|
||||
\ to anyone on the LAN/mesh who can resolve teddy.hubris.network; undocumented in inventory.yaml until\
|
||||
\ 2026-07-06 (drift-caught)"
|
||||
notes:
|
||||
- Docker host for TeddyCloud (ghcr.io/toniebox-reverse-engineering/teddycloud), a self-hosted reimplementation
|
||||
of the Toniebox cloud backend. Debian 12 (bookworm).
|
||||
- 1 core / 1 GiB RAM / 512 MiB swap / 16 GiB rootfs (local-lvm).
|
||||
- "Predates the client-enrollment convention \u2014 undocumented in inventory.yaml until 2026-07-06, when\
|
||||
\ Oikos's drift detector (oikos/drift.py) caught pve_id 131 live on hubris (`pct list`) with no inventory\
|
||||
\ entry. Static IP assigned 2026-07-05 during the strong migration (was picking up 192.168.8.243 via\
|
||||
\ DHCP before that \u2014 see hosts/strong.md's 2026-07-05 changelog)."
|
||||
- "No age_pubkey / homelab-context enrollment \u2014 not a homelab CLI client, just a docker-compose app\
|
||||
\ container. Not a required follow-up unless it needs secrets."
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
@@ -5,7 +5,6 @@ name: trmnl
|
||||
kind: lxc
|
||||
os: linux
|
||||
role: trmnl-middleware
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 128
|
||||
lan_ip: 192.168.8.211
|
||||
@@ -22,7 +21,7 @@ services_hosted:
|
||||
backend: trmnl
|
||||
url: https://trmnl.hubris.network
|
||||
note: self-hosted middleware for TRMNL e-ink plugins (polled by TRMNL cloud)
|
||||
doc_page: knowledge/wiki/containers/128-trmnl.md
|
||||
config_repo: dtoro/terminalito
|
||||
see_also:
|
||||
- containers/128-trmnl.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -5,7 +5,6 @@ name: zimaos
|
||||
kind: vm
|
||||
os: linux
|
||||
role: nas-frontend-eval
|
||||
state: active
|
||||
host: hubris
|
||||
pve_id: 100
|
||||
lan_ip: 192.168.8.195
|
||||
@@ -21,6 +20,7 @@ services_hosted:
|
||||
- name: zimaos
|
||||
url: https://zimaos.hubris.network
|
||||
backend: zimaos
|
||||
doc_page: knowledge/wiki/vms/100-zimaos.md
|
||||
see_also:
|
||||
- vms/100-zimaos.md
|
||||
mcp_endpoint: https://mcp.hubris.network/mcp
|
||||
secrets_issuance_endpoint: https://secrets.hubris.network/issue
|
||||
|
||||
@@ -23,7 +23,7 @@ The app repo at `/opt/<thing>` is the working tree, but the deploy tooling (`web
|
||||
- ~~`192.168.8.230` (claudio-bot — destroyed 2026-06-04)~~
|
||||
- `192.168.8.136` ([mule-images (120)](../containers/120-mule-images.md))
|
||||
- `192.168.8.77` ([hubris host](../hosts/hubris.md) — backup-library)
|
||||
- ~~`192.168.8.190` ([plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology))~~ (destroyed 2026-06-28)
|
||||
- ~~`192.168.8.190` ([plato (126)](../containers/126-plato.md))~~ (destroyed 2026-06-28)
|
||||
- `192.168.8.211` ([trmnl (128)](../containers/128-trmnl.md) — terminalito)
|
||||
|
||||
**Don't strip these when editing app.ini.**
|
||||
@@ -38,19 +38,18 @@ The app repo at `/opt/<thing>` is the working tree, but the deploy tooling (`web
|
||||
| `dtoro/gitea-customizations` | [gitea (104)](../containers/104-gitea.md) `/var/lib/gitea/custom/` | A | `http://127.0.0.1:9797/deploy` (loopback) | (orig) | `systemctl restart gitea` if templates changed |
|
||||
| `dtoro/mule-image` | [mule-images (120)](../containers/120-mule-images.md) `/opt/mule-image/` | B | `http://192.168.8.136:9797/deploy` | 6 | `docker compose up -d --build` |
|
||||
| `dtoro/Artifacto` | [apps (105)](../containers/105-apps.md) `/opt/artifacto/` | B | `http://192.168.8.205:9798/deploy` | 7 | `docker compose up -d --build` |
|
||||
| ~~`dtoro/Plato`~~ | ~~[plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology) `/opt/plato/app/`~~ (destroyed 2026-06-28) | ⊘ | `http://192.168.8.190:9799/deploy` (dead) | 8 (removed) | Repo archived — LXC destroyed |
|
||||
| `dtoro/claudio-bot` | ~~[claudio-bot (123)](../containers/archive/123-claudio-bot.md)~~ (destroyed 2026-06-04) | ⊘ | `http://192.168.8.230:9797/deploy` (dead) | (archived) | Repo archived — LXC destroyed |
|
||||
| ~~`dtoro/Plato`~~ | ~~[plato (126)](../containers/126-plato.md) `/opt/plato/app/`~~ (destroyed 2026-06-28) | ⊘ | `http://192.168.8.190:9799/deploy` (dead) | 8 (removed) | Repo archived — LXC destroyed |
|
||||
| `dtoro/claudio-bot` | ~~[claudio-bot (123)](../containers/123-claudio-bot.md)~~ (destroyed 2026-06-04) | ⊘ | `http://192.168.8.230:9797/deploy` (dead) | (archived) | Repo archived — LXC destroyed |
|
||||
| `dtoro/backup-library` | [hubris host](../hosts/hubris.md) `/opt/backup-library/` | A | `http://192.168.8.77:9798/deploy` | (orig) | runs `deploy.sh` (preserves admin-edited `/etc/restic/include-*.list`) |
|
||||
| `dtoro/Homelab-Docs` → homelab-mcp | [apps (105)](../containers/105-apps.md) `/opt/homelab-mcp/` | B | `http://192.168.8.205:9811/deploy` | 10 | reinstalls `homelab-mcp.service` + restart |
|
||||
| `dtoro/Homelab-Docs` → secrets-issuance | [apps (105)](../containers/105-apps.md) `/opt/secrets-issuance/` | B | `http://192.168.8.205:9821/deploy` | 11 | reinstalls `secrets-issuance.service` + restart |
|
||||
| `dtoro/terminalito` | [trmnl (128)](../containers/128-trmnl.md) `/opt/terminalito/` | B | `http://192.168.8.211:9797/deploy` | 12 | reinstalls units + `systemctl restart trmnl-plugins` |
|
||||
| `dtoro/Homelab-Docs` → oikos-console | [apps (105)](../containers/105-apps.md) `/opt/oikos-console/` | B | `http://192.168.8.205:9831/deploy` | 14 | reinstalls `oikos-console.service` + restart — see [oikos/console/deploy/README.md](../../../oikos/console/deploy/README.md) |
|
||||
|
||||
> Note: `dtoro/Homelab-Docs` has **three webhooks** firing on the same push.
|
||||
> Note: `dtoro/Homelab-Docs` has **two webhooks** firing on the same push.
|
||||
> Each owns its own clone on LXC 105. They don't conflict because each
|
||||
> deploy.sh only touches its own service unit + venv.
|
||||
|
||||
> **Not yet wired:** `dtoro/claudio-monitor` (push, then `/opt/claudio-monitor/scripts/deploy.sh` manually). The former authentik LXC (124) is destroyed — Authentik runs on the [VPS](../../../hosts/netbird-vps.yaml). DNS moved to [Technitium on dns (107)](../containers/107-dns.md).
|
||||
> **Not yet wired:** `dtoro/claudio-monitor` (push, then `/opt/claudio-monitor/scripts/deploy.sh` manually). The former authentik LXC (124) is destroyed — Authentik runs on the [VPS](../hosts/netbird-vps.md). DNS moved to [Technitium on dns (107)](../containers/107-dns.md).
|
||||
|
||||
## When you change a tracked config
|
||||
|
||||
@@ -117,7 +116,7 @@ If you're not sure what's already lurking, run `homelab apt-audit --fleet` and l
|
||||
- [Gitea (104)](../containers/104-gitea.md) — webhook source for all of these
|
||||
- [Caddy (121)](../containers/121-caddy.md), [apps (105)](../containers/105-apps.md), [mule-images (120)](../containers/120-mule-images.md), [hubris host](../hosts/hubris.md) — webhook targets
|
||||
- [Backups (disabled)](backups.md)
|
||||
- [Operations cheatsheet](../../../.agents/operations/commands.md) — `homelab apt-audit` / `homelab apt-upgrade` reference
|
||||
- [Operations cheatsheet](../operations/commands.md) — `homelab apt-audit` / `homelab apt-upgrade` reference
|
||||
|
||||
## Changelog
|
||||
|
||||
@@ -131,7 +130,7 @@ Webhook id 12 on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy` on [
|
||||
Webhook ids 10 + 11 on `dtoro/Homelab-Docs` (ports `9811` + `9821` on [apps (105)](../containers/105-apps.md)). Two webhooks on one repo — each owns its own clone (`/opt/homelab-mcp`, `/opt/secrets-issuance`) and only restarts its own service. See [homelab-context](homelab-context.md) for why both services live in one repo.
|
||||
|
||||
### 2026-05-13 — Plato pipeline added
|
||||
Webhook id 8 on `dtoro/Plato` (port `9799` on [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology)). `app.ini` `ALLOWED_HOST_LIST` extended to include `192.168.8.190`.
|
||||
Webhook id 8 on `dtoro/Plato` (port `9799` on [plato (126)](../containers/126-plato.md)). `app.ini` `ALLOWED_HOST_LIST` extended to include `192.168.8.190`.
|
||||
|
||||
### 2026-04-28 — wiki entry created
|
||||
Initial documentation. Six active pipelines.
|
||||
@@ -1,30 +1,6 @@
|
||||
# Backups — restic on external drive (DEPRECATED — superseded)
|
||||
# Backups — restic on external drive (DISABLED)
|
||||
|
||||
> **DEPRECATED 2026-07-01.** Superseded by the **rclone → Proton Drive** off-host mirror on
|
||||
> [LXC 132 `rclone`](../containers/132-rclone.md). That job finally closes the off-host / 3-2-1 gap
|
||||
> this page flagged for months. The restic-on-USB job below is kept for archaeology; it has been
|
||||
> **DISABLED since 2026-04-22** and is not coming back in its old form.
|
||||
|
||||
## Current backup — rclone → Proton Drive (LXC 132)
|
||||
|
||||
- **Where:** [LXC 132 `rclone`](../containers/132-rclone.md) (`192.168.8.214`), `/mnt/library`
|
||||
mounted **read-only**.
|
||||
- **What:** plain `rclone sync` (Proton mirrors local; browsable files, no versioning) of the
|
||||
folders listed in `/etc/rclone-backup/folders.list`, to `proton:library-backup/…`.
|
||||
- **When:** monthly — `rclone-backup.timer` (`OnCalendar=*-*-01 03:00`).
|
||||
- **UI:** rclone Web GUI on `192.168.8.214:5572` (LAN-only, no auth).
|
||||
- **Encryption:** Proton's built-in E2E (no rclone `crypt` overlay).
|
||||
- **Runs / logs:** `/var/log/rclone-backup/` + `runs.jsonl`.
|
||||
- **Still a single off-host target** (Proton only). Not yet a full 3-2-1 (no second independent
|
||||
copy), but strictly better than the previous "no off-host copy at all."
|
||||
|
||||
See [132-rclone](../containers/132-rclone.md) for the full design.
|
||||
|
||||
---
|
||||
|
||||
## Legacy — restic on external drive (DISABLED 2026-04-22)
|
||||
|
||||
Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabled 2026-04-22** as part of the [hubris crash-loop A/B test](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||||
Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabled 2026-04-22** as part of the [hubris crash-loop A/B test](../investigations/2026-04-21-hubris-crash-loop.md).
|
||||
|
||||
## Status
|
||||
|
||||
@@ -36,7 +12,7 @@ Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabl
|
||||
|
||||
Fstab entry commented out. USB drive de-authorized and physically removed. `backup-library-deploy.service` left enabled (harmless webhook receiver).
|
||||
|
||||
**Reason:** the host hang recurred 2026-04-22 18:42 after 30h despite the `cpu-epp` fix, the UAS blacklist, and mount-on-demand. User wants to confirm host stability without the drive at all (was stable 33 days before the drive arrived). See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||||
**Reason:** the host hang recurred 2026-04-22 18:42 after 30h despite the `cpu-epp` fix, the UAS blacklist, and mount-on-demand. User wants to confirm host stability without the drive at all (was stable 33 days before the drive arrived). See [investigation](../investigations/2026-04-21-hubris-crash-loop.md).
|
||||
|
||||
**To re-enable:** uncomment fstab line, `systemctl enable --now` the four timers, re-attach drive.
|
||||
|
||||
@@ -105,13 +81,13 @@ Runbook at `/usr/share/doc/backup-library/RECOVERY.md` (or in the repo at `doc/R
|
||||
|
||||
## Known SPOF
|
||||
|
||||
Single drive. RECOVERY.md flags the 3-2-1 gap. Mitigations (second drive, cloud repo via `restic copy`) were not implemented before this job was retired — the **off-host copy is now provided by [rclone → Proton Drive (LXC 132)](../containers/132-rclone.md)** instead. A second independent copy is still outstanding.
|
||||
Single drive. RECOVERY.md flags the 3-2-1 gap. Mitigations (second drive, cloud repo via `restic copy`) are not yet implemented.
|
||||
|
||||
## Drive history
|
||||
|
||||
The `Silicon Motion Portable SSD` (vid:pid `090c:2320`) drops under sustained heavy writes through a hub chain. Bypass all hubs / use a rear motherboard USB 3 port if attaching it again.
|
||||
|
||||
After it was first attached on 2026-04-19, hubris crashed twice in 2.5 days (46h then 12h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic, OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1. The UAS blacklist + mount-on-demand mitigations didn't fully eliminate it (recurrence 2026-04-22), prompting drive removal as the cleaner test. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||||
After it was first attached on 2026-04-19, hubris crashed twice in 2.5 days (46h then 12h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic, OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1. The UAS blacklist + mount-on-demand mitigations didn't fully eliminate it (recurrence 2026-04-22), prompting drive removal as the cleaner test. See [investigation](../investigations/2026-04-21-hubris-crash-loop.md).
|
||||
|
||||
## Thermal monitoring
|
||||
|
||||
@@ -119,21 +95,18 @@ Moved out of this repo to `dtoro/claudio-monitor` on 2026-04-21 (commit `50dc213
|
||||
|
||||
## Related
|
||||
- [Hubris host](../hosts/hubris.md)
|
||||
- ~~[claudio-bot (123)](../containers/archive/123-claudio-bot.md)~~ (destroyed 2026-06-04)
|
||||
- ~~[claudio-bot (123)](../containers/123-claudio-bot.md)~~ (destroyed 2026-06-04)
|
||||
- [Monitoring](monitoring.md)
|
||||
- [Auto-deploy](auto-deploy.md)
|
||||
- [Investigation: 2026-04-21 crash loop](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md)
|
||||
- [Investigation: 2026-04-21 crash loop](../investigations/2026-04-21-hubris-crash-loop.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-07-01 — DEPRECATED; superseded by rclone → Proton Drive (LXC 132)
|
||||
Off-host backup moved to a plain `rclone sync` mirror on the new [LXC 132 `rclone`](../containers/132-rclone.md) (`/mnt/library` → Proton Drive, monthly, LAN Web GUI). This finally provides the off-host copy the "Known SPOF" note wanted. The restic-on-USB units on hubris remain `disabled` (drive already removed 2026-04-22); page restructured to lead with the current job and demote restic to "Legacy".
|
||||
|
||||
### 2026-04-28 — wiki entry created
|
||||
Initial documentation. Status remains DISABLED.
|
||||
|
||||
### 2026-04-22 — DISABLED
|
||||
Drive removed as the A/B test in the [crash investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md). Timers disabled, fstab commented, drive de-authorized.
|
||||
Drive removed as the A/B test in the [crash investigation](../investigations/2026-04-21-hubris-crash-loop.md). Timers disabled, fstab commented, drive de-authorized.
|
||||
|
||||
### 2026-04-21 — UAS blacklist + mount-on-demand shipped; root-caused host hangs to drive
|
||||
Drive identified as the source of the hangs after hubris crashed twice in 2.5 days. UAS blacklist forces BOT; helper script toggles `/sys/bus/usb/.../authorized` so the drive is de-authorized when not backing up. Recovery drill (restore 188KB PDF + hash compare) had passed earlier. Bug fixed in `backup-library.sh`: `python3 -c '…' KEY=VAL` does NOT pass env vars — env-var prefix must precede the command. Caused false-failure even after successful backups.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user