diff --git a/archive/hermes-plans/STALE.md b/archive/hermes-plans/STALE.md new file mode 100644 index 00000000..15343fe4 --- /dev/null +++ b/archive/hermes-plans/STALE.md @@ -0,0 +1,5 @@ +# Stale — historical reference only + +These are Hermes agent planning documents from June–July 2026. All plans were +executed or superseded. Retained for narrative context on why decisions were +made. Not actively maintained. \ No newline at end of file diff --git a/archive/knowledge/index.md b/archive/knowledge/index.md index 8252ad32..5c1e98a3 100644 --- a/archive/knowledge/index.md +++ b/archive/knowledge/index.md @@ -1,14 +1,19 @@ -# Knowledge +# Knowledge (archived) -The durable, authoritative current-state documentation of the homelab: one page per node and per -cross-cutting system, synthesized from live state and evidence. Structure and rules are in -[the knowledge schema](../.agents/domains/knowledge/schema.md). +> **Status: Stale — historical reference only, last updated 2026-07-06.** +> The DB is now the single source of truth for all structured data and +> knowledge (bootstrapped from `seeds/knowledge.yaml`). Use MCP +> `search_knowledge` / `get_entity_knowledge` for live queries. +> +> Infrastructure architecture docs have been moved to +> [`docs/infrastructure/`](../../docs/infrastructure/). See `MOVED.md` there +> for what moved where. | Section | What it covers | |---------|----------------| -| [wiki/hosts/](wiki/hosts/index.md) | Proxmox host narratives — `hubris`, `strong`. | -| [wiki/containers/](wiki/containers/index.md) | LXC fleet — one page per container, plus the master table and archaeology. | -| [wiki/vms/](wiki/vms/index.md) | Virtual machines — ZimaOS, Home Assistant OS. | -| [wiki/infrastructure/](wiki/infrastructure/index.md) | Cross-cutting systems — DNS, ingress, mesh, storage, auth, monitoring, generated topology. | +| [hosts/](hosts/index.md) | Proxmox host narratives — `hubris`, `strong`. | +| [containers/](containers/index.md) | LXC fleet — one page per container, plus the master table and archaeology. | +| [vms/](vms/index.md) | Virtual machines — ZimaOS, Home Assistant OS. | +| [infrastructure/](infrastructure/index.md) | Cross-cutting systems — **docs moved to `docs/infrastructure/`**. | +| [GLOSSARY.md](GLOSSARY.md) | **Moved to `docs/GLOSSARY.md`**. | | [sources/](sources/index.md) | External reference docs and the pointer to incident evidence. | -| [GLOSSARY.md](GLOSSARY.md) | Term definitions. | diff --git a/archive/knowledge/infrastructure/MOVED.md b/archive/knowledge/infrastructure/MOVED.md new file mode 100644 index 00000000..5087c0d7 --- /dev/null +++ b/archive/knowledge/infrastructure/MOVED.md @@ -0,0 +1,20 @@ +# Moved to `docs/infrastructure/` + +The infrastructure architecture docs have been moved to +[`docs/infrastructure/`](../../docs/infrastructure/): + +- [network.md](../../docs/infrastructure/network.md) +- [dns.md](../../docs/infrastructure/dns.md) +- [mesh.md](../../docs/infrastructure/mesh.md) +- [ssh-access.md](../../docs/infrastructure/ssh-access.md) +- [ingress.md](../../docs/infrastructure/ingress.md) +- [media-permissions.md](../../docs/infrastructure/media-permissions.md) +- [backups.md](../../docs/infrastructure/backups.md) +- [homelab-context.md](../../docs/infrastructure/homelab-context.md) +- [auto-deploy.md](../../docs/infrastructure/auto-deploy.md) +- [vps-hardening.md](../../docs/infrastructure/vps-hardening.md) +- [monitoring.md](../../docs/infrastructure/monitoring.md) +- [oikos-check-lifecycle.md](../../docs/infrastructure/oikos-check-lifecycle.md) + +The copies here are kept for archive continuity but are NOT the source of +truth — update the docs/ copies instead. \ No newline at end of file diff --git a/archive/ledger/2026-07.jsonl b/archive/ledger/2026-07.jsonl deleted file mode 100644 index b91313b3..00000000 --- a/archive/ledger/2026-07.jsonl +++ /dev/null @@ -1,5 +0,0 @@ -{"ts": "2026-07-06T11:05:35+00:00", "agent": "mac-mini", "entity": "host:teddycloud", "action": "activate", "risk": "config_mutation", "verification": "homelab node teddycloud relations", "result": "ok"} -{"ts": "2026-07-06T11:15:04+00:00", "agent": "mac-mini", "entity": "repo:Homelab-Docs", "action": "register-webhook", "risk": "config_mutation", "result": "ok", "notes": "webhook id 14 for oikos-console deploy"} -{"ts": "2026-07-06T11:29:56+00:00", "agent": "mac-mini", "entity": "service:caddy", "action": "add-site-block", "risk": "config_mutation", "verification": "curl -s https://git.hubris.network (unrelated route still healthy after reload)", "result": "ok", "notes": "oikos.hubris.network -> 192.168.8.205:8091, Authentik-gated, in dtoro/caddy-conf@c195142"} -{"ts": "2026-07-06T11:40:28+00:00", "agent": "mac-mini", "entity": "host:dns", "action": "add-record", "risk": "config_mutation", "verification": "dig @192.168.8.2 +short oikos.hubris.network", "result": "ok", "notes": "oikos.hubris.network A -> 192.168.8.175 (Caddy LAN IP), via Technitium API, no token persisted"} -{"ts": "2026-07-06T11:57:21+00:00", "agent": "mac-mini", "entity": "host:apps", "action": "deploy-oikos-console", "risk": "config_mutation", "verification": "curl http://127.0.0.1:8091/ on apps -> 200; https://oikos.hubris.network/ -> 302 (Authentik gate)", "result": "ok"} diff --git a/archive/oikos-cards/cards/host-apps.md b/archive/oikos-cards/cards/host-apps.md deleted file mode 100644 index 4afe8ac6..00000000 --- a/archive/oikos-cards/cards/host-apps.md +++ /dev/null @@ -1,21 +0,0 @@ -# apps (host:apps) - -- kind: lxc (LXC 105) -- state: active -- runs-on: host:hubris -- role: docker-apps -- address: 192.168.8.205 (mesh: tailscale:apps) -- mounts: /mnt/library -- doc: knowledge/wiki/containers/105-apps.md -- secrets: enrolled (age key present) - -## Blast radius -- impacts: service:artifacto, service:homelab_mcp, service:secrets_issuance -- affected by: host:hubris, mount:/mnt/library, repo:dtoro/Artifacto, repo:dtoro/Homelab-Docs -- full blast radius: service:artifacto, service:homelab_mcp, service:secrets_issuance - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- 2026-07-06T11:57:21+00:00 deploy-oikos-console (config_mutation) — ok diff --git a/archive/oikos-cards/cards/host-arriman.md b/archive/oikos-cards/cards/host-arriman.md deleted file mode 100644 index 5db5481a..00000000 --- a/archive/oikos-cards/cards/host-arriman.md +++ /dev/null @@ -1,20 +0,0 @@ -# arriman (host:arriman) - -- kind: lxc (LXC 122) -- state: active -- runs-on: host:strong -- role: arr-stack -- address: 192.168.8.245 (mesh: tailscale:arr) -- mounts: /mnt/media_local -- doc: knowledge/wiki/containers/122-arriman.md - -## Blast radius -- impacts: service:arr_stack -- affected by: host:strong, mount:/mnt/media_local -- full blast radius: service:arr_stack - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-auth-outpost.md b/archive/oikos-cards/cards/host-auth-outpost.md deleted file mode 100644 index 525c654e..00000000 --- a/archive/oikos-cards/cards/host-auth-outpost.md +++ /dev/null @@ -1,18 +0,0 @@ -# auth-outpost (host:auth-outpost) - -- kind: lxc (LXC 106) -- state: active -- runs-on: host:hubris -- role: authentik-gateway -- address: 192.168.8.6 -- doc: knowledge/wiki/containers/106-auth-outpost.md - -## Blast radius -- impacts: (none) -- affected by: host:hubris - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-caddy.md b/archive/oikos-cards/cards/host-caddy.md deleted file mode 100644 index ee80bcad..00000000 --- a/archive/oikos-cards/cards/host-caddy.md +++ /dev/null @@ -1,19 +0,0 @@ -# caddy (host:caddy) - -- kind: lxc (LXC 121) -- state: active -- runs-on: host:hubris -- role: reverse-proxy -- address: 192.168.8.175 -- doc: knowledge/wiki/containers/121-caddy.md - -## Blast radius -- impacts: service:caddy -- affected by: host:hubris, repo:dtoro/caddy-conf -- full blast radius: service:caddy - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-dns.md b/archive/oikos-cards/cards/host-dns.md deleted file mode 100644 index 0b76c3f2..00000000 --- a/archive/oikos-cards/cards/host-dns.md +++ /dev/null @@ -1,19 +0,0 @@ -# dns (host:dns) - -- kind: lxc (LXC 107) -- state: active -- runs-on: host:hubris -- role: dns-server -- address: 192.168.8.2 -- doc: knowledge/wiki/containers/107-dns.md - -## Blast radius -- impacts: service:dns -- affected by: host:hubris -- full blast radius: service:dns - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- 2026-07-06T11:40:28+00:00 add-record (config_mutation) — ok diff --git a/archive/oikos-cards/cards/host-elementsynapse.md b/archive/oikos-cards/cards/host-elementsynapse.md deleted file mode 100644 index 6509c967..00000000 --- a/archive/oikos-cards/cards/host-elementsynapse.md +++ /dev/null @@ -1,19 +0,0 @@ -# elementsynapse (host:elementsynapse) - -- kind: lxc (LXC 118) -- state: active -- runs-on: host:strong -- role: matrix-server -- address: 192.168.8.242 -- doc: knowledge/wiki/containers/118-elementsynapse.md - -## Blast radius -- impacts: service:matrix -- affected by: host:strong -- full blast radius: service:matrix - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-gitea.md b/archive/oikos-cards/cards/host-gitea.md deleted file mode 100644 index a77eb7f4..00000000 --- a/archive/oikos-cards/cards/host-gitea.md +++ /dev/null @@ -1,20 +0,0 @@ -# gitea (host:gitea) - -- kind: lxc (LXC 104) -- state: active -- runs-on: host:hubris -- role: git-server -- address: 192.168.8.121 (mesh: tailscale:gitea) -- mounts: /mnt/library -- doc: knowledge/wiki/containers/104-gitea.md - -## Blast radius -- impacts: service:gitea -- affected by: host:hubris, mount:/mnt/library, repo:dtoro/gitea-customizations -- full blast radius: service:gitea - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-grimmory.md b/archive/oikos-cards/cards/host-grimmory.md deleted file mode 100644 index 706377c6..00000000 --- a/archive/oikos-cards/cards/host-grimmory.md +++ /dev/null @@ -1,20 +0,0 @@ -# grimmory (host:grimmory) - -- kind: lxc (LXC 130) -- state: active -- runs-on: host:strong -- role: book-library -- address: 192.168.8.247 -- mounts: /mnt/media_local -- doc: knowledge/wiki/containers/130-grimmory.md -- secrets: enrolled (age key present) - -## Blast radius -- impacts: (none) -- affected by: host:strong, mount:/mnt/media_local - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-haos.md b/archive/oikos-cards/cards/host-haos.md deleted file mode 100644 index 45f2abaf..00000000 --- a/archive/oikos-cards/cards/host-haos.md +++ /dev/null @@ -1,19 +0,0 @@ -# haos (host:haos) - -- kind: vm (VM 108) -- state: active -- runs-on: host:hubris -- role: home-automation -- address: 192.168.8.101 (mesh: tailscale:homeassistant) -- doc: knowledge/wiki/vms/108-haos.md - -## Blast radius -- impacts: service:haos -- affected by: host:hubris -- full blast radius: service:haos - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-house.md b/archive/oikos-cards/cards/host-house.md deleted file mode 100644 index 5f3808c6..00000000 --- a/archive/oikos-cards/cards/host-house.md +++ /dev/null @@ -1,19 +0,0 @@ -# house (host:house) - -- kind: lxc (LXC 129) -- state: active -- runs-on: host:strong -- role: family-planner -- address: 192.168.8.244 -- doc: knowledge/wiki/containers/129-house.md -- secrets: enrolled (age key present) - -## Blast radius -- impacts: (none) -- affected by: host:strong - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-hubris.md b/archive/oikos-cards/cards/host-hubris.md deleted file mode 100644 index 57236b61..00000000 --- a/archive/oikos-cards/cards/host-hubris.md +++ /dev/null @@ -1,20 +0,0 @@ -# hubris (host:hubris) - -- kind: proxmox-host -- state: active -- role: hypervisor -- address: 192.168.8.77 (mesh: netbird:proxmox-server.netbird.selfhosted) -- mounts: /mnt/library -- doc: knowledge/wiki/hosts/hubris.md -- secrets: enrolled (age key present) - -## Blast radius -- impacts: host:apps, host:auth-outpost, host:caddy, host:dns, host:gitea, host:haos, host:mule-images, host:nextcloud, host:nfs-export, host:paperless, host:sophia, host:teddycloud, host:trmnl, host:zimaos, service:proxmox_ui -- affected by: mount:/mnt/library -- full blast radius: host:apps, host:auth-outpost, host:caddy, host:dns, host:gitea, host:haos, host:mule-images, host:nextcloud, host:nfs-export, host:paperless, host:sophia, host:teddycloud, host:trmnl, host:zimaos, service:artifacto, service:caddy, service:dns, service:gitea, service:haos, service:homelab_mcp, service:nextcloud, service:paperless, service:photos, service:proxmox_ui, service:secrets_issuance, service:teddycloud, service:trmnl, service:zimaos - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-jellyfin.md b/archive/oikos-cards/cards/host-jellyfin.md deleted file mode 100644 index 178c58d5..00000000 --- a/archive/oikos-cards/cards/host-jellyfin.md +++ /dev/null @@ -1,20 +0,0 @@ -# jellyfin (host:jellyfin) - -- kind: lxc (LXC 101) -- state: active -- runs-on: host:strong -- role: media-server -- address: 192.168.8.246 (mesh: tailscale:jellyfin) -- mounts: /mnt/media_local -- doc: knowledge/wiki/containers/101-jellyfin.md - -## Blast radius -- impacts: service:jellyfin -- affected by: host:strong, mount:/mnt/media_local -- full blast radius: service:jellyfin - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-mac-mini.md b/archive/oikos-cards/cards/host-mac-mini.md deleted file mode 100644 index 7f683d32..00000000 --- a/archive/oikos-cards/cards/host-mac-mini.md +++ /dev/null @@ -1,17 +0,0 @@ -# mac-mini (host:mac-mini) - -- kind: workstation -- state: active -- role: dev -- address: 192.168.178.182 (mesh: netbird:mac-mini-234-17.netbird.selfhosted) -- secrets: enrolled (age key present) - -## Blast radius -- impacts: (none) -- affected by: (none) - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-mule-images.md b/archive/oikos-cards/cards/host-mule-images.md deleted file mode 100644 index c1c98cac..00000000 --- a/archive/oikos-cards/cards/host-mule-images.md +++ /dev/null @@ -1,20 +0,0 @@ -# mule-images (host:mule-images) - -- kind: lxc (LXC 120) -- state: active -- runs-on: host:hubris -- role: photo-management -- address: 192.168.8.136 (mesh: tailscale:muleimage) -- mounts: /mnt/library -- doc: knowledge/wiki/containers/120-mule-images.md - -## Blast radius -- impacts: service:photos -- affected by: host:hubris, mount:/mnt/library, repo:dtoro/mule-image -- full blast radius: service:photos - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-netbird-vps.md b/archive/oikos-cards/cards/host-netbird-vps.md deleted file mode 100644 index 511beccc..00000000 --- a/archive/oikos-cards/cards/host-netbird-vps.md +++ /dev/null @@ -1,17 +0,0 @@ -# netbird-vps (host:netbird-vps) - -- kind: external -- state: active -- role: netbird-mgmt -- address: (mesh: netbird:netbird-ionos.netbird.selfhosted) - -## Blast radius -- impacts: service:authentik -- affected by: (none) -- full blast radius: service:authentik - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-nextcloud.md b/archive/oikos-cards/cards/host-nextcloud.md deleted file mode 100644 index 2cd8e3be..00000000 --- a/archive/oikos-cards/cards/host-nextcloud.md +++ /dev/null @@ -1,20 +0,0 @@ -# nextcloud (host:nextcloud) - -- kind: lxc (LXC 114) -- state: active -- runs-on: host:hubris -- role: file-sync -- address: 192.168.8.224 (mesh: tailscale:nextcloud) -- mounts: /mnt/library -- doc: knowledge/wiki/containers/114-nextcloud.md - -## Blast radius -- impacts: service:nextcloud -- affected by: host:hubris, mount:/mnt/library -- full blast radius: service:nextcloud - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-nfs-export.md b/archive/oikos-cards/cards/host-nfs-export.md deleted file mode 100644 index 3731f135..00000000 --- a/archive/oikos-cards/cards/host-nfs-export.md +++ /dev/null @@ -1,18 +0,0 @@ -# nfs-export (host:nfs-export) - -- kind: lxc (LXC 102) -- state: active -- runs-on: host:hubris -- role: storage-export -- address: 192.168.8.200 -- doc: knowledge/wiki/containers/102-nfs-export.md - -## Blast radius -- impacts: (none) -- affected by: host:hubris - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-paperless.md b/archive/oikos-cards/cards/host-paperless.md deleted file mode 100644 index f9935515..00000000 --- a/archive/oikos-cards/cards/host-paperless.md +++ /dev/null @@ -1,20 +0,0 @@ -# paperless (host:paperless) - -- kind: lxc (LXC 103) -- state: active -- runs-on: host:hubris -- role: document-archive -- address: 192.168.8.130 (mesh: tailscale:paperless) -- mounts: /mnt/library -- doc: knowledge/wiki/containers/103-paperless.md - -## Blast radius -- impacts: service:paperless -- affected by: host:hubris, mount:/mnt/library -- full blast radius: service:paperless - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-rclone.md b/archive/oikos-cards/cards/host-rclone.md deleted file mode 100644 index 60b05e68..00000000 --- a/archive/oikos-cards/cards/host-rclone.md +++ /dev/null @@ -1,17 +0,0 @@ -# rclone (host:rclone) - -- kind: lxc -- state: active -- role: backup -- address: (mesh: netbird:rclone.netbird.selfhosted) -- secrets: enrolled (age key present) - -## Blast radius -- impacts: (none) -- affected by: (none) - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-republic-laptop.md b/archive/oikos-cards/cards/host-republic-laptop.md deleted file mode 100644 index f70c178d..00000000 --- a/archive/oikos-cards/cards/host-republic-laptop.md +++ /dev/null @@ -1,16 +0,0 @@ -# republic-laptop (host:republic-laptop) - -- kind: workstation -- state: active -- role: primary-dev -- address: (mesh: netbird:republic-laptop.netbird.selfhosted) - -## Blast radius -- impacts: (none) -- affected by: (none) - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-romm.md b/archive/oikos-cards/cards/host-romm.md deleted file mode 100644 index 31fd618d..00000000 --- a/archive/oikos-cards/cards/host-romm.md +++ /dev/null @@ -1,19 +0,0 @@ -# romm (host:romm) - -- kind: lxc (LXC 134) -- state: active -- runs-on: host:strong -- role: rom-manager -- address: 192.168.8.249 -- mounts: /mnt/media_local -- doc: knowledge/wiki/containers/134-romm.md - -## Blast radius -- impacts: (none) -- affected by: host:strong, mount:/mnt/media_local - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-seanime.md b/archive/oikos-cards/cards/host-seanime.md deleted file mode 100644 index 155d9917..00000000 --- a/archive/oikos-cards/cards/host-seanime.md +++ /dev/null @@ -1,19 +0,0 @@ -# seanime (host:seanime) - -- kind: lxc (LXC 133) -- state: active -- runs-on: host:strong -- role: anime-media-server -- address: 192.168.8.248 -- mounts: /mnt/media_local/anime -- doc: knowledge/wiki/containers/133-seanime.md - -## Blast radius -- impacts: (none) -- affected by: host:strong, mount:/mnt/media_local/anime - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-sophia.md b/archive/oikos-cards/cards/host-sophia.md deleted file mode 100644 index 196efeff..00000000 --- a/archive/oikos-cards/cards/host-sophia.md +++ /dev/null @@ -1,19 +0,0 @@ -# sophia (host:sophia) - -- kind: lxc (LXC 119) -- state: active -- runs-on: host:hubris -- role: workshop -- address: 192.168.8.109 (mesh: tailscale:sophia) -- mounts: /mnt/library -- doc: knowledge/wiki/containers/119-sophia.md - -## Blast radius -- impacts: (none) -- affected by: host:hubris, mount:/mnt/library - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-strong.md b/archive/oikos-cards/cards/host-strong.md deleted file mode 100644 index 9b921c13..00000000 --- a/archive/oikos-cards/cards/host-strong.md +++ /dev/null @@ -1,19 +0,0 @@ -# strong (host:strong) - -- kind: proxmox-host -- state: active -- role: hypervisor -- address: 192.168.178.181 -- doc: knowledge/wiki/hosts/strong.md -- secrets: enrolled (age key present) - -## Blast radius -- impacts: host:arriman, host:elementsynapse, host:grimmory, host:house, host:jellyfin, host:romm, host:seanime -- affected by: (none) -- full blast radius: host:arriman, host:elementsynapse, host:grimmory, host:house, host:jellyfin, host:romm, host:seanime, service:arr_stack, service:jellyfin, service:matrix - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-teddycloud.md b/archive/oikos-cards/cards/host-teddycloud.md deleted file mode 100644 index 57bc9928..00000000 --- a/archive/oikos-cards/cards/host-teddycloud.md +++ /dev/null @@ -1,20 +0,0 @@ -# teddycloud (host:teddycloud) - -- kind: lxc (LXC 131) -- state: active -- runs-on: host:hubris -- role: teddycloud -- address: 192.168.8.150 -- mounts: /mnt/library -- doc: knowledge/wiki/containers/131-teddycloud.md - -## Blast radius -- impacts: service:teddycloud -- affected by: host:hubris, mount:/mnt/library -- full blast radius: service:teddycloud - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- 2026-07-06T11:05:35+00:00 activate (config_mutation) — ok diff --git a/archive/oikos-cards/cards/host-trmnl.md b/archive/oikos-cards/cards/host-trmnl.md deleted file mode 100644 index 741b975c..00000000 --- a/archive/oikos-cards/cards/host-trmnl.md +++ /dev/null @@ -1,19 +0,0 @@ -# trmnl (host:trmnl) - -- kind: lxc (LXC 128) -- state: active -- runs-on: host:hubris -- role: trmnl-middleware -- address: 192.168.8.211 -- doc: knowledge/wiki/containers/128-trmnl.md - -## Blast radius -- impacts: service:trmnl -- affected by: host:hubris, repo:dtoro/terminalito -- full blast radius: service:trmnl - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/host-zimaos.md b/archive/oikos-cards/cards/host-zimaos.md deleted file mode 100644 index b249fa48..00000000 --- a/archive/oikos-cards/cards/host-zimaos.md +++ /dev/null @@ -1,19 +0,0 @@ -# zimaos (host:zimaos) - -- kind: vm (VM 100) -- state: active -- runs-on: host:hubris -- role: nas-frontend-eval -- address: 192.168.8.195 -- doc: knowledge/wiki/vms/100-zimaos.md - -## Blast radius -- impacts: service:zimaos -- affected by: host:hubris -- full blast radius: service:zimaos - -## Safe actions -- see the services this host runs for action-level risk classes - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-arr_stack.md b/archive/oikos-cards/cards/service-arr_stack.md deleted file mode 100644 index 2c09e3b0..00000000 --- a/archive/oikos-cards/cards/service-arr_stack.md +++ /dev/null @@ -1,17 +0,0 @@ -# arr_stack (service:arr_stack) - -- backend: host:arriman -- doc: knowledge/wiki/containers/122-arriman.md - -## Blast radius -- impacts: (none) -- affected by: host:arriman - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-artifacto.md b/archive/oikos-cards/cards/service-artifacto.md deleted file mode 100644 index fc857a31..00000000 --- a/archive/oikos-cards/cards/service-artifacto.md +++ /dev/null @@ -1,20 +0,0 @@ -# artifacto (service:artifacto) - -- backend: host:apps -- url: https://artifacto.hubris.network -- doc: knowledge/wiki/containers/105-apps.md -- config repo: dtoro/Artifacto - -## Blast radius -- impacts: (none) -- affected by: host:apps - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-authentik.md b/archive/oikos-cards/cards/service-authentik.md deleted file mode 100644 index 93ec04b9..00000000 --- a/archive/oikos-cards/cards/service-authentik.md +++ /dev/null @@ -1,19 +0,0 @@ -# authentik (service:authentik) - -- backend: host:netbird-vps -- url: https://auth.hubris.network -- doc: knowledge/wiki/containers/106-auth-outpost.md -- risk notes: SSO provider — outage locks login to OIDC/forward-auth services - -## Blast radius -- impacts: (none) -- affected by: host:netbird-vps - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-caddy.md b/archive/oikos-cards/cards/service-caddy.md deleted file mode 100644 index a26093da..00000000 --- a/archive/oikos-cards/cards/service-caddy.md +++ /dev/null @@ -1,20 +0,0 @@ -# caddy (service:caddy) - -- backend: host:caddy -- doc: knowledge/wiki/containers/121-caddy.md -- config repo: dtoro/caddy-conf -- risk notes: wide blast radius — every *.hubris.network route rides on it (see oikos/policy.yaml service_overrides) - -## Blast radius -- impacts: (none) -- affected by: host:caddy - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — config_mutation (approval: operator) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- 2026-07-06T11:29:56+00:00 add-site-block (config_mutation) — ok diff --git a/archive/oikos-cards/cards/service-dns.md b/archive/oikos-cards/cards/service-dns.md deleted file mode 100644 index a7a3d9e8..00000000 --- a/archive/oikos-cards/cards/service-dns.md +++ /dev/null @@ -1,18 +0,0 @@ -# dns (service:dns) - -- backend: host:dns -- doc: knowledge/wiki/containers/107-dns.md -- risk notes: LAN-wide resolver — misconfig breaks name resolution for every client - -## Blast radius -- impacts: (none) -- affected by: host:dns - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-gitea.md b/archive/oikos-cards/cards/service-gitea.md deleted file mode 100644 index 89dca938..00000000 --- a/archive/oikos-cards/cards/service-gitea.md +++ /dev/null @@ -1,21 +0,0 @@ -# gitea (service:gitea) - -- backend: host:gitea -- url: https://git.hubris.network -- doc: knowledge/wiki/containers/104-gitea.md -- config repo: dtoro/gitea-customizations -- risk notes: hosts all config repos + deploy webhooks; outage blocks auto-deploy and sync - -## Blast radius -- impacts: (none) -- affected by: host:gitea - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-haos.md b/archive/oikos-cards/cards/service-haos.md deleted file mode 100644 index 4c3401bb..00000000 --- a/archive/oikos-cards/cards/service-haos.md +++ /dev/null @@ -1,17 +0,0 @@ -# haos (service:haos) - -- backend: host:haos -- doc: knowledge/wiki/vms/108-haos.md - -## Blast radius -- impacts: (none) -- affected by: host:haos - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-homelab_mcp.md b/archive/oikos-cards/cards/service-homelab_mcp.md deleted file mode 100644 index fb5a1569..00000000 --- a/archive/oikos-cards/cards/service-homelab_mcp.md +++ /dev/null @@ -1,21 +0,0 @@ -# homelab_mcp (service:homelab_mcp) - -- backend: host:apps -- url: https://mcp.hubris.network/mcp -- doc: knowledge/wiki/infrastructure/homelab-context.md -- config repo: dtoro/Homelab-Docs -- risk notes: agents' primary read surface — outage degrades every agent to grepping the clone - -## Blast radius -- impacts: (none) -- affected by: host:apps - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-jellyfin.md b/archive/oikos-cards/cards/service-jellyfin.md deleted file mode 100644 index e12e8507..00000000 --- a/archive/oikos-cards/cards/service-jellyfin.md +++ /dev/null @@ -1,19 +0,0 @@ -# jellyfin (service:jellyfin) - -- backend: host:jellyfin -- url: https://media.hubris.network -- doc: knowledge/wiki/containers/101-jellyfin.md -- risk notes: native Authentik OIDC via SSO-Auth plugin, no Caddy forward-auth gate; VAAPI transcode depends on GPU passthrough on strong - -## Blast radius -- impacts: (none) -- affected by: host:jellyfin - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-matrix.md b/archive/oikos-cards/cards/service-matrix.md deleted file mode 100644 index d321ee2a..00000000 --- a/archive/oikos-cards/cards/service-matrix.md +++ /dev/null @@ -1,19 +0,0 @@ -# matrix (service:matrix) - -- backend: host:elementsynapse -- url: https://matrix.hubris.network -- doc: knowledge/wiki/containers/118-elementsynapse.md -- risk notes: alert/approval channel for Oikos — outage silences agent escalation - -## Blast radius -- impacts: (none) -- affected by: host:elementsynapse - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-nextcloud.md b/archive/oikos-cards/cards/service-nextcloud.md deleted file mode 100644 index 995644b3..00000000 --- a/archive/oikos-cards/cards/service-nextcloud.md +++ /dev/null @@ -1,18 +0,0 @@ -# nextcloud (service:nextcloud) - -- backend: host:nextcloud -- url: https://cloud.hubris.network -- doc: knowledge/wiki/containers/114-nextcloud.md - -## Blast radius -- impacts: (none) -- affected by: host:nextcloud - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-paperless.md b/archive/oikos-cards/cards/service-paperless.md deleted file mode 100644 index cc02065b..00000000 --- a/archive/oikos-cards/cards/service-paperless.md +++ /dev/null @@ -1,19 +0,0 @@ -# paperless (service:paperless) - -- backend: host:paperless -- url: https://paperless.hubris.network -- doc: knowledge/wiki/containers/103-paperless.md -- risk notes: document archive — treat data as irreplaceable; DB operations are destructive-class - -## Blast radius -- impacts: (none) -- affected by: host:paperless - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-photos.md b/archive/oikos-cards/cards/service-photos.md deleted file mode 100644 index a8f94425..00000000 --- a/archive/oikos-cards/cards/service-photos.md +++ /dev/null @@ -1,20 +0,0 @@ -# photos (service:photos) - -- backend: host:mule-images -- url: https://photos.hubris.network -- doc: knowledge/wiki/containers/120-mule-images.md -- config repo: dtoro/mule-image - -## Blast radius -- impacts: (none) -- affected by: host:mule-images - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-proxmox_ui.md b/archive/oikos-cards/cards/service-proxmox_ui.md deleted file mode 100644 index a83de8c4..00000000 --- a/archive/oikos-cards/cards/service-proxmox_ui.md +++ /dev/null @@ -1,19 +0,0 @@ -# proxmox_ui (service:proxmox_ui) - -- backend: host:hubris -- url: https://proxmox.hubris.network -- doc: knowledge/wiki/hosts/hubris.md -- risk notes: hypervisor UI — changes here affect every guest on the node - -## Blast radius -- impacts: (none) -- affected by: host:hubris - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-secrets_issuance.md b/archive/oikos-cards/cards/service-secrets_issuance.md deleted file mode 100644 index 1178cbe3..00000000 --- a/archive/oikos-cards/cards/service-secrets_issuance.md +++ /dev/null @@ -1,21 +0,0 @@ -# secrets_issuance (service:secrets_issuance) - -- backend: host:apps -- url: https://secrets.hubris.network/issue -- doc: .agents/operations/agent-enrollment.md -- config repo: dtoro/Homelab-Docs -- risk notes: identity issuance — any change is security-sensitive; key operations are destructive-class - -## Blast radius -- impacts: (none) -- affected by: host:apps - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-teddycloud.md b/archive/oikos-cards/cards/service-teddycloud.md deleted file mode 100644 index 1533be8c..00000000 --- a/archive/oikos-cards/cards/service-teddycloud.md +++ /dev/null @@ -1,19 +0,0 @@ -# teddycloud (service:teddycloud) - -- backend: host:teddycloud -- url: https://teddy.hubris.network -- doc: knowledge/wiki/containers/131-teddycloud.md -- risk notes: no Caddy forward-auth gate (unlike sab.hubris.network on the same Caddyfile) — reachable to anyone on the LAN/mesh who can resolve teddy.hubris.network; undocumented in inventory.yaml until 2026-07-06 (drift-caught) - -## Blast radius -- impacts: (none) -- affected by: host:teddycloud - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-trmnl.md b/archive/oikos-cards/cards/service-trmnl.md deleted file mode 100644 index 29e9540a..00000000 --- a/archive/oikos-cards/cards/service-trmnl.md +++ /dev/null @@ -1,20 +0,0 @@ -# trmnl (service:trmnl) - -- backend: host:trmnl -- url: https://trmnl.hubris.network -- doc: knowledge/wiki/containers/128-trmnl.md -- config repo: dtoro/terminalito - -## Blast radius -- impacts: (none) -- affected by: host:trmnl - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) -- edit-config-and-deploy — config_mutation (approval: operator) - -## Recent changes -- (none yet) diff --git a/archive/oikos-cards/cards/service-zimaos.md b/archive/oikos-cards/cards/service-zimaos.md deleted file mode 100644 index 6c8f048a..00000000 --- a/archive/oikos-cards/cards/service-zimaos.md +++ /dev/null @@ -1,18 +0,0 @@ -# zimaos (service:zimaos) - -- backend: host:zimaos -- url: https://zimaos.hubris.network -- doc: knowledge/wiki/vms/100-zimaos.md - -## Blast radius -- impacts: (none) -- affected by: host:zimaos - -## Safe actions -- health-check — read_only (approval: none) -- view-logs — read_only (approval: none) -- view-docs — read_only (approval: none) -- restart — reversible_low (approval: none) - -## Recent changes -- (none yet) diff --git a/archive/secrets-issuance/STALE.md b/archive/secrets-issuance/STALE.md new file mode 100644 index 00000000..976b42ef --- /dev/null +++ b/archive/secrets-issuance/STALE.md @@ -0,0 +1,5 @@ +# Stale — historical reference only + +Python-era secrets-issuance HTTP service. Not ported to Go; functionality +superseded by the Oikos enrollment flow and Infisical. Retained for protocol +design reference. \ No newline at end of file diff --git a/archive/secrets-sops-backup/MOVED.md b/archive/secrets-sops-backup/MOVED.md new file mode 100644 index 00000000..3af51f99 --- /dev/null +++ b/archive/secrets-sops-backup/MOVED.md @@ -0,0 +1,10 @@ +# Docs moved to `docs/secrets/` + +The secret management procedures and rotation runbook have been moved to +[`docs/secrets/`](../../docs/secrets/): + +- [README.md](../../docs/secrets/README.md) — SOPS/Infisical conventions +- [rotation.md](../../docs/secrets/rotation.md) — rotation runbook + +The encrypted `.yaml` files here are kept as DR fallback only. Do not update +them — use Infisical for live secret management. \ No newline at end of file diff --git a/docs/GLOSSARY.md b/docs/GLOSSARY.md new file mode 100644 index 00000000..9b68dc0c --- /dev/null +++ b/docs/GLOSSARY.md @@ -0,0 +1,33 @@ +# Glossary + +Terms and abbreviations used throughout the homelab wiki. + +| Term | Meaning | +|------|---------| +| **Authentik** | SSO/identity provider. Core runs on the VPS; forward-auth outpost at LXC 106 on hubris | +| **Caddy** | Reverse proxy (LXC 121). Terminates TLS for every `*.hubris.network` hostname | +| **Caveman** | Terse communication standard for agent responses — no filler, keep substance | +| **Forward-auth** | Caddy snippet that delegates authentication to an Authentik outpost. Protects web UIs like qBit, SABnzbd | +| **Gitea** | Git server at `git.hubris.network`. Hosts all tracked config repos | +| **Gluetun** | WireGuard VPN sidecar on arriman. All \*arr traffic routes through it | +| **HAOS** | Home Assistant Operating System. VM 108 on hubris | +| **Hubris** | Primary Proxmox VE node (GMKtec NucBox M6 Ultra). PVE hostname, cluster member 1 | +| **LXC** | Linux Container (Proxmox). VM-like isolation without a full OS kernel | +| **LVM-thin** | Thin-provisioned logical volume manager. Used for all container/VM storage | +| **MCP** | Model Context Protocol (MCP server at `mcp.hubris.network`). Structured tools for agents to query homelab state | +| **Mesh** | Overlay VPN for off-LAN connectivity. Netbird is current; Tailscale is legacy | +| **Netbird** | Preferred mesh VPN. VPS hosts the management plane; all homelab nodes are members | +| **OIDC** | OpenID Connect. Protocol used by Authentik for SSO login flows | +| **Oikos** | Agent operating model ([.agents/OIKOS.md](../.agents/OIKOS.md)). OODA loop, risk classes, policy, ontology | +| **PVE** | Proxmox Virtual Environment — the hypervisor on both hubris and strong | +| **SOPS** | `sops` — Mozilla SOPS. Encrypts secrets with age keys so they live in the git repo | +| **Strong** | Secondary Proxmox VE node. Cluster member 2 (hostname `strong`, nickname ludo/ludo-mini) | +| **Traefik** | Reverse proxy on IONOS VPS. Serves `*.hubris.network` to the public internet | +| **VAAPI** | Video Acceleration API. Intel/AMD GPU-based hardware transcode for Jellyfin | +| **VPS** | Virtual Private Server at IONOS (`82.165.190.79`). Runs Authentik core + Netbird management | +| **\\*arr** | Media automation suite: Sonarr (TV), Radarr (movies), Lidarr (music), Prowlarr (indexer), Bazarr (subtitles), Readarr (books — not in use) | + +## See also + +- [Infrastructure index](wiki/infrastructure/index.md) — cross-cutting systems each with their own doc page +- [OIKOS operating model](../.agents/OIKOS.md) — agent policy, risk classes, lifecycle \ No newline at end of file diff --git a/docs/index.md b/docs/index.md index 0eb82d5a..58c94422 100644 --- a/docs/index.md +++ b/docs/index.md @@ -1,15 +1,18 @@ # Docs -Long-form reference material for the Oikos platform. Operational state and -topology live in the DB (seeded from `seeds/`); these docs cover decisions, -procedures, and the system model. +Long-form reference material for the Oikos platform and homelab infrastructure. +Operational state and topology live in the DB (seeded from `seeds/`); these docs +cover architecture, decisions, procedures, and the system model. | Path | Contents | | ---- | -------- | | [adr/](adr/README.md) | Architecture Decision Records (numbered, append-only) | +| [infrastructure/](infrastructure/) | Homelab infrastructure architecture — network, DNS, mesh, ingress, backups, SSH, VPS hardening, auto-deploy, monitoring | | [mbse/](mbse/README.md) | Model-Based Systems Engineering views of the platform | | [mascot/](mascot/README.md) | MBSE subsystem model for the desktop mascot (planned) | | [operations/](operations/README.md) | Operator runbooks (deploy, rollback, recovery) | +| [secrets/](secrets/README.md) | Secret management procedures and rotation runbook | +| [GLOSSARY.md](GLOSSARY.md) | Term definitions across the homelab | For agent orientation see [AGENTS.md](../AGENTS.md); for the operating model see [.agents/OIKOS.md](../.agents/OIKOS.md); for development see diff --git a/docs/infrastructure/auto-deploy.md b/docs/infrastructure/auto-deploy.md new file mode 100644 index 00000000..e1d24c94 --- /dev/null +++ b/docs/infrastructure/auto-deploy.md @@ -0,0 +1,149 @@ +# Auto-deploy — gitea-webhook pipelines + +Several configs and apps in the lab live in `dtoro/*` repos on [gitea (104)](containers/104-gitea.md) and auto-redeploy on push. All pipelines follow one of two shapes. + +## Two shapes + +### Shape A — checkout IS the working tree (config repos) + +`/etc/` or `/var/lib//...` is itself a `git clone`. Push triggers `git pull` + a reload command. Used for pure-config repos where re-cloning is cheap. + +### Shape B — receiver outside the app repo (compose stacks) + +The app repo at `/opt/` is the working tree, but the deploy tooling (`webhook.py`, `deploy.sh`, systemd unit) lives in a sibling `/opt/-deploy/` so the app repo stays portable. Push triggers `git pull` + `docker compose up -d --build`. Returns 202 immediately and runs the build in a daemon thread because docker builds exceed gitea's request timeout. + +## Common + +- All receivers validate `X-Gitea-Signature` HMAC-SHA256 against a per-pipeline secret in `/etc/-deploy/secret`. +- All filter to `refs/heads/main` (or `master` for older repos). Gitea's "test delivery" button sends `ref=main` (without `refs/heads/`) — those will log "ignoring ref main" and 204. Real pushes work. **Don't "fix" the ref filter to accept both** — it'd also accept PR merges from side branches that got fast-forwarded. +- Gitea's `app.ini` `[webhook] ALLOWED_HOST_LIST` must include every receiver IP. Currently: + - `127.0.0.1` (gitea customizations on [LXC 104](containers/104-gitea.md)) + - `192.168.8.175` ([caddy (121)](containers/121-caddy.md)) + - `192.168.8.205` ([apps (105)](containers/105-apps.md) — Artifacto) + - ~~`192.168.8.230` (claudio-bot — destroyed 2026-06-04)~~ + - `192.168.8.136` ([mule-images (120)](containers/120-mule-images.md)) + - `192.168.8.77` ([hubris host](hosts/hubris.md) — backup-library) + - ~~`192.168.8.190` ([plato (126)](containers/index.md#recently-destroyed-kept-for-archaeology))~~ (destroyed 2026-06-28) + - `192.168.8.211` ([trmnl (128)](containers/128-trmnl.md) — terminalito) + + **Don't strip these when editing app.ini.** + +- Git creds for root-run deploy services live in `/etc/-deploy/git-credentials` (mode 600) and are wired via `credential.helper = store --file=/etc/-deploy/git-credentials` in the repo's `.git/config`. Necessary because the unit typically runs with `ProtectHome=true`, which blocks `/root`. + +## Pipelines + +| Repo | Target | Shape | Receiver | Webhook id | Reload action | +| ------------------------------- | -------------------------------------------- | ----- | ------------------------------------- | ---------- | ------------- | +| `dtoro/caddy-conf` | [caddy (121)](containers/121-caddy.md) `/etc/caddy/` | A | `http://192.168.8.175:9797/deploy` | 2 | `caddy validate` + `systemctl reload caddy` | +| `dtoro/gitea-customizations` | [gitea (104)](containers/104-gitea.md) `/var/lib/gitea/custom/` | A | `http://127.0.0.1:9797/deploy` (loopback) | (orig) | `systemctl restart gitea` if templates changed | +| `dtoro/mule-image` | [mule-images (120)](containers/120-mule-images.md) `/opt/mule-image/` | B | `http://192.168.8.136:9797/deploy` | 6 | `docker compose up -d --build` | +| `dtoro/Artifacto` | [apps (105)](containers/105-apps.md) `/opt/artifacto/` | B | `http://192.168.8.205:9798/deploy` | 7 | `docker compose up -d --build` | +| ~~`dtoro/Plato`~~ | ~~[plato (126)](containers/index.md#recently-destroyed-kept-for-archaeology) `/opt/plato/app/`~~ (destroyed 2026-06-28) | ⊘ | `http://192.168.8.190:9799/deploy` (dead) | 8 (removed) | Repo archived — LXC destroyed | +| `dtoro/claudio-bot` | ~~[claudio-bot (123)](containers/archive/123-claudio-bot.md)~~ (destroyed 2026-06-04) | ⊘ | `http://192.168.8.230:9797/deploy` (dead) | (archived) | Repo archived — LXC destroyed | +| `dtoro/backup-library` | [hubris host](hosts/hubris.md) `/opt/backup-library/` | A | `http://192.168.8.77:9798/deploy` | (orig) | runs `deploy.sh` (preserves admin-edited `/etc/restic/include-*.list`) | +| `dtoro/Homelab-Docs` → homelab-mcp | [apps (105)](containers/105-apps.md) `/opt/homelab-mcp/` | B | `http://192.168.8.205:9811/deploy` | 10 (deprecated) | ~~reinstalls `homelab-mcp.service` + restart~~ → replaced by Go Docker stack on mac-mini | +| `dtoro/Homelab-Docs` → secrets-issuance | [apps (105)](containers/105-apps.md) `/opt/secrets-issuance/` | B | `http://192.168.8.205:9821/deploy` | 11 (deprecated) | ~~reinstalls `secrets-issuance.service` + restart~~ → replaced by `internal/secrets/` Go package | +| `dtoro/terminalito` | [trmnl (128)](containers/128-trmnl.md) `/opt/terminalito/` | B | `http://192.168.8.211:9797/deploy` | 12 | reinstalls units + `systemctl restart trmnl-plugins` | +| `dtoro/Homelab-Docs` → oikos-console | [apps (105)](containers/105-apps.md) `/opt/oikos-console/` | B | `http://192.168.8.205:9831/deploy` | 14 | reinstalls `oikos-console.service` + restart — see [oikos/console/deploy/README.md](../../../oikos/console/deploy/README.md) | + +> Note: `dtoro/Homelab-Docs` has **three webhooks** firing on the same push. +> Each owns its own clone on LXC 105. They don't conflict because each +> deploy.sh only touches its own service unit + venv. + +> **Not yet wired:** `dtoro/claudio-monitor` (push, then `/opt/claudio-monitor/scripts/deploy.sh` manually). The former authentik LXC (124) is destroyed — Authentik runs on the [VPS](../../../hosts/netbird-vps.yaml). DNS moved to [Technitium on dns (107)](containers/107-dns.md). + +## When you change a tracked config + +Always commit + push. Local-only edits drift. Common ones: + +- `/etc/caddy/Caddyfile` ↔ `dtoro/caddy-conf` (auto-deploys) +- `/var/lib/gitea/custom/` ↔ `dtoro/gitea-customizations` (auto-deploys) +- `/opt/artifacto/` ↔ `dtoro/Artifacto` (auto-deploys) +- `/opt/mule-image/` ↔ `dtoro/mule-image` (auto-deploys) +- ~~`/opt/plato/app/` ↔ `dtoro/Plato`~~ (destroyed 2026-06-28) +- ~~`/opt/claudio-bot/` ↔ `dtoro/claudio-bot`~~ (destroyed 2026-06-04) +- `/opt/backup-library/` ↔ `dtoro/backup-library` (auto-deploys) +- `/opt/homelab-mcp/` + `/opt/secrets-issuance/` ↔ `dtoro/Homelab-Docs` (auto-deploys both, see [homelab-context](homelab-context.md)) + +## Per-pipeline notes / gotchas + +### caddy-conf +- Repo includes `scripts/webhook/install.sh`. Editing the systemd unit *inside the repo* does **not** auto-reinstall — re-run `install.sh` manually after unit edits. +- The unit has `ReadWritePaths=/etc/caddy` — load-bearing (`ProtectSystem=full` would otherwise block `git pull`). + +### gitea-customizations +- Receiver is on **loopback** (`127.0.0.1:9797`), not the LXC IP. +- Online3DViewer binary assets are NOT tracked; `deploy.sh` fetches them on first run. + +### mule-image / Artifacto +- Async deploy (returns 202) — gitea would otherwise time out the request. Logs: `pct exec -- journalctl -u -deploy-webhook -f`. +- **Cloning from inside the LXC must use the internal gitea IP** (`http://192.168.8.121:3000/...`). `https://git.hubris.network` hits a connection reset from inside [apps (105)](containers/105-apps.md) (Caddy routing / TLS hairpin not configured for this LXC). Configured `origin` on the in-LXC checkout is the internal URL. +- Manual deploy: `pct exec -- /opt/-deploy/deploy.sh`. +- Health: `pct exec -- curl -s http://127.0.0.1:/health` → `ok`. +- Setup tokens used to register the webhook (e.g., `artifacto-deploy-setup`, `artifacto-deploy-setup-2`, `artifacto-cleanup` on user `dtoro`) need manual revocation in the Gitea UI → Settings → Applications → Manage Access Tokens. Gitea's `/users/{u}/tokens` endpoints require basic auth (not bearer), so cleanup couldn't be automated. + +### backup-library +- Currently the only deploy that targets the host directly (`192.168.8.77:9798`). +- `deploy.sh` is careful to preserve admin edits to `/etc/restic/include-*.list` — canonical source is `config/` in the repo, but the install path is treated as authoritative once `deploy.sh` has run. + +### homelab-mcp / secrets-issuance +- Both ride a single push to `dtoro/Homelab-Docs`. Two clones on LXC 105 + (`/opt/homelab-mcp`, `/opt/secrets-issuance`) — each is an independent + Shape-B target with its own webhook receiver. +- The deploy script restarts the service it just updated. Because the + webhook receiver itself is a separate systemd unit (`*-deploy.service`), + it does NOT restart itself — but `deploy.sh` running `systemctl + restart homelab-mcp-deploy.service` (or the secrets-issuance one) + would create a kill-self loop. The current `deploy.sh` is careful + to only restart the main service. +- Both services consume `/opt/homelab-context` for their runtime data + (inventory, secret recipient lookup). That clone is **the same clone + every other client has** — kept fresh by `homelab-context-sync.timer`, + not by these webhooks. + +## Custom-built binaries that overlap apt-managed paths + +If a pipeline (or any out-of-band build) drops a binary into a path that an apt package also owns — most commonly `/usr/bin/` — then the next `apt upgrade` of the corresponding package will silently clobber the custom build. That's exactly how [LXC 121 caddy](containers/121-caddy.md) went down for ~10 min on 2026-05-21: an xcaddy build with `caddy-dns/ionos` lived at `/usr/bin/caddy` and Debian's caddy 2.11.2→2.11.3 apt upgrade replaced it with a vanilla 2.11.3 that couldn't parse the Caddyfile. + +Two patterns are acceptable, pick one when authoring a pipeline that ships a non-apt binary: + +1. **Ship the build as a `.deb` with an epoch-bumped version.** Use `dpkg-deb --build` (or `nfpm`) to package the binary as `Package: `, `Version: 1:-hubris`. The epoch (`1:`) means it beats any non-epoch upstream version regardless of point bumps, so `apt upgrade` is a no-op for that package. Used by caddy: `caddy 1:2.11.3-hubris1` (see commit `2026-05-21` in [121-caddy.md](containers/121-caddy.md)). + +2. **Hold the apt package.** `apt-mark hold ` on the LXC during pipeline install; apt will refuse to upgrade it. Simpler than `.deb` packaging but: (a) the hold flag isn't preserved by `dpkg -i` of a new version, (b) it's invisible unless you check `apt-mark showhold`, (c) you have to remember to `apt-mark unhold` when you intentionally want a new version. The new `homelab apt-audit` subcommand surfaces holds across the fleet so they don't get forgotten. + +If you're not sure what's already lurking, run `homelab apt-audit --fleet` and look at the `NONAPT` column — that's a count of binaries in `/usr/bin/{caddy,docker,jellyfin}` + `/usr/local/bin/*` that no apt package owns. + +## Related +- [Gitea (104)](containers/104-gitea.md) — webhook source for all of these +- [Caddy (121)](containers/121-caddy.md), [apps (105)](containers/105-apps.md), [mule-images (120)](containers/120-mule-images.md), [hubris host](hosts/hubris.md) — webhook targets +- [Backups (disabled)](backups.md) +- [Operations cheatsheet](../../../.agents/operations/commands.md) — `homelab apt-audit` / `homelab apt-upgrade` reference + +## Changelog + +### 2026-06-28 — Plato pipeline decommissioned +LXC 126 destroyed, webhook id 8 on `dtoro/Plato` removed. `192.168.8.190` removed from gitea `app.ini` `ALLOWED_HOST_LIST`. + +### 2026-06-24 — terminalito pipeline added +Webhook id 12 on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy` on [trmnl (128)](containers/128-trmnl.md). Shape B (`server/deploy/webhook.py` receiver, in-repo `server/deploy/deploy.sh`; secret `/etc/terminalito-deploy/secret`). `app.ini` `ALLOWED_HOST_LIST` extended with `192.168.8.211`. Verified end-to-end with a push. Repo-local `credential.helper` in `/opt/terminalito/.git/config` (the unit can't read root's global git config). + +### 2026-05-20 — homelab-mcp + secrets-issuance pipelines added +Webhook ids 10 + 11 on `dtoro/Homelab-Docs` (ports `9811` + `9821` on [apps (105)](containers/105-apps.md)). Two webhooks on one repo — each owns its own clone (`/opt/homelab-mcp`, `/opt/secrets-issuance`) and only restarts its own service. See [homelab-context](homelab-context.md) for why both services live in one repo. + +### 2026-05-13 — Plato pipeline added +Webhook id 8 on `dtoro/Plato` (port `9799` on [plato (126)](containers/index.md#recently-destroyed-kept-for-archaeology)). `app.ini` `ALLOWED_HOST_LIST` extended to include `192.168.8.190`. + +### 2026-04-28 — wiki entry created +Initial documentation. Six active pipelines. + +### 2026-04-22 — Artifacto pipeline added +Webhook id 7 on `dtoro/Artifacto` (port 9798 on apps). `app.ini` `ALLOWED_HOST_LIST` extended. + +### 2026-06-04 — claudio-bot pipeline decommissioned +LXC 123 destroyed, `dtoro/claudio-bot` archived. Webhook port 9797 dead. + +### 2026-04-21 — mule-image + claudio-bot pipelines added +Webhook id 6; receiver on apps' sibling `/opt/mule-deploy/`. Same shape used for claudio-bot. + +### 2026-04-20 — caddy-conf + gitea-customizations + backup-library pipelines shipped +Initial three. Set the conventions everything else follows. diff --git a/docs/infrastructure/backups.md b/docs/infrastructure/backups.md new file mode 100644 index 00000000..5f5fbc6c --- /dev/null +++ b/docs/infrastructure/backups.md @@ -0,0 +1,142 @@ +# Backups — restic on external drive (DEPRECATED — superseded) + +> **DEPRECATED 2026-07-01.** Superseded by the **rclone → Proton Drive** off-host mirror on +> [LXC 132 `rclone`](containers/132-rclone.md). That job finally closes the off-host / 3-2-1 gap +> this page flagged for months. The restic-on-USB job below is kept for archaeology; it has been +> **DISABLED since 2026-04-22** and is not coming back in its old form. + +## Current backup — rclone → Proton Drive (LXC 132) + +- **Where:** [LXC 132 `rclone`](containers/132-rclone.md) (`192.168.8.214`), `/mnt/library` + mounted **read-only**. +- **What:** plain `rclone sync` (Proton mirrors local; browsable files, no versioning) of the + folders listed in `/etc/rclone-backup/folders.list`, to `proton:library-backup/…`. +- **When:** monthly — `rclone-backup.timer` (`OnCalendar=*-*-01 03:00`). +- **UI:** rclone Web GUI on `192.168.8.214:5572` (LAN-only, no auth). +- **Encryption:** Proton's built-in E2E (no rclone `crypt` overlay). +- **Runs / logs:** `/var/log/rclone-backup/` + `runs.jsonl`. +- **Still a single off-host target** (Proton only). Not yet a full 3-2-1 (no second independent + copy), but strictly better than the previous "no off-host copy at all." + +See [132-rclone](containers/132-rclone.md) for the full design. + +--- + +## Legacy — restic on external drive (DISABLED 2026-04-22) + +Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabled 2026-04-22** as part of the [hubris crash-loop A/B test](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md). + +## Status + +**DISABLED 2026-04-22.** All four timers `systemctl disable --now`'d: +- `backup-library@homecloud.timer` +- `backup-library@images.timer` +- `backup-library@small.timer` +- `backup-library-check.timer` + +Fstab entry commented out. USB drive de-authorized and physically removed. `backup-library-deploy.service` left enabled (harmless webhook receiver). + +**Reason:** the host hang recurred 2026-04-22 18:42 after 30h despite the `cpu-epp` fix, the UAS blacklist, and mount-on-demand. User wants to confirm host stability without the drive at all (was stable 33 days before the drive arrived). See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md). + +**To re-enable:** uncomment fstab line, `systemctl enable --now` the four timers, re-attach drive. + +## Design + +Monthly rolling snapshots onto a 2 TB external USB drive. Chunked across the month so no single run pushes the Samsung 990 EVO Plus 4 TB (which backs `/mnt/library`) into thermal danger. + +Retention per tag: `--keep-last 3 --keep-monthly 12 --keep-yearly 3` with `--group-by host,tags,paths`. + +## Components + +- **Repo:** `dtoro/backup-library` +- **Checkout:** `/opt/backup-library` on the [hubris host](hosts/hubris.md) +- **Auto-deploys** via gitea webhook → `http://192.168.8.77:9798/deploy`. See [auto-deploy](auto-deploy.md). [Gitea (104)](containers/104-gitea.md) `app.ini` `ALLOWED_HOST_LIST` includes `192.168.8.77` for this. +- **Restic repo:** `/mnt/backup/restic-library`. Passphrase `/etc/restic/passphrase` (mode 600). **Escrow in password manager — loss = permanent data loss.** +- **External drive:** `/dev/sda1` ext4 label `backup-library` UUID `ff46e775-1ba1-4892-82c9-e5cac5be933a`. Fstab uses `noauto` + `nofail,x-systemd.device-timeout=10s,errors=remount-ro`. + +## Mount-on-demand + +`/usr/local/sbin/backup-usb.sh attach|detach|status`. All three backup units (`backup-library.service`, `backup-library@.service`, `backup-library-check.service`) have `ExecStartPre=backup-usb.sh attach` and `ExecStopPost=backup-usb.sh detach`. The helper toggles `/sys/bus/usb/devices/*/authorized` by matching vendor:product `090c:2320`, then mounts/unmounts `/mnt/backup`. Drive is de-authorized when not backing up — no UAS keepalive, no kernel error-recovery paths firing against a flaky bridge. + +## UAS blacklist + +`/etc/modprobe.d/usb-storage-quirks.conf`: +``` +options usb-storage quirks=090c:2320:u +``` +Forces Bulk-Only Transport (BOT) instead of UAS for the SMI bridge. Confirm with `dmesg | grep "UAS is ignored"`. + +## Schedule + +Three timers, one per chunk, staggered ~10 days apart so each disk zone gets a long cooldown: + +| Timer | When | Include list | Approx size | +| ---------------------------------- | -------------- | ------------------------------------ | ----------- | +| `backup-library@homecloud.timer` | day 1 / month | `/etc/restic/include-homecloud.list` | ~315 G | +| `backup-library@images.timer` | day 10 / month | `/etc/restic/include-images.list` | ~103 G | +| `backup-library@small.timer` | day 20 / month | `/etc/restic/include-small.list` (docs / books / music / notes / repos / marimo / heaper) | ~11 G | + +Snapshots tagged `chunk-` so forget/prune treats each series independently. + +Ad-hoc full run (kept for manual use): `systemctl start backup-library.service` (no arg → uses `/etc/restic/include.list`, tag `monthly`). + +Yearly integrity: `backup-library-check.timer` (`OnCalendar=yearly`) runs full `restic check --read-data`. + +## Thermal caps + +Baked into the systemd units: +- `IOReadBandwidthMax=/mnt/library 50M` +- `IOWriteBandwidthMax=/mnt/backup 30M` +- `--read-concurrency=1` on restic. + +## Wrapper + +`/usr/local/sbin/backup-library.sh` — preflight → unlock → backup → forget/prune (`--group-by host,tags,paths`) → `check --read-data-subset=5%` → notify. Takes optional `` arg or `GROUP=` env. + +## Notifications + +~~POST to claudio-bot (123) `http://192.168.8.230:9090/notify`~~ — IPC endpoint dead since 2026-06-04. When backups are re-enabled, wire notifications to Hermes `send_message` via Matrix instead. + +`OnFailure=notify-failure@%n.service` on the backup unit fires a synchronous notify as belt-and-suspenders for cases where the wrapper itself died before reaching its own notify. + +## Recovery + +Runbook at `/usr/share/doc/backup-library/RECOVERY.md` (or in the repo at `doc/RECOVERY.md`). Covers `restic snapshots/ls/find/restore/mount`, uid/gid gotcha, cross-host recovery. + +## Known SPOF + +Single drive. RECOVERY.md flags the 3-2-1 gap. Mitigations (second drive, cloud repo via `restic copy`) were not implemented before this job was retired — the **off-host copy is now provided by [rclone → Proton Drive (LXC 132)](containers/132-rclone.md)** instead. A second independent copy is still outstanding. + +## Drive history + +The `Silicon Motion Portable SSD` (vid:pid `090c:2320`) drops under sustained heavy writes through a hub chain. Bypass all hubs / use a rear motherboard USB 3 port if attaching it again. + +After it was first attached on 2026-04-19, hubris crashed twice in 2.5 days (46h then 12h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic, OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1. The UAS blacklist + mount-on-demand mitigations didn't fully eliminate it (recurrence 2026-04-22), prompting drive removal as the cleaner test. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md). + +## Thermal monitoring + +Moved out of this repo to `dtoro/claudio-monitor` on 2026-04-21 (commit `50dc213`). See [monitoring](monitoring.md). + +## Related +- [Hubris host](hosts/hubris.md) +- ~~[claudio-bot (123)](containers/archive/123-claudio-bot.md)~~ (destroyed 2026-06-04) +- [Monitoring](monitoring.md) +- [Auto-deploy](auto-deploy.md) +- [Investigation: 2026-04-21 crash loop](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md) + +## Changelog + +### 2026-07-01 — DEPRECATED; superseded by rclone → Proton Drive (LXC 132) +Off-host backup moved to a plain `rclone sync` mirror on the new [LXC 132 `rclone`](containers/132-rclone.md) (`/mnt/library` → Proton Drive, monthly, LAN Web GUI). This finally provides the off-host copy the "Known SPOF" note wanted. The restic-on-USB units on hubris remain `disabled` (drive already removed 2026-04-22); page restructured to lead with the current job and demote restic to "Legacy". + +### 2026-04-28 — wiki entry created +Initial documentation. Status remains DISABLED. + +### 2026-04-22 — DISABLED +Drive removed as the A/B test in the [crash investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md). Timers disabled, fstab commented, drive de-authorized. + +### 2026-04-21 — UAS blacklist + mount-on-demand shipped; root-caused host hangs to drive +Drive identified as the source of the hangs after hubris crashed twice in 2.5 days. UAS blacklist forces BOT; helper script toggles `/sys/bus/usb/.../authorized` so the drive is de-authorized when not backing up. Recovery drill (restore 188KB PDF + hash compare) had passed earlier. Bug fixed in `backup-library.sh`: `python3 -c '…' KEY=VAL` does NOT pass env vars — env-var prefix must precede the command. Caused false-failure even after successful backups. + +### 2026-04-20 — deployed; redesigned for thermal-gentleness +Initial deploy. First backup attempt died at 18:56 (USB drive dropped off the bus during heavy writes); after re-plugging, restic resumed and completed at 21:21. diff --git a/docs/infrastructure/dns.md b/docs/infrastructure/dns.md new file mode 100644 index 00000000..f093883f --- /dev/null +++ b/docs/infrastructure/dns.md @@ -0,0 +1,137 @@ +# DNS — split-horizon `*.hubris.network` + +LAN clients resolve `*.hubris.network` to the [Caddy reverse proxy](containers/121-caddy.md) (`192.168.8.175`). Public clients resolve to the IONOS VPS (`82.165.190.79`) via an IONOS wildcard, where they hit the [VPS traefik public ingress](ingress.md). + +There is **no wildcard on the LAN side**. Every subdomain needs an explicit entry. + +## Components + +- **Authoritative public DNS:** IONOS. `*.hubris.network → 82.165.190.79` (was `74.118.126.4` until 2026-04-22). +- **LAN authoritative for `hubris.network` records:** [Technitium DNS](https://technitium.com) on [dns (107)](containers/107-dns.md) at `192.168.8.2:53`. Syncs A records to the NetBird managed DNS zone via cron (see [dns-sync.py](../../../scripts/dns-sync.py)). Formerly dnsmasq on [authentik (124)](containers/106-auth-outpost.md) (decommissioned 2026-06-04). +- **PVE host** (`192.168.8.77`): resolver is the local Netbird daemon at `100.122.38.109:53`, which forwards to the LAN/upstream and learns hubris.network answers via that path. `netbird status` says "Nameservers: 0/0 Available" — confirming netbird does NOT manage a hubris.network zone; it just caches whatever the system resolver returns. +- **Some LXCs** keep router DNS (`192.168.8.1`) or Tailscale MagicDNS (`100.100.100.100`), both of which return the public IONOS A record. Those LXCs need either a `/etc/hosts` override or local dnsmasq — see [mesh migration](mesh.md) for which technique applies where. + +## Live entries (as of 2026-06-04) + +``` +address=/auth.hubris.network/82.165.190.79 # → VPS, not Caddy (Authentik migrated 2026-05-31) +address=/sso.hubris.network/192.168.8.175 # → Caddy → LAN forward-auth outpost (106); added 2026-06-01 +address=/git.hubris.network/192.168.8.175 +address=/media.hubris.network/192.168.8.175 +address=/paperless.hubris.network/192.168.8.175 +address=/books.hubris.network/192.168.8.175 +address=/home.hubris.network/192.168.8.175 +address=/cloud.hubris.network/192.168.8.175 +address=/matrix.hubris.network/192.168.8.175 +address=/proxmox.hubris.network/192.168.8.175 +address=/docker.hubris.network/192.168.8.175 +address=/jellyseerr.hubris.network/192.168.8.175 +address=/qbit.hubris.network/192.168.8.175 +address=/sab.hubris.network/192.168.8.175 +address=/blog.hubris.network/192.168.8.175 +address=/photos.hubris.network/192.168.8.175 +address=/photos-new.hubris.network/192.168.8.175 +address=/artifacto.hubris.network/192.168.8.175 +address=/zimaos.hubris.network/192.168.8.175 +address=/nfs-export.hubris.network/192.168.8.200 +``` + +Note: `nfs-export.hubris.network` is the only `.hubris.network` entry that points to a non-HTTP service (NFSv4 on port 2049). It bypasses [caddy (121)](containers/121-caddy.md) because NFS is L4, not HTTP — Caddy has nothing to do. + +## Why split-horizon + +The IONOS wildcard points at the VPS for public ingress (per-host routers in [VPS traefik](ingress.md)). The VPS only routes hostnames it knows — anything else 404s. So LAN clients pointing at the public IP are a dead end for any service that isn't explicitly published. The [Technitium DNS](dns.md) override on `192.168.8.2` keeps LAN traffic on the home Caddy. + +## The gotcha that cost a debug session (2026-04-22) + +Creating a new Caddyfile site block is necessary but **not sufficient**. Without the Technitium entry on [dns (107)](containers/107-dns.md), LAN queries fall through to upstream, get the public IONOS answer, and time out. Symptom: "subdomain doesn't load" even though Caddy config + cert are fine. + +## Recipe — adding a new subdomain + +1. Edit `/etc/caddy/Caddyfile` on [caddy (121)](containers/121-caddy.md), commit + push to `dtoro/caddy-conf`. Webhook reloads caddy. See [auto-deploy](auto-deploy.md). +2. Add the A record in the [Technitium UI](http://192.168.8.2) at `dns (107)` — the NetBird managed DNS zone sync picks it up within ~10 minutes via cron. Or add directly to the NetBird managed zone via API if you need it faster. +3. Verify: `dig @192.168.8.2 +short .hubris.network` → `192.168.8.175`. +4. On macOS clients, flush: `sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder`. + +> The Technitium config on LXC 107 is the single source of truth. Never hand-edit the NetBird managed zone directly — the [`scripts/dns-sync.py`](../../../scripts/dns-sync.py) cron on 107 reconciles them and reaps stale records. See [dns.md changelog 2026-06-03](#2026-06-03--single-authoring-source-technitium--netbird-managed-zone-sync). + +## Public path — what does and doesn't follow the LAN map + +- Hostnames published in [VPS traefik dynamic config](ingress.md) (currently `artifacto.hubris.network`, `blog.hubris.network`) reach a real backend over the netbird mesh. +- Anything else with a `*.hubris.network` URL hits the VPS but isn't routed anywhere — returns 404. +- `netbird.hubris.network` is its own thing — TCP passthrough at the VPS, served by netbird-proxy. Doesn't follow the file-provider router pattern. + +## Long-term plan + +Either: +- Move split-horizon DNS to the LAN router so `*.hubris.network → 192.168.8.175` is answered for every LAN client. Eliminates per-LXC overrides. +- Or, once the [Tailscale → Netbird migration](mesh.md) completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain. + +## Related +- [Caddy (121)](containers/121-caddy.md) — every LAN entry points here +- [Ingress (VPS traefik)](ingress.md) — public-side counterpart +- [Mesh migration](mesh.md) — per-LXC DNS workarounds during the transition +- [DNS server (107)](containers/107-dns.md) — Technitium, current DNS authority + +## Changelog + +### 2026-06-28 — `plato.hubris.network` removed +Plato (LXC 126) decommissioned. Technitium entry deleted; dns-sync cron reaped the NetBird managed zone record. + +### 2026-06-17 — Fritz!Box DNSv4 server set to Technitium; old limitation resolved +Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs — the limitation noted below is resolved. Configured at Fritz!Box Internet → Filter → DNS Server → DNSv4 Server = `192.168.8.2` (User-defined). +Authentik LXC 124 (192.168.8.180) destroyed — Authentik runs on VPS, DNS on Technitium (107). +- Caddy: `auth.hubris.network`, `authentik` snippet, and `sso.hubris.network` all proxied to VPS +- `header_up Host auth.hubris.network` added to strip `:443` from upstream Host header +- All 9 LXCs' /etc/hosts updated: `auth.hubris.network → 192.168.8.175` (Caddy proxy) +- Inventory: removed `hosts.authentik`, renamed `dnsmasq` service → `dns` +- Docs: `124-authentik.md` deleted; dns.md references updated to Technitium (107) +The "delete NetBird managed zone → forward everything to Technitium" plan was **abandoned** — NetBird's DNS defeats it: it **won't apply a nameserver group that contains the peer's own mesh IP** (the Mac's `100.122.234.17` → `Nameservers: 0/0 Available`), and nameserver-group forwarding to Technitium never actually took effect for mesh peers (the **managed zone was doing all the real work**; disabling it broke all mesh resolution). So the model is now: + +- **Technitium (`192.168.8.2`) is the single place you author DNS** (UI/API, MX/SPF/CAA, full zone). +- A **sync job on [dns (107)](containers/107-dns.md)** (`/opt/dns-sync/sync.py`, cron */10) reconciles Technitium's named A-records → the **NetBird managed DNS zone** via the NetBird API (`/api/dns/zones/{id}/records`, PAT in sops `secrets/netbird-pat.yaml`). Mesh peers keep using the managed zone (which works); non-mesh LAN clients query Technitium directly; undefined names fall to the public IONOS wildcard (`.79`) — correct. +- This **killed the manual drift** that caused the whole `auth`/`sso`/`nfs-export` saga. Never hand-edit the NetBird managed zone again — edit Technitium; the sync propagates. + +**Cleanup done same day:** removed the inert Mac-Mini Technitium secondary (mesh-only, served nobody); reverted the primary's `zoneTransfer=Allow`; fixed `home-lab-dns` group → `[192.168.8.2]` (dropped the self-referencing Mac IP → now `1/1 Available`); deleted the vestigial `Proxmox Names` group. + +> Reference: [scripts/dns-sync.py](../../../scripts/dns-sync.py). The sync's source of truth is Technitium; it **deletes** NetBird records absent from Technitium (so obsolete names like `files`, `photos-new` get reaped). + +### 2026-06-06 — dns-sync cron finally installed (had been dormant since 2026-06-04 deployment) +The `dns-sync.py` script on LXC 107 had been placed at `/opt/dns-sync/sync.py` on 2026-06-04 but **no crontab was configured** — the sync had never run automatically. The NetBird managed DNS zone was only in sync because manual runs happened during incident debugging. + +**Fixed:** added `/etc/cron.d/dns-sync` (`*/10 * * * * root python3 /opt/dns-sync/sync.py >> /var/log/dns-sync.log 2>&1`). + +Also added a Caddy backend health check cron on hubris (`/etc/cron.d/caddy-backend-health`) that runs `scripts/check-caddy-backends.sh` every 10 minutes. + +### 2026-06-02 — 8 LXCs moved from DHCP to static IP +All LXCs that Caddy reverse-proxies to by IP were on `ip=dhcp` and could float on reboot (arriman got a different lease mid-session and broke). Fixed via `pct set` + in-LXC `/etc/network/interfaces`. Affected: 101 jellyfin, 103 paperless, 104 gitea, 105 apps, 114 nextcloud, 118 elementsynapse, 120 mule-images, 121 caddy, 122 arriman. See [arriman changelog](containers/122-arriman.md#changelog). + +### 2026-06-01 — dnsmasq replaced by Technitium on [dns (107)](containers/107-dns.md); LXC 124 retired +Split-horizon DNS moved off [124](containers/106-auth-outpost.md) to a dedicated **Technitium** LXC at **`192.168.8.2`** (zone: specific A overrides + wildcard→VPS + replicated MX/SPF/CAA). NetBird `home-lab-dns` nameserver group cut over to `192.168.8.2` (with `.180` as a now-dead fallback). dnsmasq stopped, all names verified via Technitium, **LXC 124 shut down**. **Caveat:** the [NetBird managed DNS zone](containers/106-auth-outpost.md) still answers most app names *directly* (bypassing the nameserver group) — three overlapping DNS sources remain; see the single-source-of-truth decision (Phase 4). **Action needed:** update router DHCP DNS from the dead `.180` → `192.168.8.2` for any plain-LAN (non-mesh) clients. + +### 2026-05-31 — `auth.hubris.network` re-pointed to the VPS (`82.165.190.79`) +Authentik migrated off LXC 124 onto the VPS (see [investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)). The dnsmasq entry changed from `192.168.8.175` (home Caddy) to `82.165.190.79` (VPS traefik). This is the first LAN entry that intentionally points at the VPS rather than Caddy — `auth` is now a genuinely public service served directly from the VPS. **Gotcha logged:** the NetBird per-client resolver (`100.122.255.254`) caches dnsmasq answers and does **not** clear on `netbird down/up`; clients needed `/etc/hosts` overrides or `resolvectl flush-caches` to pick up the change. Since the service is now fully public, the long-term cleaner option is to drop the override entirely and let it fall through to the IONOS wildcard (which also points at the VPS). + +### 2026-05-14 — `nfs-export.hubris.network` added (direct, non-HTTP) +NFSv4 export server [nfs-export (102)](containers/102-nfs-export.md) at `192.168.8.200`. Direct entry, not Caddy-fronted — NFS is L4, no HTTP reverse-proxy meaningful. + +### 2026-05-14 — `zimaos.hubris.network` added (Caddy-fronted, standard pattern) +New LAN entry for [100-zimaos](../vms/100-zimaos.md) → [caddy (121)](containers/121-caddy.md) → `192.168.8.195`. Briefly pointed direct-to-VM during install for the initial smoke-test, then re-pointed once a Caddyfile block was added (`reverse_proxy 192.168.8.195` + IONOS DNS-01 TLS). + +### 2026-05-13 — `plato.hubris.network` added; `files.hubris.network` removed +New LAN-only entry for [plato (126)](containers/index.md#recently-destroyed-kept-for-archaeology). Same day, the `files.hubris.network` entry for the just-decommissioned seafile experiment was dropped; queries now fall through to the public IONOS answer (no LAN backend). + +### 2026-05-12 — `files.hubris.network` added (since removed 2026-05-13) +Originally added for the seafile (LXC 125) Nextcloud-replacement evaluation. Pointed at 192.168.8.175 (Caddy reverse-proxied to 192.168.8.185:80). Entry removed when the experiment was torn down a day later. + +### 2026-04-28 — wiki entry created +Initial documentation. 16 active entries. + +### 2026-04-22 — IONOS wildcard moved 74.118.126.4 → 82.165.190.79 +Public path now lands on the VPS traefik, not the old yunohost. Necessary for the [public ingress](ingress.md) pattern. The LAN dead-end semantics didn't change — public DNS still doesn't help LAN clients reach LAN-only services. + +### 2026-04-22 — three caddy sites without DNS entries (jellyseerr, qbit, sab) +Caddy + certs were working but LAN resolution failed because the dnsmasq lines weren't added. Lesson recorded; entries added later that day. + +### 2026-04-21 — dnsmasq stood up on LXC 124 +Co-located with Authentik. Initial entries cover everything routed through Caddy. diff --git a/docs/infrastructure/homelab-context.md b/docs/infrastructure/homelab-context.md new file mode 100644 index 00000000..82512858 --- /dev/null +++ b/docs/infrastructure/homelab-context.md @@ -0,0 +1,143 @@ +# Homelab context distribution + +The cross-client context-and-secrets system that makes every agent (Claude +Code, Hermes Agent, future MCP-capable clients) on every machine in the lab +self-locating and able to read the same source of truth. + +Operational walkthrough for enrolling a new client lives in +[operations/agent-enrollment.md](../../../.agents/operations/agent-enrollment.md); this +page is the architecture reference. + +## What's where + +| Piece | Host | Path | Role | +| --- | --- | --- | --- | +| Source of truth | [gitea (104)](containers/104-gitea.md) | `dtoro/Homelab-Docs.git` | Inventory + wiki + service code | +| Per-client clone | every enrolled client | `/opt/homelab-context/` | Read by `homelab` CLI, MCP server, Hermes Agent | +| `homelab` CLI | every enrolled client | `/usr/local/bin/homelab` → `/opt/homelab-context/bin/homelab` (symlink) | Operator surface for enroll/secret/ssh/pct | +| Per-client age key | every enrolled client | `/etc/age/key.txt` (0600 root) | Decrypts SOPS-encrypted secrets the client is a recipient on | +| MCP server | [apps (105)](containers/105-apps.md) | `homelab-mcp.service` on port 9810 (https://mcp.hubris.network/mcp) | 14 tools: 8 context (get_host, search_docs, …) + 5 read-only management (get_service_status, tail_log, …) + list_my_secrets | +| Secrets-issuance | [apps (105)](containers/105-apps.md) | `secrets-issuance.service` on port 9820 (https://secrets.hubris.network/issue) | Generates per-client age keypair on first bootstrap; idempotent; admin-token-gated `/revoke` | +| Sync timer | every enrolled client | `homelab-context-sync.timer` (Linux) / `network.hubris.homelab-context-sync.plist` (macOS) | `git pull --ff-only` every 5 min | +| Encrypted secrets | `dtoro/Homelab-Docs` | `secrets/*.yaml` (SOPS+age) | Recipients declared in `.sops.yaml` | +| Read-only context PAT | `dtoro` Gitea user, scope `read:repository` | given to operators out-of-band | Bootstrap-only — for the initial clone before SOPS works | +| Write-scoped PAT | `secrets/gitea-pat.yaml` (SOPS) | `homelab refresh-creds` swaps it into `/etc/homelab-context/git-credentials` | All post-bootstrap pushes (client lifecycle, wiki edits) | + +## Data flow + +``` + dtoro/Homelab-Docs (gitea) + │ + ┌────────── push ────────┤ ◀── git push (write PAT or SSH) + │ │ + │ ┌────── push ──────┘ + │ │ │ + │ │ ▼ webhook (push event) + │ │ ┌─── homelab-mcp-deploy ──── (LXC 105:9811) + │ │ └─── secrets-issuance-deploy (LXC 105:9821) + │ │ │ + │ │ ▼ + │ │ git pull → deploy.sh → restart service + │ │ + │ └── on every client: + │ timer (5 min) → git pull --ff-only into /opt/homelab-context + │ + ▼ + homelab CLI / MCP server reads /opt/homelab-context for everything +``` + +## Why two clones on LXC 105 + +The MCP server and secrets-issuance each have their own clone +(`/opt/homelab-mcp`, `/opt/secrets-issuance`) **in addition to** +`/opt/homelab-context`. Reasons: + +- The deploy webhook for each service updates its own clone, runs + `deploy.sh` from there, and re-installs the systemd unit. Mixing this + with the client-context clone would create a circular dependency + (deploy reinstalls the unit that pulled it). +- The MCP server reads its data from `/opt/homelab-context` (the same path + every client uses) so changes to inventory propagate identically. Code + changes live in `/opt/homelab-mcp` and trigger a service restart. + +## Mesh / network gates + +- Both services bind `0.0.0.0:`. The trust boundary is + `MESH_SUBNETS` in the service's environment + nftables (planned). Today + `MESH_SUBNETS=100.122.0.0/16,100.64.0.0/10,192.168.8.0/24` — Netbird + + Tailscale + the homelab LAN. Adjust if the LAN ever has untrusted + devices. +- Caddy fronts both with Let's Encrypt certs via the IONOS DNS challenge: + `mcp.hubris.network` → `192.168.8.205:9810`, + `secrets.hubris.network` → `192.168.8.205:9820`. Off-LAN clients on + Netbird reach them via the `192.168.8.0/24` network resource routed + through the PVE peer ([mesh.md](mesh.md)). +- Clients with default-public DNS (workstations not on Netbird, LXCs + using router DNS) need a `/etc/hosts` override pointing + `mcp.hubris.network` and `secrets.hubris.network` at the caddy LXC + (`192.168.8.175`) — same caveat as every other `*.hubris.network` + service, see [dns.md](dns.md). + +## Secrets model + +- Each enrolled client gets one **age private key** issued by + secrets-issuance on first bootstrap. The key file stays root-only on + the client; the public key is committed to `inventory.yaml` (and + becomes a recipient on SOPS-encrypted files via `.sops.yaml`). +- `secrets/*.yaml` are SOPS+age. Recipients per file are pinned in + `.sops.yaml` `creation_rules` by `path_regex`. Re-encrypting a file is + `sops updatekeys -y secrets/.yaml`. +- The MCP server's `list_my_secrets(caller_pubkey)` tool returns only + secret *names* a given pubkey can decrypt — the server never sees + plaintext. Decryption is local-on-client (`homelab secret ` + shells out to `sops -d` with the client's key). +- The "all-clients" secrets (`hello.yaml` for the bootstrap decrypt + test, `gitea-pat.yaml` for the write-scoped PAT) are auto-granted to + every newly enrolled client by `homelab client add --finalize-pubkey` + (which appends the pubkey to the matching `.sops.yaml` rule and runs + `sops updatekeys`). +- **Removal does not erase past disclosure.** Revoking a client via + `homelab client remove` shreds the issuance-side key, denylists the + hostname, removes them from the recipient list, and re-keys all + shared secrets — but anything they already decrypted to disk is out of + your control. Rotate the underlying credential if compromise is + suspected. + +## Why this design + +- **One source of truth** keeps inventory, code, secrets, and docs + versioned together. A `git log` of `inventory.yaml` is the history of + the homelab. +- **Per-client age keys** scale better than a shared admin secret — + removing a client is a real revocation (for new ciphertext), not just + removing them from a wiki page. +- **MCP layer over the same clone** gives MCP-capable agents structured + query (`find_service`, `search_docs`) without forcing non-MCP tools to + go without — anything can still `cat` the markdown. +- **Sync timer rather than push fan-out** keeps the failure mode + contained: one client's webhook outage doesn't block a push from + landing on the others. Sub-5-min staleness is fine for docs and rare + enough for secrets that we don't need lower latency. + +## Related + +- [Operations: agent enrollment](../../../.agents/operations/agent-enrollment.md) — the + step-by-step for adding a new client +- [Auto-deploy](auto-deploy.md) — the `homelab-mcp` + `secrets-issuance` + pipelines (and the rest of the lab's webhook pipelines) +- [Mesh](mesh.md) — Netbird / Tailscale paths and the `192.168.8.0/24` + network resource +- [Apps (105)](containers/105-apps.md) — where both services run +- [Gitea (104)](containers/104-gitea.md) — the source of truth + +## Changelog + +### 2026-05-20 — system live across hubris, apps, republic-laptop +Phase 1 of the [cross-client context plan](../../../README.md) merged. Three +clients enrolled end-to-end: PAT-based bootstrap, age-key issuance, SOPS +decrypt verified on each. Webhook auto-deploy for both LXC 105 services +wired (hook ids 10 + 11). `homelab refresh-creds` + atomic +`client add --finalize-pubkey` grant flow live so new clients are one +ceremony instead of four manual steps. Outstanding: bootstrap mac-mini +(macOS, exercises launchd) + ludo-mini + the remaining LXCs; +Hermes Agent integration so the agent uses inventory at chat-time. diff --git a/docs/infrastructure/ingress.md b/docs/infrastructure/ingress.md new file mode 100644 index 00000000..c61d4139 --- /dev/null +++ b/docs/infrastructure/ingress.md @@ -0,0 +1,103 @@ +# Public ingress — VPS traefik + cert mirror + +How home services reach the open internet without exposing the home network. Two-stage pattern: traefik on the IONOS VPS terminates TLS at the public edge, then reverse-proxies over the netbird mesh to home Caddy / direct backends. + +## The shape + +``` +Public client + │ *.hubris.network → 82.165.190.79 (IONOS wildcard) + ▼ +[VPS traefik] priority 1: HostSNI(*) → netbird-proxy:8443 ← netbird control plane + priority 10: per-host HTTP routers ← home services + │ HTTP over netbird mesh + ▼ +[Home backend on 192.168.8.x] +``` + +LAN clients resolve via the [Technitium DNS on dns (107)](dns.md) → `192.168.8.175` → home [Caddy (121)](containers/121-caddy.md), unchanged. The two paths are independent. + +## Why this shape + +- VPS traefik already has a HostSNI(*) TCP passthrough at priority 1 so netbird's own ingress (`netbird.hubris.network`, future `*.proxy.hubris.network`) is unaffected. +- Per-hostname HTTP file-provider routers at priority 10 win over the passthrough for the listed hosts and let traefik terminate TLS itself for those. +- Traefik's own ACME (`letsencrypt` resolver) fails on this box: HostSNI(*) grabs TLS-ALPN-01 challenges before traefik's `allowACMEByPass=true` can respond. Solution: home Caddy obtains certs via IONOS DNS-01 (no conflict) and the VPS *mirrors* the result over. + +## Components + +### On the VPS (`82.165.190.79`) + +- `/opt/traefik-dynamic.yaml` — file-watched dynamic config. One `http.routers.-public` + one `http.services.-public` per exposed service, plus one entry in the top-level `tls.certificates` list per hostname. +- Let's-Encrypt-effective directory: `/letsencrypt/` inside the traefik container, backed by the host-side docker volume `opt_netbird_traefik_letsencrypt`. +- Backups before edits: `cp /opt/traefik-dynamic.yaml /opt/traefik-dynamic.yaml.bak.$(date +%s)` — several bak files live alongside. + +### On the PVE host (`192.168.8.77`) + +- `/usr/local/bin/hubris-public-cert-sync.sh` — runs daily via `hubris-public-cert-sync.timer`. Maps source hostname → VPS cert filenames in a bash assoc array. For each mapping: `pct pull` cert+key from [Caddy (121)](containers/121-caddy.md)'s store, diff against the VPS copy, scp only on change. +- Filenames are stable per host so the dynamic.yaml never needs editing on renewal — traefik file-watches and hot-reloads the cert. + +## Services currently exposed + +| Hostname | Path scope | Backend | Middlewares | Cert files on VPS | +| ------------------------------ | -------------------------------- | -------------------------------- | -------------------------------------------- | ------------------------------------------ | +| `artifacto.hubris.network` | `/p/*`, `/static/*`, `/healthz` | `192.168.8.205:3100` | `artifacto-strip-sso` + `artifacto-ratelimit` (50 rps / 100 burst) | `fullchain.crt` / `privkey.key` | +| `blog.hubris.network` | whole host | `192.168.8.205:8080` | `blog-ratelimit` (100 rps / 200 burst) | `blog.fullchain.crt` / `blog.privkey.key` | +| `trmnl.hubris.network` | whole host | `192.168.8.211:9851` ([trmnl 128](containers/128-trmnl.md)) | `trmnl-ratelimit` (20 rps / 40 burst) | `trmnl.fullchain.crt` / `trmnl.privkey.key` | +| `house.hubris.network` | whole host | `192.168.8.212:3000` ([house 129](containers/129-house.md)) | `house-ratelimit` (30 rps / 60 burst) | `house.fullchain.crt` / `house.privkey.key` | + +`artifacto-strip-sso` blanks inbound `X-Authentik-*` and `X-Artifacto-Gateway` so external clients can't spoof the SSO auto-login header contract. Path split is enforced at the VPS router rule, not by home Caddy. See [Artifacto on apps (105)](containers/105-apps.md). + +### `auth.hubris.network` — different pattern (local container, not cert-mirror) + +Since 2026-05-31 [Authentik runs on the VPS itself](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md), so `auth.hubris.network` is served by a **local Docker container**, not proxied to a home backend. It therefore does **not** use the file-provider + cert-mirror pattern above: + +- Routed via traefik **Docker provider labels** on the `authentik-server` service (`/opt/docker-compose.yml`), not `traefik-dynamic.yaml`. +- TLS via traefik's own `letsencrypt` resolver (works here because it's a normal HTTP router, not the HostSNI passthrough). +- Traefik reaches it over the `auth` Docker network (`172.30.1.0/24`); Postgres/Redis on that net are isolated from the netbird containers. +- Admin UI is IP-gated: an `admin-allowlist` ipAllowList middleware on `PathPrefix(/if/admin/)` (currently `5.61.168.0/24`). Login/flow endpoints stay public. + +No cert-mirror entry and no `hubris-public-cert-sync.sh` mapping is needed for `auth`. + +## Recipe — exposing another service + +1. Ensure home Caddy on [LXC 121](containers/121-caddy.md) already serves the hostname (cert exists at `/var/lib/caddy/.local/share/caddy/certificates/acme-v02.api.letsencrypt.org-directory//`). +2. Add an entry to `HOSTS` in `/usr/local/bin/hubris-public-cert-sync.sh` mapping the hostname → VPS filenames. Run once: `systemctl start hubris-public-cert-sync.service`. Confirm the cert landed. +3. Edit `/opt/traefik-dynamic.yaml` on the VPS: + - Add to `tls.certificates`: paths `/letsencrypt/` and `/letsencrypt/`. + - Add `http.routers.-public`: `rule: 'Host(\`\`)'` (or with path matchers if scope-gating), `entryPoints: [websecure]`, `priority: 10`, `tls: {}`, `service: -public`, `middlewares: [...]`. + - Add a ratelimit middleware under `http.middlewares` if wanted. + - Add `http.services.-public.loadBalancer.servers[0].url: 'http://:'`. +4. Verify: + ``` + ssh root@100.122.165.149 'curl -skI --resolve :443:127.0.0.1 https:///' # 2xx/3xx + curl -skI --resolve :443: https:/// # same + ``` +5. **No DNS edit needed** — the IONOS wildcard already points at the VPS. + +## What does NOT follow this pattern + +- `netbird.hubris.network` (and any future `*.proxy.hubris.network`) uses the netbird-proxy / HostSNI passthrough path. Netbird handles its own cert via ACME cleanly because it *is* the passthrough target. + +## Related +- [DNS split-horizon](dns.md) +- [Caddy (121)](containers/121-caddy.md) — cert source, internal counterpart +- [Mesh migration](mesh.md) — netbird is the transport between VPS and home +- [VPS hardening](vps-hardening.md) — fail2ban / nftables that the access logs feed +- [Artifacto on apps (105)](containers/105-apps.md) — first publicly-exposed service + +## Changelog + +### 2026-06-24 — `trmnl.hubris.network` exposed +TRMNL plugins middleware on [trmnl (128)](containers/128-trmnl.md). File-provider router `trmnl-public` → `192.168.8.211:9851`, `trmnl-ratelimit` (20 rps / 40 burst), cert mirrored as `trmnl.fullchain.crt`/`trmnl.privkey.key`. Verified live from the internet (200 with token / 401 without). It was provisioned during a mesh outage — the `home-lab-network` (192.168.8.0/24) route had no active routing peer because the **mac-mini routing peer's netbird was down** (all home-backed public services 504'd). Bringing netbird up on mac-mini restored the route; no traefik change was needed. + +### 2026-05-31 — `auth.hubris.network` now served locally on the VPS +Authentik migrated onto the VPS ([investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)). Unlike the home-backed services above, `auth` is a local container routed via traefik Docker-provider labels with traefik-managed Let's Encrypt — no cert-mirror, no `traefik-dynamic.yaml` router. Admin UI gated by an ipAllowList middleware. Traefik gained a second Docker network (`auth`, `172.30.1.0/24`) to reach it while keeping its DB/Redis isolated from the netbird stack. + +### 2026-04-28 — wiki entry created +Initial documentation. + +### 2026-04-23 — `blog.hubris.network` exposed +WriteFreely on [apps (105)](containers/105-apps.md). Whole host is public. + +### 2026-04-22 — pattern established with Artifacto +First service through the file-provider router. IONOS wildcard moved to the VPS this day. Cert mirror script + timer deployed on the PVE host. diff --git a/docs/infrastructure/media-permissions.md b/docs/infrastructure/media-permissions.md new file mode 100644 index 00000000..4c912a76 --- /dev/null +++ b/docs/infrastructure/media-permissions.md @@ -0,0 +1,101 @@ +# Media permissions — `media` GID 10000 + +Standard for any LXC reading/writing `/mnt/library` on [hubris](hosts/hubris.md). Applied 2026-04-20. + +## Standard + +Every LXC that mounts `/mnt/library` participates in a shared `media` group with **GID 10000**. Shared subtrees are owned by that group with the setgid bit (`drwxrwsr-x`, mode `2775`), so new files auto-inherit the right group regardless of which container wrote them. + +## Why + +`/mnt/library` is a cross-container storage pool. \*arr writes, jellyfin reads, mulita scans, paperless ingests. Without a shared group, each container sees files as `nobody:nogroup` (unprivileged) or `www-data` (privileged 1:1) and the permission web collapses into one-off chmods. GID 10000 bridges privileged and unprivileged containers. + +## Onboarding a new LXC + +1. `pct set -mp0 /mnt/library,mp=/mnt/library` (if not already mounted). +2. Inside the container: + ``` + groupadd -g 10000 media + usermod -aG media # for every user that needs library access + ``` +3. If the container is **unprivileged** (check `pct config | grep unprivileged`), append this idmap block to `/etc/pve/lxc/.conf` (back up first): + ``` + lxc.idmap: u 0 100000 65536 + lxc.idmap: g 0 100000 10000 + lxc.idmap: g 10000 10000 1 + lxc.idmap: g 10001 110001 55535 + ``` + Then `pct stop && pct start `. +4. For systemd services running with `User=root` (not typical), add a drop-in with `SupplementaryGroups=media`. Systemd skips `initgroups()` for `User=root`. +5. `pct exec` sessions don't get supplementary groups (no initgroups). Use `sudo -i` or `su - ` inside the container to verify membership interactively. Real services use `initgroups` and work correctly. + +## State snapshot + +### Host + +- Group `media` GID 10000 exists. +- `/etc/subgid` has `root:100000:65536` AND `root:10000:1` (second line required for unprivileged LXCs to receive GID 10000). +- Shared subtrees owned `:media` mode `2775` (drwxrwsr-x, setgid): + - `movies`, `tv`, `music`, `anime`, `podcasts` — jellyfin libraries + - `audiobooks`, `audiobookshelf-metadata`, `books`, `comics` — audiobookshelf / grimmory + - `downloads` — \*arr stack output + - `images` — photoprism / immich / mulita + - `roms` — emu frontends + - `syncthing` — empty subtree, retained for archaeology (LXC 109 destroyed 2026-05-14) +- Container-specific subtrees intentionally **not** migrated (keep their own owner:group): + - `documents` (paperless, `www-data:www-data 750`) + - `homecloud` (nextcloud — its own permission model, easy to break) + - `marimo` (marimo venv) — *LXC since destroyed; review whether subtree still serves a purpose* + - `notes`, `sophia` (single-container use); `heaper` — orphaned data subtree (LXC since destroyed 2026-05-14, 224 MiB retained) + - `repos` (owner UID 102 GID 105 from inside [gitea](containers/104-gitea.md) — don't touch) + +### LXCs with media-group membership + +| ID | Name | Priv | Media-group members | +| --- | --------------------------------------------- | ---- | --------------------------------------------- | +| 101 | [jellyfin](containers/101-jellyfin.md) | **unpriv + idmap** | jellyfin | +| 103 | [paperless](containers/103-paperless.md) | priv | www-data | +| 104 | [gitea](containers/104-gitea.md) | priv | www-data, gitea | +| 105 | [apps](containers/105-apps.md) | priv | www-data | +| 114 | [nextcloud](containers/114-nextcloud.md) | priv | www-data | +| 119 | [sophia](containers/119-sophia.md) | priv | www-data | +| 120 | [mule-images](containers/120-mule-images.md) | priv | www-data | +| 122 | [arriman](containers/122-arriman.md) | priv | www-data, audiobookshelf, radarr, sonarr, lidarr, prowlarr, qbittorrent, bazarr, jellyseerr, mylar, jackett, overseerr, plex, arr | +| 130 | [grimmory](containers/130-grimmory.md) | priv | Docker container uses `GROUP_ID=10000` env var (linuxserver pattern) — no in-LXC group needed | +| 132 | [rclone](containers/132-rclone.md) | priv | **read-only** mount; runs as root → reads all subtrees. No media group needed | + +> Some entries from earlier snapshots — 100 (arr-yunohost), 107 (marimo), 109 (syncthing), 110 (photoprism), 112 (immich), 116 (heaper) — referenced LXCs that have since been destroyed. See [containers/index](containers/index.md#recently-destroyed-kept-for-archaeology). + +Config backups: `/root/101.conf.bak.*`, `/root/109.conf.bak.*` (109 destroyed 2026-05-14). + +## Gotchas + +- **[apps (105)](containers/105-apps.md) and [grimmory (130)](containers/130-grimmory.md) are Docker hosts.** Adding `media` to the LXC alone is *not* enough for Docker containers inside. Each Docker container needs its GID passed in explicitly: `--group-add 10000`, `user: ":10000"`, or `GROUP_ID=10000` (linuxserver images) in compose. Grimmory, audiobookshelf-in-docker, etc. need this per-container. +- **`pct exec` does NOT run initgroups.** So `pct exec -- id` shows only the primary group. For interactive verification, use `pct exec -- sudo -i -u root id` or `su - -c id`. Real systemd services work fine. +- **systemd `User=root`** skips initgroups — explicit `SupplementaryGroups=media` drop-in needed. +- **`pct restore`** or template rebuilds wipe in-container group membership and unprivileged-LXC idmap blocks. Re-apply from this page. +- **`/etc/subgid`** must retain both `root:100000:65536` AND `root:10000:1`. Dropping the second breaks startup of any unprivileged LXC with the idmap block. +- **\*arr "Set Permissions" options** can override the setgid inheritance by explicitly chown'ing files. Leave those off, or set the group to `media`. Relevant to Sonarr/Radarr/qBittorrent on [arriman (122)](containers/122-arriman.md). +- **Nextcloud** files under `/mnt/library/homecloud` are deliberately NOT in the media group. NC manages its own permission model. See [nextcloud (114)](containers/114-nextcloud.md). +- **\*arr stack on arriman** required `MEDIACENTER_GID=10000` (not 13000) in `.env` because s6-setuidgid only honors the primary PGID; `group_add:` doesn't propagate. See [arriman (122)](containers/122-arriman.md#changelog). + +## Related +- [Hubris host](hosts/hubris.md) +- All container pages list whether they're in the standard + +## Changelog + +### 2026-05-14 — LXC 109 (syncthing) destroyed +Removed the syncthing row from the membership table and the syncthing-as-`User=root` example from the onboarding section. `/mnt/library/syncthing` subtree was already empty and retained as an empty dir. + +### 2026-05-14 — LXC 116 (heaper) destroyed +Removed the heaper row from the LXC membership table and noted the orphaned `/mnt/library/heaper` subtree (224 MiB retained). See [host changelog](hosts/hubris.md#changelog). + +### 2026-04-28 — wiki entry created +Initial documentation. + +### 2026-04-26 — `MEDIACENTER_GID` fix on [arriman (122)](containers/122-arriman.md) +qBit was erroring every torrent with "Permission denied" because `MEDIACENTER_GID=13000` was set as a supplementary GID via `group_add:`. Changed to 10000 (primary GID); fix described above is now standard. + +### 2026-04-20 — standard rolled out +GID 10000 hostgroup, idmap blocks for unprivileged LXCs, setgid 2775 on shared subtrees, `media` membership for service users in every participating LXC. diff --git a/docs/infrastructure/mesh.md b/docs/infrastructure/mesh.md new file mode 100644 index 00000000..6cf7878d --- /dev/null +++ b/docs/infrastructure/mesh.md @@ -0,0 +1,194 @@ +# Mesh — Tailscale → Netbird migration + +The hubris fleet is migrating from Tailscale to Netbird. Netbird is the target end-state. In-progress as of 2026-04-21. + +## Current state + +- **PVE host** uses Netbird (`wt0`, `100.122.38.109/16`). Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets `*.hubris.network → 192.168.8.175` via the system resolver chain. +- **Netbird mgmt host** (`82.165.190.79`, FQDN `inspiring-ramanujan.netbird.selfhosted`, NB IP `100.122.165.149`) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes the homelab network (`192.168.8.0/24`) via the PVE peer. This gives the mgmt host LAN access *and* split-horizon DNS for `*.hubris.network`. Useful independently of any Authentik integration. +- **Most LXCs** still run Tailscale or use router DNS (`192.168.8.1`) / Tailscale MagicDNS (`100.100.100.100`), both of which return the *public* IONOS A record `*.hubris.network → 82.165.190.79`. The VPS only routes hostnames it actually publishes (today, `artifacto` + `blog`), so this path is a dead end for any LAN-only service. + +## ICE / STUN / TURN + +**Today** (post-2026-05-21 migration): + +- External STUN servers (Google + Cloudflare) declared under `server.Stuns` in `/opt/management.json`. Embedded STUN is no longer in use. +- coturn runs on the VPS host (apt-installed, systemd-managed). TURN URL `turn:netbird.hubris.network:3478?transport=tcp` is advertised to peers via mgmt's `TURNConfig.Turns` block. Long-term credentials at `user=netbird:`. +- For non-symmetric peers, ICE picks direct `srflx/srflx` (P2P). For symmetric-NAT peers, ICE falls through to TURN-relay before falling back to the WSS relay (`rels://netbird.hubris.network:443`). + +**IONOS port-3478 caveat** (load-bearing — undocumented before 2026-05-21): + +IONOS upstream filters STUN-class traffic on port 3478 **for both UDP and TCP** by default. The UDP block was already known (verified 2026-05-10 with `tcpdump -ni any udp port 3478`). The TCP block was discovered 2026-05-21 when TURN-allocate retries timed out from a symmetric-NAT laptop: TCP handshake completed (carried by kernel-only SYN exchange) but data packets never reached the VPS's `ens6` interface. + +**Resolution**: operator must add an inbound firewall exception in the IONOS dashboard (Cloud Panel → Networks → Firewall policy on the VPS) for inbound TCP 3478 (and UDP 3478 if STUN-via-VPS-listener is ever needed; we use external STUN so this is optional). Without that exception, coturn is unreachable from outside even though host nftables + listener look fine. + +**Verifying TURN works** end-to-end from an outside peer: + +```python +# python3 +import socket, struct, secrets +s = socket.create_connection(("netbird.hubris.network", 3478), timeout=10) +tid = secrets.token_bytes(12) +attrs = struct.pack("!HHI", 0x0019, 4, 17 << 24) # REQUESTED-TRANSPORT = UDP +msg = struct.pack("!HHI", 0x0003, len(attrs), 0x2112A442) + tid + attrs # Allocate +s.sendall(msg) +print(s.recv(4096).hex()) # expect 120 bytes starting 0x0113 (Allocate Error = auth challenge) +``` + +A 120-byte `0x0113`-typed response means coturn is reachable and responding. An indefinite TCP timeout (after `connected from ...` prints) means the IONOS rule was reverted or narrowed. + +**If a peer is still on the WSS relay after this** — verify in `netbird status -d`: `Connection type: P2P` with srflx/srflx is the cone-NAT happy path; `Connection type: P2P` with relay/srflx (or relay/prflx) means the peer's network is symmetric-NAT and it's using TURN. Falling back to `Relayed (rels://...)` only happens if TURN allocate also fails — investigate the IONOS exception first. + +**Old combined-server note (history, kept for context):** + +Pre-migration, the bundled `netbirdio/netbird-server` combined image silently ignored both `server.turns:` and top-level `TURNConfig:` YAML, so TURN was non-functional. That's why the architectural migration to the canonical multi-container stack (`mgmt + signal + relay + coturn`) happened. See the 2026-05-21 changelog entry below and the [Netbird-combined-no-TURN finding](https://github.com/netbirdio/netbird/issues) in homelab memory for the full discovery. + +## Consequence — every LXC wired to Authentik needs an internal override + +Until each LXC is migrated to Netbird, anything that needs to reach `auth.hubris.network` (Authentik), `cloud.hubris.network` (Nextcloud), etc., must override the public answer with `192.168.8.175`. + +Two techniques. Pick by HTTP-client behavior. + +### A) `/etc/hosts` override + +Works for libc `getaddrinfo` clients: curl, wget, most Go/Python/Ruby apps, gitea. + +- Place the line **outside** the `# --- BEGIN PVE ---` / `# --- END PVE ---` markers. Proxmox rewrites everything inside that block on every container start. +- Belt-and-suspenders: `/etc/systemd/system/hubris-hosts-override.service` (oneshot, enabled, idempotent). + +### B) Local dnsmasq + +Required for clients that bypass `/etc/hosts`. **Nextcloud (PHP Guzzle + `OC\Http\Client\DnsPinMiddleware`) is one** — uses `dns_get_record()`, not `getaddrinfo`. Likely candidates for the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting. + +Recipe: +``` +apt install dnsmasq + +cat > /etc/dnsmasq.d/hubris-internal.conf < --nameserver "127.0.0.1 192.168.8.1 1.1.1.1" +# Then update /etc/resolv.conf inside the LXC too. +``` + +### Known overrides applied + +| LXC | Technique | Notes | +| ------------------------------------------ | ---------------------------------------- | ----- | +| [104 (gitea)](containers/104-gitea.md) | `/etc/hosts` + `hubris-hosts-override.service` | Standard | +| [114 (nextcloud)](containers/114-nextcloud.md) | local dnsmasq @ 127.0.0.1 + resolv.conf reorder | Guzzle bypasses /etc/hosts | +| [105 (apps)](containers/105-apps.md), inner containers | `extra_hosts:` in compose | Booklore, mulita, WriteFreely each ship with this | + +## Adding new LXCs + +- Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris. +- When wiring a new app into Authentik: `cat /etc/resolv.conf` on the target LXC. If nameserver is `192.168.8.1` or `100.100.100.100`, add the hosts override. If it's the netbird daemon IP, skip. + +## Long-term fix + +Either: +- Split-horizon DNS at LAN/router level so `*.hubris.network → 192.168.8.175` for every LAN client. Eliminates all per-LXC overrides. +- Once Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain on the PVE host. + +## CRITICAL — never `docker compose up` Portainer-managed stacks + +[apps (105)](containers/105-apps.md) runs multiple stacks deployed via the Portainer UI (`/var/lib/docker/volumes/portainer_data/_data/compose//`). Running `docker compose up -d ` from the host shell on these triggers a recreate of OTHER services in the stack — labels drift, compose treats the project as reconciling, containers get destroyed and remade. **This wiped Booklore's mariadb data on 2026-04-22** (bind mount `./mariadb/config` re-initialized fresh). + +Recipe for container-config changes (e.g. adding `extra_hosts`) on Portainer-managed stacks: +1. Edit the compose in Portainer UI → **Stacks → → Editor → Update stack**. Portainer handles the recreate cleanly with its own state tracking. +2. Do NOT edit `/var/lib/docker/volumes/portainer_data/_data/compose//docker-compose.yml` directly. +3. For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the app's built-in OIDC over container-level `extra_hosts` when possible. + +## Related +- [DNS split-horizon](dns.md) +- [Authentik (124)](containers/106-auth-outpost.md) — the IdP that triggers most of these overrides +- [Nextcloud (114)](containers/114-nextcloud.md) — example of Technique B +- [Gitea (104)](containers/104-gitea.md) — example of Technique A +- [Public ingress (VPS traefik)](ingress.md) — uses the same mesh as transport + +## Changelog + +### 2026-05-31 (later) — Authentik moved to the VPS; mesh-dependency for auth eliminated (supersedes the band-aid below) +The earlier same-day fix routed `auth.hubris.network` through VPS Traefik → Caddy → LXC 124 **over the mesh**. That restored service but re-created the original fragility: if the mesh is dark when management restarts, the `192.168.8.175` backend is unreachable and management crash-loops again (the "Bootstrap note" in the entry below). That note is now **obsolete** — Authentik was migrated onto the VPS itself, so OIDC no longer touches the mesh. The `auth-authentik` → `192.168.8.175` route and its `skip-verify` transport were removed from `/opt/traefik-dynamic.yaml`; `auth.hubris.network` is now served by a local `authentik-server` container via Traefik Docker-provider labels, and netbird-mgmt has `depends_on: authentik-server: condition: service_healthy`. The socat / reverse-SSH bootstrap dance is no longer needed. Full detail: [2026-05-31 Authentik VPS migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md). + +### 2026-05-31 — Netbird mesh recovered; auth.hubris.network exposed via VPS Traefik + +**Symptom:** `netbird-mgmt` crash-looped for ~9 days (since 2026-05-21 migration). All peers showed `Connecting`, management returned `404 (Not Found)` for gRPC → `EOF` on startup. + +**Root cause:** The 2026-05-21 migration configured `management.json` with `OIDCConfigEndpoint: https://auth.hubris.network/...`. The public DNS for `auth.hubris.network` already pointed to `82.165.190.79` (VPS), but the VPS Traefik had no route for that host → management got EOF on every boot, crash-looped. + +**Fix:** +1. Added `auth-hubris` router to `/opt/traefik-dynamic.yaml` — `Host(auth.hubris.network)` with `certResolver: letsencrypt` → service `auth-authentik`. +2. Service backend: `https://192.168.8.175` (Caddy on hubris LAN) via the Netbird mesh. +3. ServersTransport `skip-verify` with `serverName: auth.hubris.network` + `insecureSkipVerify: true` so Traefik sends the correct SNI to Caddy. +4. Upgraded VPS Netbird client `0.69.0 → 0.71.3` via apt (`apt install netbird=0.71.3`). +5. Mesh fully recovered; management connected to peers within ~1 min. + +**Bootstrap note (if mesh is dark and management must restart):** If the WireGuard tunnels are fully dead AND management needs to restart (after e.g. a VPS reboot with no peer handshakes), the Traefik backend `192.168.8.175` will be unreachable and management will crash again. Recovery: temporarily open VPS port 22 via IONOS console (`nft insert rule inet hubris-fw input iifname "ens6" tcp dport 22 accept`), SSH in, run `socat TCP-LISTEN:8443,bind=172.30.0.1,reuseaddr,fork TCP:127.0.0.1:8443 &`, then from hubris `ssh -f -N -R 127.0.0.1:8443:192.168.8.175:443 root@82.165.190.79` — this bootstraps one management start, after which the mesh self-heals. + +**Architecture after this change:** `auth.hubris.network` is publicly accessible (HTTPS via VPS Traefik → Caddy on hubris → Authentik LXC 124). External devices authenticating to Netbird now hit this public path. Phase 6 (Authentik as Netbird IdP) is complete and live. + +### 2026-05-21 — VPS migrated combined → vanilla netbird stack with external TURN +The combined `netbirdio/netbird-server` image was replaced with the canonical multi-container deploy (`netbirdio/management:0.71.3` + `signal:0.71.3` + `relay:0.71.3` + `dashboard:latest` + host coturn) on `/opt/docker-compose.yml`. Driver: combined image silently ignored external `TURNConfig` so symmetric-NAT peers couldn't use TURN. + +Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on [LXC 124](containers/106-auth-outpost.md), upgrading mgmt to 0.71.3. The `store.db` schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old `opt_netbird_data` volume into the new `mgmt_data` volume). Pre-cutover backups at `/root/netbird-*.tgz` on the VPS, ~857 MB, retained for ~7d. + +Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478 inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe). + +The new Authentik provider for NetBird is `Client type: Public` (PKCE-only). Confidential would break the dashboard SPA's token exchange. The Device Code grant flow is wired (see [containers/124-authentik.md](containers/106-auth-outpost.md#device-code-grant--configured-2026-05-21)) so interactive `netbird up` works — `--setup-key` is no longer required for new peers. + +**Post-migration JWT-issuer gotcha on existing peers** (cost ~30 min to diagnose 2026-05-21): + +Existing peers — registered against the old combined image's embedded Dex IdP at `https://netbird.hubris.network/oauth2` — cache the OLD expected SSH-JWT issuer in the netbird daemon's in-memory state. After the migration, incoming `netbird ssh` connections were rejected with: + +``` +JWT authentication failed: validate token ( + expected issuer=https://netbird.hubris.network/oauth2, + audiences=[netbird-dashboard netbird-cli], + actual issuer=https://auth.hubris.network/application/o/netbird/, + audience=netbird-dashboard +) +``` + +Neither `systemctl restart netbird` nor `netbird down && netbird up` clears the cache. Root cause: in `client/internal/engine_ssh.go`, `updateSSH()` bails out with `if e.sshServer != nil { return nil }` whenever the SSH server is already running, so mgmt-pushed JWT config updates are silently ignored. Only a full daemon-process tear-down lets the SSH server re-initialize with the new validator config: + +``` +sudo systemctl stop netbird +sleep 3 +sudo systemctl start netbird +``` + +After that, `grep -iE "issuer|audience" /var/log/netbird/client.log | tail` shows the new Authentik issuer. Run this on every existing peer (PVE host + every LXC + every workstation) once after a future IdP swap. + +**Username gotcha (related):** `netbird ssh` defaults the remote username to the LOCAL one (e.g. `dtoro` from the operator's laptop). Hubris + the LXCs only have `root`, so the JWT is accepted but the session immediately fails with `user dtoro not found`. Always use the explicit `root@` prefix when invoking netbird-ssh manually: + +``` +netbird ssh -p 22022 root@proxmox-server.netbird.selfhosted +``` + +The `homelab` CLI handles this automatically via the per-host `ssh.user` field in `inventory.yaml` (defaults to `root`; set explicitly only for workstations whose login user isn't `root`). + +Open follow-up: TURN-over-TLS on TCP 5349 (cert via certbot or extract Traefik's acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient for current usage. + +### 2026-05-10 — ICE direct p2p restored (external STUN swap) +All peers were `Connection type: Relayed` because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in `/opt/config.yaml`. After `netbird down/up`, peers now report `P2P` with srflx/host candidates. Backup of pre-change config at `/opt/config.yaml.bak-20260510-185050`. Fixes a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers). + +### 2026-04-28 — wiki entry created +Initial documentation. + +### 2026-04-22 — Booklore mariadb data wiped (lesson recorded) +The "never docker compose up Portainer-managed stacks" rule comes from this. See [apps (105)](containers/105-apps.md#changelog). + +### 2026-04-22 — netbird mgmt host joined its own mesh +`82.165.190.79` is now a peer (`100.122.165.149`). LAN access + split-horizon DNS via PVE peer. See [VPS hardening](vps-hardening.md). + +### 2026-04-21 — overrides applied to gitea (104) and nextcloud (114) +Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle. diff --git a/docs/infrastructure/monitoring.md b/docs/infrastructure/monitoring.md new file mode 100644 index 00000000..721d6148 --- /dev/null +++ b/docs/infrastructure/monitoring.md @@ -0,0 +1,60 @@ +# Monitoring — Hermes health watchdog + +Homelab health monitoring via Hermes Agent on mac-mini. Replaced the legacy +`claudio-monitor` + `claudio-bot` IPC pipeline on 2026-06-04. + +## Current approach + +Two layers: + +1. **On-demand:** ask Hermes "how's the homelab?" or run `homelab health` — loads + the `homelab-hardware-health` skill, checks hardware temps, LXC resources, + service reachability, and apt/docker drift across all hosts. + +2. **Cron watchdog:** `homelab-health-watchdog` runs every 15 minutes via Hermes + cron. Silent when healthy. When thresholds breach, sends an actionable alert + to Matrix (`@dtoro:avispero`) with options the user can reply to directly + (e.g. "resize rootfs", "investigate", "snooze 24h"). Hermes takes action on + the selected option via SSH. + +Thresholds: LXC disk >80% warn/>90% critical, NVMe >60°C/>70°C, CPU >70°C/>80°C, +apt >10/>50 upgradable, services down. + +Home Assistant pulls PVE metrics independently via its Proxmox VE integration +(unaffected by this change). + +## Legacy: claudio-monitor (deprecated 2026-06-04) + +The old system was a bash watchdog on hubris (`claudio-monitor.timer`, every 5 +min) that POSTed alerts to a Matrix bot (`@claudio:avispero`) via an IPC server +on LXC 123:9090. All components decommissioned: + +| Component | Fate | +|-----------|------| +| LXC 123 (claudio-bot) | Destroyed 2026-06-04 | +| `dtoro/claudio-bot` | Archived (read-only) on Gitea | +| `dtoro/claudio-monitor` | Archived (read-only) on Gitea | +| `claudio-monitor.timer` | Disabled on hubris | +| `/opt/claudio-monitor/` | Still on hubris (cleanup pending) | +| `/etc/claudio-monitor/` | Still on hubris (cleanup pending) | + +For the full deprecation plan, see `plans/2026-06-04_130000-deprecate-claudio-bot.md`. + +## Related pages +- [Hubris host](hosts/hubris.md) +- [HAOS VM (108)](../vms/108-haos.md) +- [Backups (disabled)](backups.md) +- [Homelab context distribution](homelab-context.md) + +## Changelog + +### 2026-06-04 — migrated to Hermes health watchdog +claudio-monitor + claudio-bot IPC pipeline replaced by Hermes-native monitoring. +On-demand `homelab health` via extended skill; 15-min cron watchdog with actionable +Matrix alerts. LXC 123 destroyed, repos archived. + +### 2026-04-28 — wiki entry created +Initial documentation. + +### 2026-04-21 — claudio-monitor stood up; thermal-watch removed +General health monitor with per-LXC checks. MQTT/REST push paths ripped out. diff --git a/docs/infrastructure/network.md b/docs/infrastructure/network.md new file mode 100644 index 00000000..517be947 --- /dev/null +++ b/docs/infrastructure/network.md @@ -0,0 +1,88 @@ +# Network + +Physical and logical network topology for the homelab. + +## Why + +The homelab runs on a dedicated internal subnet (`192.168.8.0/24`) isolated from the main household LAN (`192.168.178.0/24`). Isolation is enforced at Proxmox: LXC/VM traffic is bridged only on the internal `vmbr0` bridge; Proxmox routes packets out to Fritz!Box via `vmbr1`. The main LAN cannot reach homelab services directly without a Fritz!Box static route (which is configured to allow inbound). + +Fritz!OS 8.x does not support second IP networks on LAN ports, so Proxmox (`hubris`) acts as the subnet router rather than the Fritz!Box. + +## Hardware + +| Device | Role | +|---|---| +| Fritz!Box 7590 | Main router / ISP gateway (`192.168.178.1`) | +| SODOLA 5-Port 2.5Gbit Managed | Homelab switch — flat L2, all ports native | +| hubris (Proxmox) | Subnet router — routes between `192.168.8.0/24` and `192.168.178.0/24` | + +## Topology + +``` +ISP + └── Fritz!Box 7590 (192.168.178.1) + │ static route: 192.168.8.0/24 → 192.168.178.10 + │ + └── SODOLA 5-Port 2.5Gbit + ├── Port 1 uplink → Fritz!Box LAN + ├── Port 2 hubris eno1 → vmbr1 (192.168.178.10) + ├── Port 3 [device] + ├── Port 4 [device] + └── Port 5 spare + +hubris internal bridges: + vmbr1 192.168.178.10/24 eno1 (uplink, DHCP-reserved) gateway 192.168.178.1 + vmbr0 192.168.8.77/24 no physical port (internal) + 192.168.8.1/24 alias — LXC default gateway + ├── all 16 LXCs + └── HAOS VM +``` + +## Subnets + +| Subnet | Gateway | Purpose | +|---|---|---| +| `192.168.178.0/24` | `192.168.178.1` | Household LAN — laptops, phones, Fritz!Box DHCP | +| `192.168.8.0/24` | `192.168.8.1` (Proxmox `vmbr0` alias) | Homelab — all LXCs and VMs | + +## DHCP + +- **Household (`192.168.178.x`)**: Fritz!Box built-in DHCP. Proxmox `vmbr1` has a reservation: MAC `84:47:09:6b:e7:58` → `192.168.178.10`. +- **Homelab (`192.168.8.x`)**: Technitium on [CT 107](containers/107-dns.md) at `192.168.8.2`. Range `192.168.8.241–192.168.8.254`, gateway `192.168.8.1`, DNS `192.168.8.2`. + +Static IPs span `.101–.239` (all LXCs, VMs, and workstations). DHCP pool narrowed to `.241–.254` (2026-06-03) to avoid overlap and IP conflicts. + +## DNS + +Split-horizon DNS for `*.hubris.network` served by Technitium on [CT 107](containers/107-dns.md) at `192.168.8.2:53`. See [dns.md](dns.md) for full detail. + +## Routing + +Proxmox has `net.ipv4.ip_forward=1` (already enabled by PVE). Packets from LXCs on `vmbr0` destined for the internet exit via `vmbr1` → Fritz!Box. Fritz!Box masquerades all outbound WAN traffic. Fritz!Box has a static route (`192.168.8.0/24 → 192.168.178.10`) so return traffic reaches the LXCs. + +No NAT on Proxmox — traffic flows without double-NAT. + +## Remote access + +- **NetBird mesh** — primary path for remote administration. Authenticated via [Authentik on the VPS](../../../vps/). +- **Tailscale** — legacy, being phased out. See [mesh.md](mesh.md). + +## Related + +- [DNS](dns.md) — split-horizon config and entry list +- [Ingress](ingress.md) — public entry points via VPS traefik +- [Mesh](mesh.md) — NetBird / Tailscale VPN overlay +- [hosts/hubris.md](hosts/hubris.md) — Proxmox host (vmbr0/vmbr1 config) +- [CT 107 — dns](containers/107-dns.md) — Technitium DNS + DHCP server + +## Changelog + +### 2026-06-17 — Fritz!Box DNSv4 server set to Technitium (192.168.8.2) +Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs. Configured in Fritz!Box at Internet → Filter → DNS Server → DNSv4 Server → "Use other DNSv4 servers" → Preferred = `192.168.8.2`. No per-device or Netbird setup needed. +Previous pool `.100–.240` overlapped with all static LXCs/VMs (` .101–.239`), creating IP conflict risk (DHCP could hand out an IP that a static service expects). Shrunk pool to `.241–.254` via Technitium API. No services re-IP'd. 11 stale DHCP leases in `.101–.110` will expire naturally. **Open:** ZimaOS (VM 100) holds DHCP lease `.103` but inventory expects `.195` — needs static IP set inside VM. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md). + +### 2026-06-02 — Executed migration; Proxmox as subnet router +Fritz!OS 8.x does not support second IP networks on LAN ports, so the final design uses Proxmox as the router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10`; `vmbr0` is a portless internal bridge with `192.168.8.1` alias as the LXC gateway. Technitium DHCP enabled for `192.168.8.100–240`. Caddy service unit was missing and recreated. See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md). + +### 2026-06-01 — Initial network doc; Slate AX retired; SODOLA switch added +Replaced the GL.iNet Slate AX sub-router with the SODOLA 5-Port 2.5Gbit managed switch. Eliminated double-NAT. See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md). diff --git a/docs/infrastructure/oikos-check-lifecycle.md b/docs/infrastructure/oikos-check-lifecycle.md new file mode 100644 index 00000000..7f0e9303 --- /dev/null +++ b/docs/infrastructure/oikos-check-lifecycle.md @@ -0,0 +1,104 @@ +# Oikos check lifecycle — how monitoring works + +This runbook covers how Oikos health checks are derived, created, and wired so +an agent (Nomos) doesn't reverse-engineer source when asked to add monitoring to +an entity — the problem that stranded session `23da10db` (2026-08-03). + +## Concepts + +- **`check_defs`** (scheduler config, table `check_defs`): the row the scheduler + reads to know *what* to probe and *when*. One per check instance. +- **`check` entity** (type `check`, slug `check:::`): the + knowledge-graph entity for that check. It carries attributes + (`check_type`, `target`, `port`, …) and `checks` edges to the probed target. +- **`monitoring` spec** on an entity type (`entity_types.monitoring_spec`): the + default list of check kinds (e.g. `[http, process]` for `service`). + - Per-entity override: set `monitoring` in the entity's attributes — + `"none"` for zero checks, `["http"]` to replace the type defaults. +- **`checkdefaults.Ensure`** (`internal/checkdefaults/defaults.go`): the + function that reads the monitoring spec, resolves host/port/URL from + attributes + relationships, and writes `check_defs` rows. Idempotent. + +## When checks are derived + +`checkdefaults.Ensure` runs in three situations (as of v0.17.1+): + +1. **Seed/deploy ingest** — `internal/db/seed.go:231`. Every entity gets its + default checks once on initial ingest. +2. **HTTP `POST /api/v1/entities` (create)** — `ensureDefaultChecks` at + `internal/httpapi/impl.go:1012`. Creating an entity via the REST API derives + its checks in the same transaction. +3. **HTTP `PATCH /api/v1/entities` (patch)** — `ensureDefaultChecks` at + `internal/httpapi/impl.go:1280`. Changing an entity's attributes (especially + `monitoring`) via the REST API regenerates its checks. +4. **MCP `create_entity`** — SAME hook. Creating an entity via the MCP tool + derives checks. (Added 2026-08-03; previously MCP had no create.) +5. **MCP `update_entity_attributes`** — SAME hook. Changing an entity's + `monitoring` attribute via MCP now regenerates checks. (Added 2026-08-03; + previously MCP updates silently skipped check derivation — the exact bug + that stranded the haos session.) + +## Check slug grammar + +``` +check:::: +``` + +Examples: `check:http:service:jellyfin:0`, `check:vm-status:vm:haos:0`, +`check:cert-expiry:cert:house.hubris.network:0`. + +## Adding monitoring to an entity + +**If the entity already exists:** + +``` +update_entity_attributes(slug="service:haos", attributes={"monitoring":["http"]}) +``` + +This regenerates checks via `checkdefaults.Ensure`. The result message tells you +how many checks were derived and whether any kinds were skipped (and why). + +**If the entity does not exist yet (a new check, ingress, cert, etc.):** + +``` +create_entity(type="check", name="HAOS http check", + slug="check:http:service:haos:0", + attributes={"check_type":"http:service","target":"service:haos","port":"8123"}) +``` + +This creates the entity AND derives its `check_defs`. Same for a new `ingress` +(`type=ingress`, monitoring `[http]`) or `cert` (`type=cert`, +monitoring `[cert-expiry]`). + +**To remove monitoring:** set `monitoring:["none"]` or transition the entity +to a terminal lifecycle state (`set_entity_state` → `deprecated`/`destroyed`). + +## Caveats + +- **A service without a `url` attribute AND without a `probe_unit` gets no + process check** (the http check covers liveness; the process check would + be redundant without an opt-in `probe_unit`). The skip is logged. +- **A service whose address comes from a `hosts` edge** may produce no checks on + initial create because the edge doesn't exist yet — the next inventory ingest + (or a later `update_entity_attributes` after the edge is created) fills it in. +- **A `not found` error from `update_entity_attributes`** means the entity + doesn't exist — use `create_entity` instead. +- **`check_defs` has target columns** (`target_id`, `target_type`). A check + entity needs a `checks` relationship (`create_relationship(source=check:…, + target=service:…, type="checks")`) so the scheduler can resolve what to + probe. `create_entity` derives the check_def; `create_relationship` links + the check entity to its target in the graph. + +## Related files + +- `internal/checkdefaults/defaults.go` — `Ensure`, `Target`, `LogResult` +- `internal/httpapi/default_checks.go` — `ensureDefaultChecks` (HTTP hook) +- `internal/db/checks.go` — `db.EnsureEntityChecks` (shared hook) +- `internal/db/seed.go` — seed-time check derivation +- `internal/mcp/tools.go` — `create_entity`, `update_entity_attributes` + +## Revision history + +- **2026-08-03:** Created after session `23da10db` stranded for lack of entity- + creation tool and unawareness of check-derivation triggers. Covers the MCP + create_entity + update_entity_attributes regen paths added same day. diff --git a/docs/infrastructure/ssh-access.md b/docs/infrastructure/ssh-access.md new file mode 100644 index 00000000..4386d361 --- /dev/null +++ b/docs/infrastructure/ssh-access.md @@ -0,0 +1,206 @@ +# SSH access + +How to reach every host in the fleet from any workstation, with LAN as +the primary path and Netbird as the automatic backup. + +## Architecture + +SSH access relies on three layers: + +1. **Homelab inventory (`inventory.yaml`)** — the single source of truth + for every host's LAN IP, Netbird addresses, SSH user, and port. +2. **Key distribution (`ssh/deploy-keys.sh`)** — deploys workstation SSH + public keys to hubris and every running LXC, so any key-authorized + workstation can log in anywhere. +3. **Config generation (`homelab ssh-config --install`)** — generates + `~/.ssh/config.d/homelab` with short hostname aliases for every host, + using LAN IPs (routed via Netbird's `192.168.8.0/24` subnet route when + off-LAN) with Netbird FQDN fallbacks (`-mesh`) for roaming + workstations. + +### How it works + +- **From on-LAN:** `ssh gitea` resolves to `192.168.8.121` directly. +- **From off-LAN (Netbird):** The same `192.168.8.121` works because + hubris routes the `192.168.8.0/24` subnet through Netbird. +- **Roaming workstations:** `ssh mac-mini-mesh` or `ssh republic-laptop-mesh` + uses the Netbird FQDN as a fallback when the workstation is off its + home subnet. + +The `homelab ssh ` CLI command also has built-in LAN probing: +it tries a 1.5s TCP connect to the LAN IP, and if that fails, falls +back to the Netbird FQDN. + +## Key distribution + +Each workstation's SSH public key lives in the repo at: +`ssh/authorized_keys/.pub` + +To deploy or re-deploy all workstation keys to hubris + every running LXC: + +```bash +# From hubris (or via homelab pct): +sudo bash /opt/homelab-context/ssh/deploy-keys.sh + +# Or from any workstation: +ssh root@192.168.8.77 "bash /opt/homelab-context/ssh/deploy-keys.sh" +``` + +This script: +- Reads all `.pub` files from `ssh/authorized_keys/` +- Adds any missing keys to `/etc/pve/priv/authorized_keys` on hubris +- For each running LXC, appends keys to `/root/.ssh/authorized_keys` +- Is idempotent — skips keys already present + +## Config generation + +To generate the SSH config on any workstation: + +```bash +homelab ssh-config --install +``` + +This writes to `~/.ssh/config.d/homelab` and ensures +`Include ~/.ssh/config.d/homelab` is present in `~/.ssh/config`. + +The config is regenerated automatically on every `homelab sync` (which +kicks the 5-minute context sync timer). + +## Adding a new workstation + +When onboarding a new machine: + +1. Hostname must match an entry in `inventory.yaml`. +2. If the workstation will be on the LAN, add its `lan_ip` to + `inventory.yaml` and push. This gives it a primary LAN entry in the + generated SSH config. +3. Enable SSH Remote Login: + - **macOS:** `sudo launchctl load -w /System/Library/LaunchDaemons/ssh.plist` + - **Linux:** `sudo systemctl enable --now sshd` +4. Generate an SSH keypair if one doesn't exist: + ```bash + ssh-keygen -t ed25519 -a 100 + ``` +5. Publish the public key to the repo: + ```bash + cp ~/.ssh/id_ed25519.pub /opt/homelab-context/ssh/authorized_keys/.pub + cd /opt/homelab-context && git add ssh/authorized_keys/ && git commit -m 'ssh: add pubkey' && git push + ``` +6. Deploy the key to all hosts: + ```bash + ssh root@192.168.8.77 "cd /opt/homelab-context && git pull --ff-only && bash ssh/deploy-keys.sh" + ``` +7. Generate the local SSH config: + ```bash + homelab ssh-config --install + ``` + +## Hosts + +### Hubris + strong (PVE cluster: `Homelab`) + +Both nodes share `/etc/pve/priv/authorized_keys` — it's Proxmox +cluster-synced, so a key added on either node is authorized on both. + +| Detail | hubris | strong | +|--------|--------|-----------| +| LAN IP | `192.168.8.77` | `192.168.178.181` | +| Cluster node name | `hubris` | `strong` (OS hostname kept as-is from install) | +| Netbird | `100.122.38.109` (`proxmox-server.netbird.selfhosted`) | not enrolled yet | +| Netbird SSH port | `22022` (mesh-only, OIDC auth) | n/a | +| SSH user | `root` | `root` | + +Authorized root keys currently deployed (cluster-wide): +- `root@hubris` (self, RSA) +- `d.toro.v@pm.me` (ed25519) — mac-mini +- `root@strong` (RSA) — strong's own key, added 2026-07-01 for the cluster join + +### LXCs + +Every LXC at `192.168.8.x` accepts root SSH via authorized_keys. Keys +are managed by `ssh/deploy-keys.sh`. SSH user is `root`. + +| LXC | Name | LAN IP | Role | +|-----|------|--------|------| +| 101 | jellyfin | `192.168.8.206` | media-server | +| 102 | nfs-export | `192.168.8.200` | storage-export | +| 103 | paperless | `192.168.8.130` | document-archive | +| 104 | gitea | `192.168.8.121` | git-server | +| 105 | apps | `192.168.8.205` | docker-apps | +| 106 | auth-outpost | `192.168.8.184` | authentik-outpost | +| 107 | dns | `192.168.8.185` | dns-helper | +| 114 | nextcloud | `192.168.8.224` | file-sync | +| 118 | elementsynapse | `192.168.8.239` | matrix-server | +| 119 | sophia | `192.168.8.157` | workshop | +| 120 | mule-images | `192.168.8.136` | photo-management | +| 121 | caddy | `192.168.8.175` | reverse-proxy | +| 122 | arriman | `192.168.8.132` | arr-stack | + +### Workstations + +| Name | OS | LAN IP | Netbird FQDN | SSH user | +|------|----|--------|--------------|----------| +| mac-mini | macOS | `192.168.8.174` | `mac-mini-234-17.netbird.selfhosted` | `dtoro` | +| republic-laptop | Linux | TBD | `republic-laptop.netbird.selfhosted` | `dtoro` | + +strong moved out of this table 2026-07-01 — it's a Proxmox host now, see the cluster table above. + +### VPS (external) + +| Detail | Value | +|--------|-------| +| Public IP | `82.165.190.79` | +| Netbird | `100.122.165.149` (FQDN: `netbird-ionos.netbird.selfhosted`) | +| SSH user | `root` | +| Access | Mesh-only — public port 22 is blocked by nftables. Key-only auth. | + +## VPS + +Access is mesh-only. From a mesh-connected peer: + +```bash +ssh root@100.122.165.149 +ssh root@netbird-ionos.netbird.selfhosted +# or via homelab: +homelab ssh netbird-vps +``` + +## Verification + +```bash +# From any workstation after running homelab ssh-config --install: +for name in hubris gitea apps sophia paperless caddy jellyfin nextcloud; do + ssh -o BatchMode=yes "$name" "hostname" && echo "$name OK" +done +``` + +## Related + +- [Mesh migration](mesh.md) +- [VPS hardening](vps-hardening.md) +- [Agent enrollment](../../../.agents/operations/agent-enrollment.md) +- [Homelab CLI](../../../bin/homelab) + +## Changelog + +### 2026-07-01 — strong reformatted to Proxmox, joined cluster; table corrected +strong moved from the Workstations table to the PVE-cluster table (was showing a stale `192.168.8.133`, never actually reachable — the real LAN IP has always been `192.168.178.181`, matching hosts/strong.yaml). Root key access bootstrapped via one-time console password, then key-only going forward. See [hosts/hubris.md#cluster](hosts/hubris.md#cluster) and [hosts/strong.md](hosts/strong.md). + +### 2026-06-02 — universal SSH reachability + +Replaced ad-hoc per-workstation SSH configs with inventory-generated +configs (`ssh/gen-config.py`, `homelab ssh-config`). Added centralized +key distribution (`ssh/deploy-keys.sh`, `ssh/authorized_keys/`). All +LXCs now accept root SSH from any workstation whose pubkey is in the +repo. mac-mini Remote Login enabled. Netbird subnet route +(192.168.8.0/24 via hubris) provides off-LAN reachability for all LAN +IPs. + +### 2026-04-28 — wiki entry created +Initial documentation. + +### 2026-04-23 — VPS SSH hardened to mesh-only +Public `:22` blocked at nftables. Key-only sshd. + +### 2026-04-22 — iMac key authorized on hubris +`d.toro.v@pm.me` added to `/etc/pve/priv/authorized_keys`. \ No newline at end of file diff --git a/docs/infrastructure/vps-hardening.md b/docs/infrastructure/vps-hardening.md new file mode 100644 index 00000000..5703b7db --- /dev/null +++ b/docs/infrastructure/vps-hardening.md @@ -0,0 +1,91 @@ +# VPS hardening — `82.165.190.79` / `100.122.165.149` + +IONOS VPS that runs the Netbird control plane and the [public ingress traefik](ingress.md). Hardened 2026-04-23 from its stock-Plesk state. + +## At a glance +- **Hostname:** `inspiring-ramanujan.82-165-190-79.plesk.page` +- **OS:** Debian 13 +- **Mesh:** netbird `100.122.165.149` (peer of the lab mesh; routes `192.168.8.0/24` via [hubris](hosts/hubris.md)). +- **Public:** `82.165.190.79` (`ens6`). +- **Public DNS:** IONOS wildcard `*.hubris.network → 82.165.190.79`. +- **Docker stack** at `/opt/docker-compose.yml`: `traefik` (TLS/ACME) + `dashboard` + `mgmt` + `signal` + `relay` + `proxy` — netbird-mgmt 0.71.3 vanilla deploy since 2026-05-21 (see [mesh.md changelog](mesh.md#changelog)). +- **Host services (outside docker):** `coturn` (TURN-TCP on :3478, long-term creds rendered into `/etc/turnserver.conf` by `homelab render-vps-configs` from sops-encrypted `secrets/turn-shared-secret.yaml`). +- **Config rendering:** `/etc/turnserver.conf` + `/opt/management.json` are generated from templates in `vps/*.tmpl` on this repo by `homelab render-vps-configs`. Secret placeholders (`{{TURN_PASSWORD}}`, `{{AUTHENTIK_CLIENT_SECRET}}`) are substituted from sops-encrypted secrets decrypted on hubris and pushed over ssh. **Do not hand-edit those two files on the VPS** — the next render will overwrite them. + +## SSH + +- Key-only (`PasswordAuthentication no`, `PermitRootLogin prohibit-password`) via drop-in at `/etc/ssh/sshd_config.d/10-hubris-hardening.conf`. Original config backed up at `/etc/ssh/sshd_config.bak.`. +- **Mesh-only**: public `:22` is dropped by the nftables firewall. SSH reaches the VPS only over `wt0`. `ListenAddress` itself is still `0.0.0.0` — gating is firewall-layer. +- Authorized root keys: PVE (`root@hubris`), Mac Mini (`d.toro.v@pm.me`). Add a new device with `ssh-copy-id root@100.122.165.149` from a mesh peer **before** disabling its access paths. + +## Firewall — nftables (`inet hubris-fw`) + +Config at `/etc/nftables.conf`, service enabled. + +- Public iface `ens6`. Wireguard iface `wt0`. +- **INPUT on `ens6`** allow-list: DHCP (67→68), rate-limited ICMP/ICMPv6, **TCP 3478** (coturn TURN-TCP, added 2026-05-21). Everything else drops. +- `wt0` fully accepted in INPUT. `lo` accepted. +- **IONOS upstream firewall** also gates inbound traffic before it reaches `ens6`. Open ports today: TCP 80/443 (traefik), UDP 51820 (netbird-proxy), TCP 3478 (coturn, added 2026-05-21). UDP 3478 is dropped by IONOS upstream regardless of local nftables. See [mesh.md ICE/STUN](mesh.md) for the STUN/TURN port matrix. +- **FORWARD chain at priority `filter-10`** (runs before Docker's FORWARD) hosts the fail2ban ban enforcement — see below. +- Set `banned4` (typed `ipv4_addr`, flag `timeout`) holds fail2ban's drops. +- Coexists with Docker's `ip nat` / `ip filter` tables (iptables-nft compat). **Do NOT `flush ruleset`** in this config — it'll wipe Docker's state too. + +## fail2ban + +- **Jail `traefik-4xx`** tails `/var/log/traefik/access.log` (bind-mounted from container). Filter at `/etc/fail2ban/filter.d/traefik-4xx.conf` matches 401/403/404/429 from `blog-public@file` or `artifacto-public@file` routers only — netbird-grpc traffic isn't considered. +- Tunables: `findtime=600, maxretry=30, bantime=3600`. +- **Action** at `/etc/fail2ban/action.d/nft-hubris.conf` adds/removes elements from `inet hubris-fw banned4` with per-element timeout. + +### CRITICAL invariant — wireguard / fail2ban + +**Bans must never affect `wt0` or wireguard UDP.** The FORWARD chain explicitly `accept`s the following *before* the ban check: +- `udp 51820` (wireguard) +- `udp 3478` (STUN) +- `ct state established,related` + +The INPUT ban rule is scoped to `iifname "ens6"`. + +Violating this takes the mesh down for every home device (they share one public IP) and the only recovery is IONOS console → `nft flush set inet hubris-fw banned4`. + +## Traefik access log + +- Written to `/var/log/traefik/access.log` on the host via a bind mount added to `/opt/docker-compose.yml` (traefik volumes include `/var/log/traefik:/logs`) plus `--accesslog.filepath=/logs/access.log`. +- CLF format. Real client IP arrives correctly because docker userland-proxy is off — see [public ingress](ingress.md). + +## Plesk / mail / FTP / Dr.Web + +Stopped and disabled (not uninstalled). All of: +`dovecot`, `dovecot.socket`, `postfix`, `postfix@-`, `pc-remote`, `xinetd`, `plesk-task-manager`, `plesk-web-socket`, `sw-cp-server`, `sw-engine`, `plesk-repaird`, `plesk-repaird.socket`, `drwebd`. + +`psa.service` is masked (was a one-shot boot bootstrap). `/etc/cron.d/plesk-backup-manager-task` renamed to `.disabled`. + +Reverse: `systemctl unmask psa; systemctl enable --now `. + +## Auto-patching + +- `unattended-upgrades` enabled (stock). +- Drop-in at `/etc/apt/apt.conf.d/52hubris-reboot.conf` sets auto-reboot at **04:00 UTC** when `/var/run/reboot-required` is set. +- Runs inside the stock `apt-daily-upgrade.timer`. + +## Recovery paths + +Ordered by preference: +1. **SSH via mesh** — primary. Any mesh peer with an authorized key. +2. **IONOS web console** (my.ionos.com → VPS → Console) — uses the system password, not SSH keys. Bypasses any firewall misconfig. +3. **IONOS rescue mode** — boot rescue, mount rootfs, edit `/etc/nftables.conf` or `/etc/ssh/sshd_config.d/10-hubris-hardening.conf` to a known-good state, reboot. + +## Related +- [Public ingress (VPS traefik)](ingress.md) +- [Mesh migration](mesh.md) — VPS as a mesh peer +- [SSH access](ssh-access.md) + +## Changelog + +### 2026-05-21 — netbird stack migrated combined → vanilla; coturn added +Replaced the `netbirdio/netbird-server` combined image with the canonical `mgmt + signal + relay + dashboard` containers (0.71.3). Added host-side `coturn` for external TURN, with nftables rule `iifname "ens6" tcp dport 3478 accept` and an IONOS upstream firewall exception. Authentik on LXC 124 now provides OIDC for the netbird dashboard. Full context in [mesh.md changelog](mesh.md#changelog). + +### 2026-04-28 — wiki entry created +Initial documentation. + +### 2026-04-23 — hardened +nftables firewall, mesh-only SSH, fail2ban traefik jail, Plesk disabled, auto-reboot 04:00 UTC, wireguard/fail2ban invariant established. diff --git a/docs/secrets/README.md b/docs/secrets/README.md new file mode 100644 index 00000000..f6a5e8a6 --- /dev/null +++ b/docs/secrets/README.md @@ -0,0 +1,63 @@ +# secrets/ + +SOPS-encrypted YAML files. The plaintext lives only in transit and in the +operator's head — committed files are always ciphertext. + +## Conventions + +- One file per logical grouping (e.g. `gitea-tokens.yaml`, `webhook-hmacs.yaml`, + `api-keys.yaml`). +- Recipients are declared in `../../.sops.yaml` by path-regex, not per-file. +- The plaintext schema inside each file is free-form YAML; the consumer code + decides what it expects (e.g. `gitea-tokens.yaml` contains + `{"": "ghp_xxx"}`). + +## How to add a secret + +```bash +# 1. Decide which clients should be able to decrypt it; edit ../../.sops.yaml to +# list their age public keys for the new path_regex. +# 2. Create the plaintext, encrypt in place: +sops -e --in-place secrets/my-thing.yaml +# 3. Commit + push. The 5-min sync propagates to every recipient. +``` + +## How to consume a secret + +```bash +# On any client that's a recipient: +homelab secret my-thing # prints plaintext +# Or programmatically: +sops -d /opt/homelab-context/secrets/my-thing.yaml +``` + +The `mcp` tool `list_my_secrets(caller_pubkey)` returns the names of secrets +the caller can decrypt. The MCP server never reads plaintext — decryption +stays client-side. + +## Granting / revoking access + +To grant a new recipient: edit `../../.sops.yaml` to add their age pubkey, then +re-key every affected file: + +```bash +sops updatekeys -y secrets/my-thing.yaml +``` + +To revoke: remove the recipient from `../../.sops.yaml` and `sops updatekeys` — +but remember this only protects future ciphertext. Past plaintext the client +already decrypted is gone from your control. Rotate the underlying credential +if compromise is suspected. + +`homelab client remove ` does the recipient removal + `updatekeys` for +you, and prints the rotation checklist as a follow-up. + +## hello.yaml — bootstrap decrypt test + +`secrets/hello.yaml` is encrypted to every enrolled client. Used by Phase 3a +verification to confirm the end-to-end decrypt path works on a freshly- +bootstrapped machine. Content is intentionally trivial: + +```yaml +greeting: hello from the homelab +``` diff --git a/docs/secrets/rotation.md b/docs/secrets/rotation.md new file mode 100644 index 00000000..53e2440b --- /dev/null +++ b/docs/secrets/rotation.md @@ -0,0 +1,100 @@ +# Secret rotation runbook (Phase 5) + +Rotation cadences per secret type. All rotation is automated via Infisical; +this runbook covers the manual verification and DR procedures. + +## Rotation schedule + +| Secret | Cadence | Method | +|--------|---------|--------| +| OpenRouter API key | 90 days | Infisical rotation policy → update `OPENROUTER_API_KEY` env | +| MCP bearer token | 30 days | Infisical random password generation | +| Approval HMAC secret | 90 days | Infisical random password generation | +| Matrix access token | 90 days | Manual (Matrix does not support automated rotation) | +| Age DR key (SOPS fallback) | Never | Static — stored offline for DR only | + +## How to rotate a secret + +### Automated (Infisical) +```bash +# Secrets managed by Infisical rotate automatically per the policy above. +# To force an immediate rotation: +infisical secrets rotate --project-id $INFISICAL_PROJECT_ID \ + --secret-name --env dev + +# Verify the new value is available: +oikos secret list +``` + +### Manual (SOPS fallback) +```bash +# If Infisical is unavailable, use the SOPS DR fallback: +sops -d secrets/.yaml + +# To rotate a SOPS secret: +sops -e --in-place secrets/.yaml # edit in place +``` + +## Rotation verification + +After any rotation, verify the consuming services still work: + +```bash +# 1. OpenRouter key: test Hermes query +curl -s -X POST http://localhost:8092/query \ + -H "Content-Type: application/json" \ + -d '{"query":"fleet health"}' + +# 2. MCP bearer token: test MCP connection +curl -s -X POST http://localhost:8092/query \ + -H "Content-Type: application/json" \ + -d '{"tool":"list_entities","args":{"limit":1}}' + +# 3. Approval HMAC: create a test execution +curl -s -X POST http://localhost:8092/query \ + -H "Content-Type: application/json" \ + -d '{"query":"restart caddy"}' +``` + +## Disaster recovery + +If Infisical is completely unavailable: + +```bash +# 1. Export SOPS DR fallback +oikos secret export-sops > /tmp/sops-dr-backup.txt + +# 2. Configure services to use SOPS fallback +# Set OIKOS_SECRETS_DIR=/opt/homelab-context/secrets +# This switches the secrets manager to SOPS-only mode. + +# 3. Restart services +docker compose restart api hermes notifier scheduler +``` + +## Restore drill + +Run monthly: + +```bash +# 1. Export all secrets from Infisical +oikos secret list + +# 2. Simulate Infisical outage: stop the container +docker compose stop infisical + +# 3. Verify SOPS fallback works +OIKOS_SECRETS_DIR=./secrets oikos secret list + +# 4. Restore Infisical +docker compose start infisical +sleep 5 + +# 5. Verify Infisical primary works again +oikos secret list +``` + +## Changelog + +### 2026-07-07 — initial rotation runbook +Phase 5 rotation cadences, verification steps, and DR restore drill.