diff --git a/.agents/skills/docs-lint/SKILL.md b/.agents/skills/docs-lint/SKILL.md index e6a8a9a..b1f296c 100644 --- a/.agents/skills/docs-lint/SKILL.md +++ b/.agents/skills/docs-lint/SKILL.md @@ -20,6 +20,6 @@ Run from the repo root: Exit code is non-zero when any violation is found, so it can gate a commit. The banned-vocabulary list mirrors `writing-style.md`; update both together if the standard changes. -> **Known baseline.** The lab carries pre-existing broken links to destroyed/archived nodes (e.g. -> `124-authentik.md`, now `106-auth-outpost`). Clean those opportunistically; do not treat the -> current count as a regression from this skill. +> **Known baseline.** `knowledge/wiki/containers/101-jellyfin.md` links into a sibling repo +> (`devops/homelab-authentik-admin`) that this checkout does not contain — expected, not a bug. +> Any other broken link is a real regression; investigate before dismissing it as baseline noise. diff --git a/.agents/skills/lifecycle-migrate-node/SKILL.md b/.agents/skills/lifecycle-migrate-node/SKILL.md index e8a6e52..b4a29e6 100644 --- a/.agents/skills/lifecycle-migrate-node/SKILL.md +++ b/.agents/skills/lifecycle-migrate-node/SKILL.md @@ -10,7 +10,7 @@ transition: "active -> migrating -> active" # Lifecycle: migrate a node Modeled on the strong Phase 1+2 migration -([plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md](../../../plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)). +([plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)). Requires (ontology): preflight + backup-verified before migrating; post-verify + Caddy backends checked + mounts checked + docs updated before returning to `active`. diff --git a/.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md b/.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md index 3d0d9b6..c6def9c 100644 --- a/.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md +++ b/.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md @@ -12,7 +12,7 @@ ## Problem statement -The Technitium DHCP server on [CT 107](containers/107-dns.md) serves `192.168.8.100–192.168.8.240`. **Every static homelab IP except hubris (`.77`) sits inside that range:** +The Technitium DHCP server on [CT 107](../../knowledge/wiki/containers/107-dns.md) serves `192.168.8.100–192.168.8.240`. **Every static homelab IP except hubris (`.77`) sits inside that range:** | Host | IP | Inside pool? | |---|---|---| diff --git a/investigations/2026-06-03-moonlight-sunshine-wifi-jitter.md b/investigations/2026-06-03-moonlight-sunshine-wifi-jitter.md index f5a6a06..256f81d 100644 --- a/investigations/2026-06-03-moonlight-sunshine-wifi-jitter.md +++ b/investigations/2026-06-03-moonlight-sunshine-wifi-jitter.md @@ -2,7 +2,7 @@ ## Summary -[`ludo-mini`](../hosts/ludo-mini.yaml) runs Sunshine as the game-streaming server; [`mac-mini`](../hosts/mac-mini.yaml) runs Moonlight as the client. Despite both machines being on the same physical subnet (192.168.178.0/24), streaming was unstable — stuttering, dropouts, and high latency. Root cause: **mac-mini is connected only via WiFi**, while ludo-mini is wired Ethernet (2.5 Gbps). WiFi throughput shows 1-second UDP dropouts and high jitter (28 ms stddev), which breaks real-time video streaming. +[`ludo-mini`](../hosts/strong.yaml) runs Sunshine as the game-streaming server; [`mac-mini`](../hosts/mac-mini.yaml) runs Moonlight as the client. Despite both machines being on the same physical subnet (192.168.178.0/24), streaming was unstable — stuttering, dropouts, and high latency. Root cause: **mac-mini is connected only via WiFi**, while ludo-mini is wired Ethernet (2.5 Gbps). WiFi throughput shows 1-second UDP dropouts and high jitter (28 ms stddev), which breaks real-time video streaming. ## Timeline diff --git a/investigations/2026-06-06-authentik-session-lifetime.md b/investigations/2026-06-06-authentik-session-lifetime.md index faf9613..76adcad 100644 --- a/investigations/2026-06-06-authentik-session-lifetime.md +++ b/investigations/2026-06-06-authentik-session-lifetime.md @@ -91,7 +91,7 @@ print("session_duration:", stage.session_duration) # → "days=30" ## Related - [Container 106 — auth-outpost](../knowledge/wiki/containers/106-auth-outpost.md) -- [Authentik VPS migration](2026-05-31-authentik-vps-migration.md) +- [Authentik VPS migration](archive/2026-05-31-authentik-vps-migration.md) - [Ingress (VPS Traefik)](../knowledge/wiki/infrastructure/ingress.md) - `.hermes/plans/2026-06-06_232200-authentik-frequent-login-fix.md` — original plan diff --git a/investigations/2026-06-06-caddyfile-truncation.md b/investigations/2026-06-06-caddyfile-truncation.md index cf3203a..516d87f 100644 --- a/investigations/2026-06-06-caddyfile-truncation.md +++ b/investigations/2026-06-06-caddyfile-truncation.md @@ -54,7 +54,7 @@ This is the same class of drift as the June 5th incidents (paperless, HAOS, apps ## Related -- [DHCP drift investigation (previous incident)](2026-06-05-homelab-dhcp-drift.md) +- DHCP drift investigation (previous incident) — not filed as its own investigation; see the [DNS sync fix](../.hermes/plans/2026-06-05_170000-prevent-dhcp-ip-drift.md) - [Caddy (121)](../knowledge/wiki/containers/121-caddy.md) - [elementsynapse (118)](../knowledge/wiki/containers/118-elementsynapse.md) - [dns-sync script](../scripts/dns-sync.py) diff --git a/investigations/archive/2026-04-21-hubris-crash-loop.md b/investigations/archive/2026-04-21-hubris-crash-loop.md index 7fbcbf3..4466649 100644 --- a/investigations/archive/2026-04-21-hubris-crash-loop.md +++ b/investigations/archive/2026-04-21-hubris-crash-loop.md @@ -2,12 +2,12 @@ ## Summary -[`hubris`](../hosts/hubris.md) hard-locked repeatedly on 2026-04-21 (silent CPU hangs, no panic, no OOM, no MCE). Two contributors identified: idle CPU sitting at ~95 °C on the `performance` governor, and a USB-attached external SSD whose UAS interaction with the AMD USB4/Thunderbolt PCIe tunnel triggered hard locks. CPU thermal addressed via `cpu-epp.service`; drive removed 2026-04-22 as an A/B test. As of 2026-04-28 the host has 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders. +[`hubris`](../../knowledge/wiki/hosts/hubris.md) hard-locked repeatedly on 2026-04-21 (silent CPU hangs, no panic, no OOM, no MCE). Two contributors identified: idle CPU sitting at ~95 °C on the `performance` governor, and a USB-attached external SSD whose UAS interaction with the AMD USB4/Thunderbolt PCIe tunnel triggered hard locks. CPU thermal addressed via `cpu-epp.service`; drive removed 2026-04-22 as an A/B test. As of 2026-04-28 the host has 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders. ## Timeline ### 2026-04-19 — drive attached -External `Silicon Motion Portable SSD` (vid:pid `090c:2320`) attached for the new restic [backup pipeline](../infrastructure/backups.md). Pre-attach uptime had been 33 days stable. +External `Silicon Motion Portable SSD` (vid:pid `090c:2320`) attached for the new restic [backup pipeline](../../knowledge/wiki/infrastructure/backups.md). Pre-attach uptime had been 33 days stable. ### 2026-04-19 → 2026-04-21 — first crashes Two hard crashes in 2.5 days (46 h then 12 h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic, OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1. @@ -24,13 +24,13 @@ Two hard crashes in 2.5 days (46 h then 12 h uptime). Kernel logs ended abruptly - **Mount-on-demand** for the drive: `/usr/local/sbin/backup-usb.sh attach|detach|status` toggles `/sys/bus/usb/devices/*/authorized` so the drive is de-authorized when no backup is running. ### 2026-04-22 — recurrence after 30 h 37 m -Same silent-cutoff signature at 18:42:08. Much longer than any pre-`cpu-epp` crash (12 h max), so `cpu-epp` helps but is not sufficient on its own. [claudio-monitor](../infrastructure/monitoring.md) showed healthy runtimes up to 43 s before the hang (no pre-crash degradation). No MCE / no RAS / pstore empty. +Same silent-cutoff signature at 18:42:08. Much longer than any pre-`cpu-epp` crash (12 h max), so `cpu-epp` helps but is not sufficient on its own. [claudio-monitor](../../knowledge/wiki/infrastructure/monitoring.md) showed healthy runtimes up to 43 s before the hang (no pre-crash degradation). No MCE / no RAS / pstore empty. ### 2026-04-22 — `cpu-epp.service` design bug fixed Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `pve-guests.service`. The hottest window of every boot (20 LXCs + 1 VM coming up) ran on the `performance` governor. Fixed: now `After=sysinit.target` + `Before=pve-guests.service`. ### 2026-04-22 — drive removed (A/B test) -User physically removed the external USB drive. [Backup timers disabled](../infrastructure/backups.md#status), fstab entry commented, drive de-authorized. Goal: confirm whether the drive + UAS + AMD USB4 PCIe-tunnel interaction is the dominant root cause. +User physically removed the external USB drive. [Backup timers disabled](../../knowledge/wiki/infrastructure/backups.md#status), fstab entry commented, drive de-authorized. Goal: confirm whether the drive + UAS + AMD USB4 PCIe-tunnel interaction is the dominant root cause. ### 2026-04-23 — SSD cooling + thermal pads installed Cold-boot baseline (3 min uptime): nvme0n1 35 °C composite / sensor1 (controller) **53 °C**; nvme1n1 36 °C composite / both sensors ≤36 °C. Lifetime warning-time counters at install: nvme0n1 709 min warn + 5 min crit; nvme1n1 778 min warn + 45 min crit — both drives had spent real time in thermal warning historically. @@ -78,9 +78,9 @@ Checked 2026-04-21. GMKtec is **not on LVFS**, so `fwupdmgr` can't update the Nu | `pcie_aspm=off pci=nomsi` | NOT applied | Reserved for if crashes recur without the drive | ## Affected nodes -- [Hubris host](../hosts/hubris.md) -- [Backups (disabled)](../infrastructure/backups.md) -- [Monitoring](../infrastructure/monitoring.md) +- [Hubris host](../../knowledge/wiki/hosts/hubris.md) +- [Backups (disabled)](../../knowledge/wiki/infrastructure/backups.md) +- [Monitoring](../../knowledge/wiki/infrastructure/monitoring.md) ## Open questions - Will the host stay up indefinitely without the drive? (Test ongoing — 3+ days as of 2026-04-28.) diff --git a/investigations/archive/2026-05-31-authentik-vps-migration.md b/investigations/archive/2026-05-31-authentik-vps-migration.md index 972e98d..ba91978 100644 --- a/investigations/archive/2026-05-31-authentik-vps-migration.md +++ b/investigations/archive/2026-05-31-authentik-vps-migration.md @@ -2,9 +2,9 @@ ## Summary -The NetBird management server (on the [VPS](../infrastructure/ingress.md)) crash-looped 1200+ times because it fetches the Authentik OIDC discovery document on startup, and Authentik was only reachable via the NetBird mesh — which was down *because* mgmt couldn't start. A classic bootstrap deadlock: **mgmt needs OIDC → OIDC needs the mesh → the mesh needs mgmt.** +The NetBird management server (on the [VPS](../../knowledge/wiki/infrastructure/ingress.md)) crash-looped 1200+ times because it fetches the Authentik OIDC discovery document on startup, and Authentik was only reachable via the NetBird mesh — which was down *because* mgmt couldn't start. A classic bootstrap deadlock: **mgmt needs OIDC → OIDC needs the mesh → the mesh needs mgmt.** -Resolved by moving Authentik off [LXC 124](../containers/124-authentik.md) onto the VPS itself, so `auth.hubris.network` resolves to a container co-located with netbird-mgmt — no mesh dependency. A `depends_on: condition: service_healthy` on the mgmt service makes the deadlock structurally impossible to recur. +Resolved by moving Authentik off [LXC 124](../../knowledge/wiki/containers/106-auth-outpost.md) onto the VPS itself, so `auth.hubris.network` resolves to a container co-located with netbird-mgmt — no mesh dependency. A `depends_on: condition: service_healthy` on the mgmt service makes the deadlock structurally impossible to recur. The full Authentik Postgres DB (all users, apps, passwords, groups) was migrated, so every gated app keeps working with no per-app reconfiguration. @@ -43,7 +43,7 @@ The real reason the browser kept hitting the *old* Authentik even after the VPS 1. **Redirect URI error.** The restored DB had redirect URIs in `REGEX` matching mode; in Authentik 2026.5.x they failed to match. Fixed by switching to `STRICT` exact matching (Django ORM, `RedirectURIMatchingMode.STRICT`). Set all four: `http://localhost:53000/` (CLI), `https://netbird.hubris.network/{peers,nb-auth,nb-silent-auth}`. 2. **Only the password field showed (no username).** NetBird passes `login_hint=` in the OAuth2 URL → Authentik pre-identifies and skips the identification stage. Expected behavior; not a bug. -3. **"Request has been denied. Unknown error."** Several overlapping causes: wrong password (reset via Django shell), reputation lockout after repeated failures (`Reputation.objects.all().delete()` — see [124-authentik](../containers/124-authentik.md)), and **broken default expression policies**. The restored DB carried 8 default policies authored in old `return`-style syntax incompatible with 2026.5.x's eval context; `ak apply_blueprints` re-applied the current defaults. +3. **"Request has been denied. Unknown error."** Several overlapping causes: wrong password (reset via Django shell), reputation lockout after repeated failures (`Reputation.objects.all().delete()` — see [124-authentik](../../knowledge/wiki/containers/106-auth-outpost.md)), and **broken default expression policies**. The restored DB carried 8 default policies authored in old `return`-style syntax incompatible with 2026.5.x's eval context; `ak apply_blueprints` re-applied the current defaults. 4. **Browser ran stale frontend JS.** Console showed `version 2026.2.2` while the backend was `2026.5.2` — because DNS still pointed at the old LXC (see DNS cutover above), not a cache issue. 5. **WebAuthn devices dead post-migration.** Passkeys are device/origin-bound and don't survive a host move. Deleted all WebAuthn devices via Django ORM; users must re-register MFA. @@ -51,7 +51,7 @@ The real reason the browser kept hitting the *old* Authentik even after the VPS | | Before | After | |---|---|---| -| Authentik host | [LXC 124](../containers/124-authentik.md) `192.168.8.180` | VPS `82.165.190.79`, `auth` Docker net `172.30.1.0/24` | +| Authentik host | [LXC 124](../../knowledge/wiki/containers/106-auth-outpost.md) `192.168.8.180` | VPS `82.165.190.79`, `auth` Docker net `172.30.1.0/24` | | Version | `2026.2.2` | `2026.5.2` | | `auth.hubris.network` (LAN) | dnsmasq → `192.168.8.175` (Caddy) | dnsmasq → `82.165.190.79` (VPS traefik) | | `auth.hubris.network` (public) | IONOS wildcard → VPS → mesh → LXC 124 | IONOS wildcard → VPS → local container | @@ -76,7 +76,7 @@ The real reason the browser kept hitting the *old* Authentik even after the VPS Forward-auth apps (Paperless, qBittorrent, Artifacto) initially still validated against LXC 124's *embedded* outpost (Caddy → `192.168.8.180:9000`) — split-brain against the frozen DB. Pointing Caddy at `https://auth.hubris.network` instead fails: VPS Traefik rewrites `X-Forwarded-Host` → outpost can't match the app → 404 (tested + reverted). -Fixed with a **dedicated LAN outpost** ([106 — auth-outpost](../containers/106-auth-outpost.md), `192.168.8.6`): `goauthentik/proxy` connects outbound to the VPS core and serves forward-auth locally; Caddy → outpost over the LAN, no Traefik, header preserved. Outpost `hubris-lan-outpost` carries the 3 proxy providers. Verified with 124-Authentik **stopped**. This was Phase 1 of the broader architecture migration (plan: VPS edge / hubris LAN core / Mac Mini redundancy). +Fixed with a **dedicated LAN outpost** ([106 — auth-outpost](../../knowledge/wiki/containers/106-auth-outpost.md), `192.168.8.6`): `goauthentik/proxy` connects outbound to the VPS core and serves forward-auth locally; Caddy → outpost over the LAN, no Traefik, header preserved. Outpost `hubris-lan-outpost` carries the 3 proxy providers. Verified with 124-Authentik **stopped**. This was Phase 1 of the broader architecture migration (plan: VPS edge / hubris LAN core / Mac Mini redundancy). ### 2026-06-05 — identification stage skip: broken "Trust me" reputation policy @@ -99,11 +99,11 @@ The policy was orphaned (no matched type data or had incompatible evaluation). R - **muli-laptop** needs `netbird down && netbird up` + `resolvectl flush-caches`. - **VPS port 22** opened for this repair; close once remote access is otherwise stable. - **Decommission LXC 124 Authentik** after a ~2-week dual-run validation. dnsmasq stays on 124 regardless (separate service). -- **Reconcile [124-authentik](../containers/124-authentik.md) provider notes** — docs describe a `Public`/PKCE provider; the migrated DB carries the `Confidential` `netbird-dashboard` client. Verify which is live and correct the page. +- **Reconcile [124-authentik](../../knowledge/wiki/containers/106-auth-outpost.md) provider notes** — docs describe a `Public`/PKCE provider; the migrated DB carries the `Confidential` `netbird-dashboard` client. Verify which is live and correct the page. - **sops-encrypt** the VPS secrets (`/opt/authentik.env`) into the `secrets/` tree. ## Related -- [124 — authentik](../containers/124-authentik.md) -- [DNS split-horizon](../infrastructure/dns.md) -- [Public ingress (VPS traefik)](../infrastructure/ingress.md) -- [Mesh migration](../infrastructure/mesh.md) +- [124 — authentik](../../knowledge/wiki/containers/106-auth-outpost.md) +- [DNS split-horizon](../../knowledge/wiki/infrastructure/dns.md) +- [Public ingress (VPS traefik)](../../knowledge/wiki/infrastructure/ingress.md) +- [Mesh migration](../../knowledge/wiki/infrastructure/mesh.md) diff --git a/knowledge/log.md b/knowledge/log.md index 508908b..c52f3d4 100644 --- a/knowledge/log.md +++ b/knowledge/log.md @@ -6,3 +6,4 @@ each page's `## Changelog` and the Oikos change ledger, not here. ## [2026-07-06] restructure | moved node/infrastructure narratives under knowledge/wiki/; references under knowledge/sources/; repointed inventory doc_page fields and gen-topology.py output. ## [2026-07-06] lint | banned-vocabulary scan of knowledge/ clean; added .agents/skills/docs-lint and knowledge/wiki/hosts/index.md. +## [2026-07-06] lint | fixed 126 pre-existing broken links (124-authentik.md rename, investigations/plans moved to archive/done, archive/ sibling depth, destroyed-node delinks); 2 remaining are an intentional cross-repo reference. diff --git a/knowledge/wiki/containers/103-paperless.md b/knowledge/wiki/containers/103-paperless.md index 20b6702..4f7c127 100644 --- a/knowledge/wiki/containers/103-paperless.md +++ b/knowledge/wiki/containers/103-paperless.md @@ -18,7 +18,7 @@ Paperless-ngx for document management. Ingests scans / PDFs from `/mnt/library/d | paperless-task-queue, paperless-scheduler, paperless-consumer | — | systemd workers | ## Auth -Behind [Authentik forward-auth](124-authentik.md). API path `/api/*` bypasses forward-auth (mobile clients can't follow the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates matching users on first SSO login. +Behind [Authentik forward-auth](106-auth-outpost.md). API path `/api/*` bypasses forward-auth (mobile clients can't follow the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true` and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates matching users on first SSO login. ## Storage - Documents at `/mnt/library/documents` (owner `www-data:www-data`, mode 750 — *not* on the `media` group, by design). @@ -27,7 +27,7 @@ Behind [Authentik forward-auth](124-authentik.md). API path `/api/*` bypasses fo - ~~Disk usage was 86.9% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15. ## Related -- [Authentik](124-authentik.md) +- [Authentik](106-auth-outpost.md) - [Caddy](121-caddy.md) - [DNS](../infrastructure/dns.md) - [Monitoring](../infrastructure/monitoring.md) diff --git a/knowledge/wiki/containers/104-gitea.md b/knowledge/wiki/containers/104-gitea.md index 95e40c9..2415d8e 100644 --- a/knowledge/wiki/containers/104-gitea.md +++ b/knowledge/wiki/containers/104-gitea.md @@ -55,7 +55,7 @@ Initial documentation. Added `192.168.8.205`. See [Artifacto auto-deploy on apps (105)](105-apps.md). ### 2026-04-21 — `/etc/hosts` override for `auth.hubris.network` added -For OIDC integration with [authentik (124)](124-authentik.md). Outside the PVE markers, with a hubris-hosts-override.service for idempotency. +For OIDC integration with [authentik (124)](106-auth-outpost.md). Outside the PVE markers, with a hubris-hosts-override.service for idempotency. ### 2026-04-20 — gitea customizations + auto-deploy pipeline shipped `dtoro/gitea-customizations` repo created; webhook receiver at loopback `:9797` validates HMAC and runs `deploy.sh`. CAD and PlantUML loaders live in `footer.tmpl`. diff --git a/knowledge/wiki/containers/105-apps.md b/knowledge/wiki/containers/105-apps.md index 3cb2130..f67ed94 100644 --- a/knowledge/wiki/containers/105-apps.md +++ b/knowledge/wiki/containers/105-apps.md @@ -100,7 +100,7 @@ Native OIDC via `[oauth.generic]` in `config/config.ini`. `host = https://auth.h ## Related - [Gitea (104)](104-gitea.md) — uses the PlantUML server - [Caddy (121)](121-caddy.md) -- [Authentik (124)](124-authentik.md) +- [Authentik (124)](106-auth-outpost.md) - [DNS](../infrastructure/dns.md) - [Auto-deploy](../infrastructure/auto-deploy.md) - [Public ingress (Artifacto + blog)](../infrastructure/ingress.md) @@ -116,7 +116,7 @@ Two new services from the [homelab-context distribution plan](../infrastructure/ `secrets-issuance.service` on `:9820` (per-client age-key provisioning). Caddy fronts both with Let's Encrypt; new vhosts on [caddy](121-caddy.md), split-horizon DNS entries on -[authentik (124)](124-authentik.md). Gitea webhook ids 10 + 11 wire +[authentik (124)](106-auth-outpost.md). Gitea webhook ids 10 + 11 wire auto-deploy. LXC is itself an enrolled context client (`/opt/homelab-context/`). diff --git a/knowledge/wiki/containers/106-auth-outpost.md b/knowledge/wiki/containers/106-auth-outpost.md index 19c16b1..27ae131 100644 --- a/knowledge/wiki/containers/106-auth-outpost.md +++ b/knowledge/wiki/containers/106-auth-outpost.md @@ -1,6 +1,6 @@ # 106 — `auth-outpost` -Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that connects outbound to the [VPS Authentik core](../../../investigations/2026-05-31-authentik-vps-migration.md) and serves forward-auth locally, so [Caddy (121)](121-caddy.md) never hairpins auth through VPS Traefik. +Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that connects outbound to the [VPS Authentik core](../../../investigations/archive/2026-05-31-authentik-vps-migration.md) and serves forward-auth locally, so [Caddy (121)](121-caddy.md) never hairpins auth through VPS Traefik. ## At a glance - **Hostname:** `auth-outpost` @@ -8,11 +8,11 @@ Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that co - **Privilege:** privileged (Docker-in-LXC, `features: nesting=1`) - **Resources:** 1 core / 512 MiB / 4 GiB rootfs - **Mounts:** none -- **Created:** 2026-06-01, Debian 13, replacing the embedded outpost on [124](124-authentik.md) +- **Created:** 2026-06-01, Debian 13, replacing the embedded outpost on [124](106-auth-outpost.md) ## Role -Runs one container — `ghcr.io/goauthentik/proxy` — that opens an outbound websocket to `https://auth.hubris.network` (the VPS core), pulls its proxy-provider config, and answers Caddy's `forward_auth` subrequests on `192.168.8.6:9000` (LAN-only bind). Because the call path is **Caddy → outpost (LAN)**, with no Traefik in between, `X-Forwarded-Host` is preserved — the failure that 404s when Caddy is pointed at `https://auth.hubris.network` directly (Traefik rewrites the header). See the [migration investigation](../../../investigations/2026-05-31-authentik-vps-migration.md). +Runs one container — `ghcr.io/goauthentik/proxy` — that opens an outbound websocket to `https://auth.hubris.network` (the VPS core), pulls its proxy-provider config, and answers Caddy's `forward_auth` subrequests on `192.168.8.6:9000` (LAN-only bind). Because the call path is **Caddy → outpost (LAN)**, with no Traefik in between, `X-Forwarded-Host` is preserved — the failure that 404s when Caddy is pointed at `https://auth.hubris.network` directly (Traefik rewrites the header). See the [migration investigation](../../../investigations/archive/2026-05-31-authentik-vps-migration.md). ## Service / port map | Service | Listen | Notes | @@ -42,10 +42,10 @@ Fix: the LAN outpost gets its **own** domain. **Lesson:** when the IdP core and the forward-auth outpost live on different hosts, the outpost needs a dedicated domain distinct from the core's — and proxy-provider `redirect_uris` must be regenerated, not just `external_host`. ## Related -- [124 — authentik](124-authentik.md) — old embedded-outpost host (now DNS-only) +- [124 — authentik](106-auth-outpost.md) — old embedded-outpost host (now DNS-only) - [Caddy (121)](121-caddy.md) — forward-auth consumer - [Ingress (VPS traefik)](../infrastructure/ingress.md) -- [Authentik VPS migration](../../../investigations/2026-05-31-authentik-vps-migration.md) +- [Authentik VPS migration](../../../investigations/archive/2026-05-31-authentik-vps-migration.md) ## Changelog @@ -53,4 +53,4 @@ Fix: the LAN outpost gets its **own** domain. VPS Authentik core `user_login` stage updated: `session_duration` changed from `seconds=0` (session cookie, cleared on browser close) to `days=30` (persistent 30-day cookie). Also set `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30` in `/opt/authentik.env` on the VPS. See [investigation](../../../investigations/2026-06-06-authentik-session-lifetime.md). ### 2026-06-01 — created; forward-auth cut over from LXC 124 -New dedicated LXC for the LAN forward-auth outpost (Phase 1 of the [architecture migration](../../../investigations/2026-05-31-authentik-vps-migration.md)). Deployed `goauthentik/proxy:2026.5.2` pointed at the VPS core; repointed Caddy `(authentik)` from `192.168.8.180:9000` → `192.168.8.6:9000`. Verified Paperless/qBittorrent/Artifacto return the SSO redirect with **124-Authentik stopped**, confirming the frozen instance is out of the path. dnsmasq stays on 124 until [DNS is relocated](124-authentik.md). +New dedicated LXC for the LAN forward-auth outpost (Phase 1 of the [architecture migration](../../../investigations/archive/2026-05-31-authentik-vps-migration.md)). Deployed `goauthentik/proxy:2026.5.2` pointed at the VPS core; repointed Caddy `(authentik)` from `192.168.8.180:9000` → `192.168.8.6:9000`. Verified Paperless/qBittorrent/Artifacto return the SSO redirect with **124-Authentik stopped**, confirming the frozen instance is out of the path. dnsmasq stays on 124 until [DNS is relocated](106-auth-outpost.md). diff --git a/knowledge/wiki/containers/107-dns.md b/knowledge/wiki/containers/107-dns.md index 74a99a6..02d09a3 100644 --- a/knowledge/wiki/containers/107-dns.md +++ b/knowledge/wiki/containers/107-dns.md @@ -1,6 +1,6 @@ # 107 — `dns` -Homelab DNS server (Technitium). Replaces the dnsmasq that lived on [124 — authentik](124-authentik.md); single-purpose, one job. +Homelab DNS server (Technitium). Replaces the dnsmasq that lived on [124 — authentik](106-auth-outpost.md); single-purpose, one job. ## At a glance - **Hostname:** `dns` @@ -42,7 +42,7 @@ Technitium also runs a DHCP server for the homelab subnet (enabled 2026-06-02): Replaces the DHCP that was previously served by the Slate AX router. Static-IP LXCs (`.101–.239`) are excluded from the pool. Pool narrowed from `.100–.240` to `.241–.254` on 2026-06-03 to eliminate IP conflict risk. ## Related -- [124 — authentik](124-authentik.md) — retired host of the old dnsmasq +- [124 — authentik](106-auth-outpost.md) — retired host of the old dnsmasq - [DNS split-horizon](../infrastructure/dns.md) - [Mesh](../infrastructure/mesh.md) @@ -55,7 +55,7 @@ Added for [trmnl (128)](128-trmnl.md) (LAN path via [Caddy (121)](121-caddy.md)) Although the 2026-06-03 changelog claimed "cron */10", **no crontab was actually configured** on the LXC. The sync was running only via ad-hoc manual invocations during incident debugging. Fixed by adding `/etc/cron.d/dns-sync`. ### 2026-06-03 — DHCP pool narrowed to `.241–.254` -Previous pool `.100–.240` overlapped with all static LXCs/VMs (`.101–.239`). Shrunk via API (`/api/dhcp/scopes/set`). 11 stale DHCP leases in `.101–.110` remain until natural expiry (2026-06-04). See [plan](../../../plans/2026-06-03-dhcp-pool-exclude-static-ips.md). +Previous pool `.100–.240` overlapped with all static LXCs/VMs (`.101–.239`). Shrunk via API (`/api/dhcp/scopes/set`). 11 stale DHCP leases in `.101–.110` remain until natural expiry (2026-06-04). See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md). ### 2026-06-03 — dns-sync added (Technitium → NetBird managed zone) This Technitium became the single DNS authoring source; `/opt/dns-sync/sync.py` (cron */10) reconciles named A-records into the NetBird managed zone via the API. Fixed previously-broken mesh names (`sso`, `nfs-export`, `mcp`, `secrets`) by adding them to the managed zone; reaped obsolete `files`/`photos-new`. See [dns.md](../infrastructure/dns.md). @@ -64,4 +64,4 @@ This Technitium became the single DNS authoring source; `/opt/dns-sync/sync.py` Enabled Technitium's built-in DHCP server for `192.168.8.0/24` (scope `homelab`, range `.100–.240`, gateway `192.168.8.1`, DNS self). Previously the Slate AX sub-router served DHCP for the homelab subnet. With the Slate AX retired and Proxmox now the subnet router, Technitium takes over DHCP. Configured via the Technitium API (`/api/dhcp/scopes/set`). DHCP LXCs kept their Slate AX leases until expiry, then renewed from Technitium. ### 2026-06-01 — created; replaced dnsmasq on 124 -Stood up Technitium at `192.168.8.2`, imported the split-horizon zone (specific A + wildcard + MX/SPF/CAA), made it the primary nameserver in the NetBird `home-lab-dns` group. Verified all names resolve with dnsmasq/124 stopped; [LXC 124 retired](124-authentik.md). +Stood up Technitium at `192.168.8.2`, imported the split-horizon zone (specific A + wildcard + MX/SPF/CAA), made it the primary nameserver in the NetBird `home-lab-dns` group. Verified all names resolve with dnsmasq/124 stopped; [LXC 124 retired](106-auth-outpost.md). diff --git a/knowledge/wiki/containers/114-nextcloud.md b/knowledge/wiki/containers/114-nextcloud.md index 3792fc3..2150431 100644 --- a/knowledge/wiki/containers/114-nextcloud.md +++ b/knowledge/wiki/containers/114-nextcloud.md @@ -11,7 +11,7 @@ Personal cloud / file collaboration. Source-of-truth for the photo libraries sur - **Public hostname:** [`cloud.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md) ## Auth -Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin` via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider --mapping-uid=nc_uid` and `--scope="openid profile email -uid"`. See [Authentik](124-authentik.md#per-app-username-override-pattern-authentik) for the full pattern. +Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin` via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider --mapping-uid=nc_uid` and `--scope="openid profile email -uid"`. See [Authentik](106-auth-outpost.md#per-app-username-override-pattern-authentik) for the full pattern. Redirect URI: `/index.php/apps/user_oidc/code` (NOT `/apps/...` — pretty URLs aren't on). @@ -74,7 +74,7 @@ Apply with `systemctl restart mariadb` (not reload — `innodb_log_file_size` ne ## Related - [mulita (120)](120-mule-images.md) — reads NC user trees + writes back via WebDAV -- [Authentik (124)](124-authentik.md) +- [Authentik (124)](106-auth-outpost.md) - [Caddy (121)](121-caddy.md) - [DNS](../infrastructure/dns.md) - [Mesh migration (DNS overrides explained)](../infrastructure/mesh.md) diff --git a/knowledge/wiki/containers/118-elementsynapse.md b/knowledge/wiki/containers/118-elementsynapse.md index e552a08..17e4334 100644 --- a/knowledge/wiki/containers/118-elementsynapse.md +++ b/knowledge/wiki/containers/118-elementsynapse.md @@ -37,7 +37,7 @@ All five bridges run as plain `docker compose` stacks under `/root/mautrix- **This LXC was destroyed on 2026-06-04.** Replaced by Hermes Agent on mac-mini. > Monitoring migrated to `homelab-hardware-health` skill + 15-min Hermes cronjob. > Repos `dtoro/claudio-bot` and `dtoro/claudio-monitor` archived (read-only) on Gitea. -> See [deprecation plan](../plans/2026-06-04_130000-deprecate-claudio-bot.md) for full details. +> See [deprecation plan](../../../../plans/done/2026-06-04_130000-deprecate-claudio-bot.md) for full details. Matrix-resident control plane. Bot account `@claudio:avispero` joined to a private room; accepts slash commands and natural language; relays infra notifications. @@ -17,7 +17,7 @@ Matrix-resident control plane. Bot account `@claudio:avispero` joined to a priva ## Stack -Repo `dtoro/claudio-bot`, checkout at `/opt/claudio-bot`, systemd unit `claudio-bot.service`. Connects to Matrix at `http://192.168.8.239:8008` (direct LAN to [synapse (118)](118-elementsynapse.md), avoids hairpin-NAT TLS issue on `matrix.hubris.network`). +Repo `dtoro/claudio-bot`, checkout at `/opt/claudio-bot`, systemd unit `claudio-bot.service`. Connects to Matrix at `http://192.168.8.239:8008` (direct LAN to [synapse (118)](../118-elementsynapse.md), avoids hairpin-NAT TLS issue on `matrix.hubris.network`). Room: `!dEUVJArVKPorxHHJZK:avispero` (invite-only, `@dtoro:avispero` allowed). @@ -43,8 +43,8 @@ Currently set to `lmstudio` → `google/gemma-4-e4b` on the Mac mini at `192.168 ## IPC `http://192.168.8.230:9090/{notify,propose,status}`. Header `X-Bot-Token` must match `/etc/claudio-bot/ipc.token`. Used by: -- [claudio-monitor on hubris](../infrastructure/monitoring.md) for edge-triggered alerts (token in `/etc/claudio-monitor/bot.token`) -- The (currently disabled) [restic backup wrapper](../infrastructure/backups.md) (token in `/etc/restic/bot.token`) +- [claudio-monitor on hubris](../../infrastructure/monitoring.md) for edge-triggered alerts (token in `/etc/claudio-monitor/bot.token`) +- The (currently disabled) [restic backup wrapper](../../infrastructure/backups.md) (token in `/etc/restic/bot.token`) > Token files at the source side **must hold the same value as `/etc/claudio-bot/ipc.token`** — rotate together. @@ -61,13 +61,13 @@ Active plugins: Push to `dtoro/claudio-bot` → gitea webhook → `http://192.168.8.230:9797/deploy` → pull + `pip install` + restart. Same shape as caddy-conf. -`app.ini` `ALLOWED_HOST_LIST` on [gitea](104-gitea.md) includes `192.168.8.230`. +`app.ini` `ALLOWED_HOST_LIST` on [gitea](../104-gitea.md) includes `192.168.8.230`. ## Related -- [elementsynapse (118)](118-elementsynapse.md) -- [Monitoring (claudio-monitor)](../infrastructure/monitoring.md) -- [Backups (disabled)](../infrastructure/backups.md) -- [Auto-deploy](../infrastructure/auto-deploy.md) +- [elementsynapse (118)](../118-elementsynapse.md) +- [Monitoring (claudio-monitor)](../../infrastructure/monitoring.md) +- [Backups (disabled)](../../infrastructure/backups.md) +- [Auto-deploy](../../infrastructure/auto-deploy.md) ## Changelog @@ -84,7 +84,7 @@ Initial documentation. `backend: lmstudio` → `google/gemma-4-e4b` on the Mac mini. Anthropic key still present so the swap is reversible by flipping the config field. ### 2026-04-21 — `monitor` plugin added -Receives events from [claudio-monitor](../infrastructure/monitoring.md). Slash commands + tools registered. See `dtoro/claudio-bot` commit `e56da25`. +Receives events from [claudio-monitor](../../infrastructure/monitoring.md). Slash commands + tools registered. See `dtoro/claudio-bot` commit `e56da25`. ### 2026-04-20 — claudio-bot deployed LXC 123 provisioned. Repo, systemd unit, Matrix wiring, `system` + `backup` plugins, IPC server. diff --git a/knowledge/wiki/containers/archive/127-mule-photos-new.md b/knowledge/wiki/containers/archive/127-mule-photos-new.md index bbf1574..0966251 100644 --- a/knowledge/wiki/containers/archive/127-mule-photos-new.md +++ b/knowledge/wiki/containers/archive/127-mule-photos-new.md @@ -1,7 +1,7 @@ # 127 — `mule-photos-new` Side-by-side **PhotoPrism M0 test** of the `dtoro/mule-image` `new` branch -at `photos-new.hubris.network`. Production [LXC 120](120-mule-images.md) keeps +at `photos-new.hubris.network`. Production [LXC 120](../120-mule-images.md) keeps running on the legacy stack at `photos.hubris.network` until M5 cutover. ## At a glance @@ -11,7 +11,7 @@ running on the legacy stack at `photos.hubris.network` until M5 cutover. - **Resources:** 6 cores / 8 GiB RAM / 40 GiB rootfs / 1 GiB swap - **Features:** `nesting=1,fuse=1,keyctl=1` - **Mounts:** *(none — see scratch copy below)* -- **Public hostname:** [`photos-new.hubris.network`](../infrastructure/dns.md) → [caddy (121)](121-caddy.md) → split (PhotoPrism `:2342`, sidecar `:8000`, Vite `:5173`) +- **Public hostname:** [`photos-new.hubris.network`](../../infrastructure/dns.md) → [caddy (121)](../121-caddy.md) → split (PhotoPrism `:2342`, sidecar `:8000`, Vite `:5173`) ## Stack (`/opt/mule-image`) @@ -65,7 +65,7 @@ unprivileged LXCs can't see through. ## Auth — Authentik OIDC -PhotoPrism's "Sign in with OIDC" button delegates to [Authentik (124)](124-authentik.md). +PhotoPrism's "Sign in with OIDC" button delegates to [Authentik (124)](../106-auth-outpost.md). - **Provider/Application slug:** `mule-photos-new` - **Issuer:** `https://auth.hubris.network/application/o/mule-photos-new/` @@ -111,7 +111,7 @@ Mirrors the LXC 120 pattern. Push to the `new` branch on [git.hubris.network/dtoro/mule-image](http://git.hubris.network/dtoro/mule-image) → webhook fires → rebuild. The legacy LXC 120 watches `main` and is unaffected. **Gitea gotcha:** the receiver IP must be in `[webhook] ALLOWED_HOST_LIST` -in `/etc/gitea/app.ini` on [LXC 104](104-gitea.md). LXC 127's +in `/etc/gitea/app.ini` on [LXC 104](../104-gitea.md). LXC 127's `192.168.8.181` was missing on first bring-up; every push delivered status 0 with the message `webhook can only call allowed HTTP servers`. Adding the IP and `systemctl restart gitea` is enough — same list is @@ -145,10 +145,10 @@ curl -sk --resolve photos-new.hubris.network:443:192.168.8.175 \ ``` > **Decommissioned 2026-05-22.** The PhotoPrism + sidecar + SvelteKit stack -> validated here was promoted into production on [LXC 120](120-mule-images.md) +> validated here was promoted into production on [LXC 120](../120-mule-images.md) > via the `Mulimage 2.0` merge (`dtoro/mule-image` commit `70dc1b6`). This > page is retained for archaeology; everything below is historic. See the -> 2026-05-22 entry in [120-mule-images.md](120-mule-images.md#changelog) for +> 2026-05-22 entry in [120-mule-images.md](../120-mule-images.md#changelog) for > the cutover detail. ## Changelog diff --git a/knowledge/wiki/containers/index.md b/knowledge/wiki/containers/index.md index 9119132..9d034ac 100644 --- a/knowledge/wiki/containers/index.md +++ b/knowledge/wiki/containers/index.md @@ -15,7 +15,7 @@ Most containers live on [`hubris`](../hosts/hubris.md). Some have been | 120 | [mule-images](120-mule-images.md) | hubris | 192.168.8.136 | priv | 6 | 12 GiB | 60 GiB | `/mnt/library` + `/dev/dri` (iGPU) | `photos.hubris.network` | running | | 121 | [caddy](121-caddy.md) | hubris | 192.168.8.175 | unpriv | 1 | 512 MiB | 6 GiB | — | (terminates all `*.hubris.network`) | running | | 122 | [arriman](122-arriman.md) | **strong** | 192.168.8.245 | priv | 4 | 8 GiB | 24 GiB | `/mnt/media_local` (via mp0) | `jellyseerr` / `qbit` / `sab` | running | -| 124 | [authentik](124-authentik.md) | hubris | 192.168.8.180 | priv | 2 | 4 GiB | 20 GiB | — | `auth.hubris.network` | running | +| 124 | [authentik](106-auth-outpost.md) | hubris | 192.168.8.180 | priv | 2 | 4 GiB | 20 GiB | — | `auth.hubris.network` | running | | 128 | [trmnl](128-trmnl.md) | hubris | 192.168.8.211 | unpriv | 1 | 768 MiB | 8 GiB | — | `trmnl.hubris.network` | running | | 129 | [house](129-house.md) | **strong** | 192.168.8.244 | unpriv | 2 | 3 GiB | 8 GiB | — | `house.hubris.network` | running | | 130 | [grimmory](130-grimmory.md) | **strong** | 192.168.8.247 | priv | 1 | 2 GiB | 16 GiB | `/mnt/media_local` (via mp0) | `books.hubris.network` | running | @@ -32,7 +32,7 @@ Most containers live on [`hubris`](../hosts/hubris.md). Some have been | 106 | flaresolverr | ~2026-04-28 | Folded into the arriman docker compose | | 116 | heaper | 2026-05-14 | Decommissioned by user; data subtree at `/mnt/library/heaper` (224 MiB) retained | | 126 | plato | 2026-06-28 | Notes/discovery workspace decommissioned; data at `/mnt/library/documents/plato` retained for archaeology | -| 123 | claudio-bot (destroyed — see [archive](archive/123-claudio-bot.md)) | 2026-06-04 | Replaced by Hermes Agent on mac-mini; monitoring migrated to `homelab-health-watchdog` cron. See [deprecation plan](../../../plans/2026-06-04_130000-deprecate-claudio-bot.md) | +| 123 | claudio-bot (destroyed — see [archive](archive/123-claudio-bot.md)) | 2026-06-04 | Replaced by Hermes Agent on mac-mini; monitoring migrated to `homelab-health-watchdog` cron. See [deprecation plan](../../../plans/done/2026-06-04_130000-deprecate-claudio-bot.md) | | 109 | syncthing | 2026-05-14 | Decommissioned by user; `/mnt/library/syncthing` was already empty | | 125 | seafile | 2026-05-13 | Seafile Pro evaluation, user disliked the product; teardown also removed `files.hubris.network` from caddy + dnsmasq | | 107 | marimo | between 2026-04-21 and 2026-04-28 | Decommissioned | @@ -45,7 +45,7 @@ Most containers live on [`hubris`](../hosts/hubris.md). Some have been ## Conventions -- All net0 are `bridge=vmbr0`, `ip=dhcp` except [124 (authentik)](124-authentik.md) which is statically `192.168.8.180/24`. Containers on [strong](../hosts/strong.md) use `bridge=vmbr1` with static IPs in the `192.168.8.240/28` range. +- All net0 are `bridge=vmbr0`, `ip=dhcp` except [124 (authentik)](106-auth-outpost.md) which is statically `192.168.8.180/24`. Containers on [strong](../hosts/strong.md) use `bridge=vmbr1` with static IPs in the `192.168.8.240/28` range. - `onboot=1` on every container — the host brings them up after `pve-guests.service`. - Bind mounts are declared as `mp0: /mnt/library,mp=/mnt/library` on hubris, or `mp0: /mnt/media_local,mp=/mnt/library` on strong. - Most containers are privileged. Unprivileged ones require an idmap block in their conf to participate in the [media GID 10000](../infrastructure/media-permissions.md) standard. diff --git a/knowledge/wiki/hosts/hubris.md b/knowledge/wiki/hosts/hubris.md index d5155b2..48d91ba 100644 --- a/knowledge/wiki/hosts/hubris.md +++ b/knowledge/wiki/hosts/hubris.md @@ -8,7 +8,7 @@ workloads still live here. As of 2026-07-01, hubris is node 1 of the 2-node ## At a glance - **Role:** Proxmox VE 9.1.2 hypervisor (kernel `6.14.11-4-pve`) - **Hardware:** GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS (Phoenix APU), 12 vCPU / ~28 GiB RAM, 2× Samsung 990 EVO Plus NVMe (one SSD primary, one for `library` LVM). 2× Realtek RTL8125 NICs (`r8169`). -- **BIOS:** 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../../../investigations/2026-04-21-hubris-crash-loop.md). +- **BIOS:** 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). - **Uplink:** `vmbr1` (slave: `eno1`) → SODOLA switch → Fritz!Box 7590. DHCP-reserved `192.168.178.10/24`, gateway `192.168.178.1`. - **Homelab bridge:** `vmbr0` — portless internal bridge, `192.168.8.77/24` + `192.168.8.1/24` alias (LXC default gateway). All 16 LXCs and the HAOS VM are on `vmbr0`. Proxmox routes between `vmbr0` and `vmbr1`; Fritz!Box has a static route `192.168.8.0/24 → 192.168.178.10`. - **WiFi:** disabled 2026-06-02 — `wlp3s0` removed from `/etc/network/interfaces`, wpa config deleted. Was used as a failover to the now-retired Slate AX AP. @@ -114,7 +114,7 @@ OpenSSH on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:220 - [Monitoring](../infrastructure/monitoring.md) - [Backups (disabled)](../infrastructure/backups.md) - [Operations cheatsheet](../../../operations/commands.md) -- [Investigation: 2026-04-21 crash loop](../../../investigations/2026-04-21-hubris-crash-loop.md) +- [Investigation: 2026-04-21 crash loop](../../../investigations/archive/2026-04-21-hubris-crash-loop.md) - [strong — Proxmox host](strong.md) ## Changelog @@ -123,7 +123,7 @@ OpenSSH on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:220 User reformatted `strong` (formerly a Linux dev workstation, `192.168.178.181`) to Proxmox VE 9.2.3. Cluster/OS hostname on that box is `strong` (left as-is from install). Bootstrapped root SSH on strong from a one-time console password (installed hubris's existing trusted key set: `root@hubris`, `d.toro.v@pm.me`), then generated a keypair on strong and pre-authorized it here (`root@strong`) so `pvecm add 192.168.8.77 --use_ssh 1` (run from strong) could join without an interactive password prompt. No cabling/routing changes needed — strong reaches hubris's corosync address (`192.168.8.77`) via the existing Fritz!Box static route. Cluster now 2 nodes, quorate, **no QDevice** (explicit choice — see [Cluster](#cluster) above for the quorum tradeoff this implies). strong hosts no guests yet; this is Phase 1 of the [library-SSD migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md), nothing further from that plan has been executed. ### 2026-06-02 — Slate AX retired; SODOLA switch added; network restructured -Replaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed. Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md) and [migration plan](../../../plans/2026-06-01-slate-ax-to-sodola-migration.md). +Replaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed. Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md) and [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md). ### 2026-05-14 — LXC 109 (syncthing) decommissioned User destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy 109 --purge` cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing` was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up. Entry moved to the "recently destroyed" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology); references stripped from [README](../../../README.md), [media-permissions](../infrastructure/media-permissions.md), [vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md). @@ -138,7 +138,7 @@ User destroyed the heaper LXC. No `116.conf.bak` left behind in `/etc/pve/lxc/`. `/etc/sysctl.d/99-bbr.conf` switches `net.ipv4.tcp_congestion_control` from `cubic` to `bbr` and `net.core.default_qdisc` from `fq_codel` to `fq`. Also bumps `rmem_max`/`wmem_max` to 64 MiB and widens `tcp_rmem`/`tcp_wmem`. `tcp_bbr` module pinned at boot via `/etc/modules-load.d/bbr.conf`. Triggered by Nextcloud client downloads from a WiFi laptop pulling ~2 MB/s despite a 152 Mbps link — server-side baseline through Caddy with BBR is ~400 MB/s single-stream loopback, so any client-perceived single-stream improvement is pure congestion-control win. Touches every LXC's outbound TCP since they all share this kernel. ### 2026-04-29 — relocated to better-ventilated spot -User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C, ACPI zone 47–49 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 60–61 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../../../investigations/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation). +User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C, ACPI zone 47–49 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 60–61 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../../../investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation). ### 2026-04-28 — Phase 1 WiFi failover Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failover, metric 200) on the GL-AXT1800-714-5G AP. Installed `wpasupplicant`+`iw`; added `wlp3s0` stanza to `/etc/network/interfaces` with `wpa-conf`; ARP isolation sysctls in `post-up`. Built `wan-failover.service` to remove the vmbr0 default route on `eno1` carrier loss, since the bridge's carrier doesn't follow `eno1` (the LXC veths keep it `1`). LXC/VM guests are still LAN-only — Phase 2 will migrate them. @@ -147,10 +147,10 @@ Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failov This wiki created. Live state at this date: 14 LXCs running (109 syncthing stopped), 1 VM, kernel `6.14.11-4-pve`, uptime 3 d 0 h post drive-removal A/B test. Compared to memory snapshot from a week ago, **destroyed**: LXC 100 (yunohost arr), 106 (flaresolverr), 107 (marimo), 110 (photoprism), 111 (karakeep), 112 (immich), 115 (reticulum). 100 + 106 destroyed per the planned 2026-04-21 \*arr migration retention; the others removed since. ### 2026-04-23 — SSD cooling + thermal pads installed -Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 60–61 °C, nvme1 38–40 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../../../investigations/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict). +Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 60–61 °C, nvme1 38–40 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../../../investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict). ### 2026-04-22 — drive removal A/B test -Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../../../investigations/2026-04-21-hubris-crash-loop.md). +Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). ### 2026-04-22 — `cpu-epp.service` ordering bug fixed Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `pve-guests.service`, so the hottest boot window (20+ guests starting on `performance`) preceded EPP application. Now `After=sysinit.target` + `Before=pve-guests.service`. @@ -159,4 +159,4 @@ Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind ` `60-crash-capture.conf`, softdog `soft_panic=1`, RuntimeWatchdog 15 s. `rasdaemon` installed and enabled. Pure silicon hangs still leave no trace; this catches everything else. ### 2026-04-21 — `cpu-epp.service` deployed -Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the [crash-loop incident](../../../investigations/2026-04-21-hubris-crash-loop.md). +Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the [crash-loop incident](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). diff --git a/knowledge/wiki/infrastructure/auto-deploy.md b/knowledge/wiki/infrastructure/auto-deploy.md index f998201..5774362 100644 --- a/knowledge/wiki/infrastructure/auto-deploy.md +++ b/knowledge/wiki/infrastructure/auto-deploy.md @@ -23,7 +23,7 @@ The app repo at `/opt/` is the working tree, but the deploy tooling (`web - ~~`192.168.8.230` (claudio-bot — destroyed 2026-06-04)~~ - `192.168.8.136` ([mule-images (120)](../containers/120-mule-images.md)) - `192.168.8.77` ([hubris host](../hosts/hubris.md) — backup-library) - - ~~`192.168.8.190` ([plato (126)](../containers/126-plato.md))~~ (destroyed 2026-06-28) + - ~~`192.168.8.190` ([plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology))~~ (destroyed 2026-06-28) - `192.168.8.211` ([trmnl (128)](../containers/128-trmnl.md) — terminalito) **Don't strip these when editing app.ini.** @@ -38,8 +38,8 @@ The app repo at `/opt/` is the working tree, but the deploy tooling (`web | `dtoro/gitea-customizations` | [gitea (104)](../containers/104-gitea.md) `/var/lib/gitea/custom/` | A | `http://127.0.0.1:9797/deploy` (loopback) | (orig) | `systemctl restart gitea` if templates changed | | `dtoro/mule-image` | [mule-images (120)](../containers/120-mule-images.md) `/opt/mule-image/` | B | `http://192.168.8.136:9797/deploy` | 6 | `docker compose up -d --build` | | `dtoro/Artifacto` | [apps (105)](../containers/105-apps.md) `/opt/artifacto/` | B | `http://192.168.8.205:9798/deploy` | 7 | `docker compose up -d --build` | -| ~~`dtoro/Plato`~~ | ~~[plato (126)](../containers/126-plato.md) `/opt/plato/app/`~~ (destroyed 2026-06-28) | ⊘ | `http://192.168.8.190:9799/deploy` (dead) | 8 (removed) | Repo archived — LXC destroyed | -| `dtoro/claudio-bot` | ~~[claudio-bot (123)](../containers/123-claudio-bot.md)~~ (destroyed 2026-06-04) | ⊘ | `http://192.168.8.230:9797/deploy` (dead) | (archived) | Repo archived — LXC destroyed | +| ~~`dtoro/Plato`~~ | ~~[plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology) `/opt/plato/app/`~~ (destroyed 2026-06-28) | ⊘ | `http://192.168.8.190:9799/deploy` (dead) | 8 (removed) | Repo archived — LXC destroyed | +| `dtoro/claudio-bot` | ~~[claudio-bot (123)](../containers/archive/123-claudio-bot.md)~~ (destroyed 2026-06-04) | ⊘ | `http://192.168.8.230:9797/deploy` (dead) | (archived) | Repo archived — LXC destroyed | | `dtoro/backup-library` | [hubris host](../hosts/hubris.md) `/opt/backup-library/` | A | `http://192.168.8.77:9798/deploy` | (orig) | runs `deploy.sh` (preserves admin-edited `/etc/restic/include-*.list`) | | `dtoro/Homelab-Docs` → homelab-mcp | [apps (105)](../containers/105-apps.md) `/opt/homelab-mcp/` | B | `http://192.168.8.205:9811/deploy` | 10 | reinstalls `homelab-mcp.service` + restart | | `dtoro/Homelab-Docs` → secrets-issuance | [apps (105)](../containers/105-apps.md) `/opt/secrets-issuance/` | B | `http://192.168.8.205:9821/deploy` | 11 | reinstalls `secrets-issuance.service` + restart | @@ -50,7 +50,7 @@ The app repo at `/opt/` is the working tree, but the deploy tooling (`web > Each owns its own clone on LXC 105. They don't conflict because each > deploy.sh only touches its own service unit + venv. -> **Not yet wired:** `dtoro/claudio-monitor` (push, then `/opt/claudio-monitor/scripts/deploy.sh` manually). The former authentik LXC (124) is destroyed — Authentik runs on the [VPS](../../../hosts/netbird-vps.md). DNS moved to [Technitium on dns (107)](../containers/107-dns.md). +> **Not yet wired:** `dtoro/claudio-monitor` (push, then `/opt/claudio-monitor/scripts/deploy.sh` manually). The former authentik LXC (124) is destroyed — Authentik runs on the [VPS](../../../hosts/netbird-vps.yaml). DNS moved to [Technitium on dns (107)](../containers/107-dns.md). ## When you change a tracked config @@ -131,7 +131,7 @@ Webhook id 12 on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy` on [ Webhook ids 10 + 11 on `dtoro/Homelab-Docs` (ports `9811` + `9821` on [apps (105)](../containers/105-apps.md)). Two webhooks on one repo — each owns its own clone (`/opt/homelab-mcp`, `/opt/secrets-issuance`) and only restarts its own service. See [homelab-context](homelab-context.md) for why both services live in one repo. ### 2026-05-13 — Plato pipeline added -Webhook id 8 on `dtoro/Plato` (port `9799` on [plato (126)](../containers/126-plato.md)). `app.ini` `ALLOWED_HOST_LIST` extended to include `192.168.8.190`. +Webhook id 8 on `dtoro/Plato` (port `9799` on [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology)). `app.ini` `ALLOWED_HOST_LIST` extended to include `192.168.8.190`. ### 2026-04-28 — wiki entry created Initial documentation. Six active pipelines. diff --git a/knowledge/wiki/infrastructure/backups.md b/knowledge/wiki/infrastructure/backups.md index 99d9bd9..2d93237 100644 --- a/knowledge/wiki/infrastructure/backups.md +++ b/knowledge/wiki/infrastructure/backups.md @@ -24,7 +24,7 @@ See [132-rclone](../containers/132-rclone.md) for the full design. ## Legacy — restic on external drive (DISABLED 2026-04-22) -Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabled 2026-04-22** as part of the [hubris crash-loop A/B test](../../../investigations/2026-04-21-hubris-crash-loop.md). +Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabled 2026-04-22** as part of the [hubris crash-loop A/B test](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). ## Status @@ -36,7 +36,7 @@ Chunked monthly restic backup of `/mnt/library`'s irreplaceable subset. **Disabl Fstab entry commented out. USB drive de-authorized and physically removed. `backup-library-deploy.service` left enabled (harmless webhook receiver). -**Reason:** the host hang recurred 2026-04-22 18:42 after 30h despite the `cpu-epp` fix, the UAS blacklist, and mount-on-demand. User wants to confirm host stability without the drive at all (was stable 33 days before the drive arrived). See [investigation](../../../investigations/2026-04-21-hubris-crash-loop.md). +**Reason:** the host hang recurred 2026-04-22 18:42 after 30h despite the `cpu-epp` fix, the UAS blacklist, and mount-on-demand. User wants to confirm host stability without the drive at all (was stable 33 days before the drive arrived). See [investigation](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). **To re-enable:** uncomment fstab line, `systemctl enable --now` the four timers, re-attach drive. @@ -111,7 +111,7 @@ Single drive. RECOVERY.md flags the 3-2-1 gap. Mitigations (second drive, cloud The `Silicon Motion Portable SSD` (vid:pid `090c:2320`) drops under sustained heavy writes through a hub chain. Bypass all hubs / use a rear motherboard USB 3 port if attaching it again. -After it was first attached on 2026-04-19, hubris crashed twice in 2.5 days (46h then 12h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic, OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1. The UAS blacklist + mount-on-demand mitigations didn't fully eliminate it (recurrence 2026-04-22), prompting drive removal as the cleaner test. See [investigation](../../../investigations/2026-04-21-hubris-crash-loop.md). +After it was first attached on 2026-04-19, hubris crashed twice in 2.5 days (46h then 12h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic, OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1. The UAS blacklist + mount-on-demand mitigations didn't fully eliminate it (recurrence 2026-04-22), prompting drive removal as the cleaner test. See [investigation](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). ## Thermal monitoring @@ -119,10 +119,10 @@ Moved out of this repo to `dtoro/claudio-monitor` on 2026-04-21 (commit `50dc213 ## Related - [Hubris host](../hosts/hubris.md) -- ~~[claudio-bot (123)](../containers/123-claudio-bot.md)~~ (destroyed 2026-06-04) +- ~~[claudio-bot (123)](../containers/archive/123-claudio-bot.md)~~ (destroyed 2026-06-04) - [Monitoring](monitoring.md) - [Auto-deploy](auto-deploy.md) -- [Investigation: 2026-04-21 crash loop](../../../investigations/2026-04-21-hubris-crash-loop.md) +- [Investigation: 2026-04-21 crash loop](../../../investigations/archive/2026-04-21-hubris-crash-loop.md) ## Changelog @@ -133,7 +133,7 @@ Off-host backup moved to a plain `rclone sync` mirror on the new [LXC 132 `rclon Initial documentation. Status remains DISABLED. ### 2026-04-22 — DISABLED -Drive removed as the A/B test in the [crash investigation](../../../investigations/2026-04-21-hubris-crash-loop.md). Timers disabled, fstab commented, drive de-authorized. +Drive removed as the A/B test in the [crash investigation](../../../investigations/archive/2026-04-21-hubris-crash-loop.md). Timers disabled, fstab commented, drive de-authorized. ### 2026-04-21 — UAS blacklist + mount-on-demand shipped; root-caused host hangs to drive Drive identified as the source of the hangs after hubris crashed twice in 2.5 days. UAS blacklist forces BOT; helper script toggles `/sys/bus/usb/.../authorized` so the drive is de-authorized when not backing up. Recovery drill (restore 188KB PDF + hash compare) had passed earlier. Bug fixed in `backup-library.sh`: `python3 -c '…' KEY=VAL` does NOT pass env vars — env-var prefix must precede the command. Caused false-failure even after successful backups. diff --git a/knowledge/wiki/infrastructure/dns.md b/knowledge/wiki/infrastructure/dns.md index c64eccd..af3bb52 100644 --- a/knowledge/wiki/infrastructure/dns.md +++ b/knowledge/wiki/infrastructure/dns.md @@ -7,7 +7,7 @@ There is **no wildcard on the LAN side**. Every subdomain needs an explicit entr ## Components - **Authoritative public DNS:** IONOS. `*.hubris.network → 82.165.190.79` (was `74.118.126.4` until 2026-04-22). -- **LAN authoritative for `hubris.network` records:** [Technitium DNS](https://technitium.com) on [dns (107)](../containers/107-dns.md) at `192.168.8.2:53`. Syncs A records to the NetBird managed DNS zone via cron (see [dns-sync.py](../../../scripts/dns-sync.py)). Formerly dnsmasq on [authentik (124)](../containers/124-authentik.md) (decommissioned 2026-06-04). +- **LAN authoritative for `hubris.network` records:** [Technitium DNS](https://technitium.com) on [dns (107)](../containers/107-dns.md) at `192.168.8.2:53`. Syncs A records to the NetBird managed DNS zone via cron (see [dns-sync.py](../../../scripts/dns-sync.py)). Formerly dnsmasq on [authentik (124)](../containers/106-auth-outpost.md) (decommissioned 2026-06-04). - **PVE host** (`192.168.8.77`): resolver is the local Netbird daemon at `100.122.38.109:53`, which forwards to the LAN/upstream and learns hubris.network answers via that path. `netbird status` says "Nameservers: 0/0 Available" — confirming netbird does NOT manage a hubris.network zone; it just caches whatever the system resolver returns. - **Some LXCs** keep router DNS (`192.168.8.1`) or Tailscale MagicDNS (`100.100.100.100`), both of which return the public IONOS A record. Those LXCs need either a `/etc/hosts` override or local dnsmasq — see [mesh migration](mesh.md) for which technique applies where. @@ -107,10 +107,10 @@ Also added a Caddy backend health check cron on hubris (`/etc/cron.d/caddy-backe All LXCs that Caddy reverse-proxies to by IP were on `ip=dhcp` and could float on reboot (arriman got a different lease mid-session and broke). Fixed via `pct set` + in-LXC `/etc/network/interfaces`. Affected: 101 jellyfin, 103 paperless, 104 gitea, 105 apps, 114 nextcloud, 118 elementsynapse, 120 mule-images, 121 caddy, 122 arriman. See [arriman changelog](../containers/122-arriman.md#changelog). ### 2026-06-01 — dnsmasq replaced by Technitium on [dns (107)](../containers/107-dns.md); LXC 124 retired -Split-horizon DNS moved off [124](../containers/124-authentik.md) to a dedicated **Technitium** LXC at **`192.168.8.2`** (zone: specific A overrides + wildcard→VPS + replicated MX/SPF/CAA). NetBird `home-lab-dns` nameserver group cut over to `192.168.8.2` (with `.180` as a now-dead fallback). dnsmasq stopped, all names verified via Technitium, **LXC 124 shut down**. **Caveat:** the [NetBird managed DNS zone](../containers/124-authentik.md) still answers most app names *directly* (bypassing the nameserver group) — three overlapping DNS sources remain; see the single-source-of-truth decision (Phase 4). **Action needed:** update router DHCP DNS from the dead `.180` → `192.168.8.2` for any plain-LAN (non-mesh) clients. +Split-horizon DNS moved off [124](../containers/106-auth-outpost.md) to a dedicated **Technitium** LXC at **`192.168.8.2`** (zone: specific A overrides + wildcard→VPS + replicated MX/SPF/CAA). NetBird `home-lab-dns` nameserver group cut over to `192.168.8.2` (with `.180` as a now-dead fallback). dnsmasq stopped, all names verified via Technitium, **LXC 124 shut down**. **Caveat:** the [NetBird managed DNS zone](../containers/106-auth-outpost.md) still answers most app names *directly* (bypassing the nameserver group) — three overlapping DNS sources remain; see the single-source-of-truth decision (Phase 4). **Action needed:** update router DHCP DNS from the dead `.180` → `192.168.8.2` for any plain-LAN (non-mesh) clients. ### 2026-05-31 — `auth.hubris.network` re-pointed to the VPS (`82.165.190.79`) -Authentik migrated off LXC 124 onto the VPS (see [investigation](../../../investigations/2026-05-31-authentik-vps-migration.md)). The dnsmasq entry changed from `192.168.8.175` (home Caddy) to `82.165.190.79` (VPS traefik). This is the first LAN entry that intentionally points at the VPS rather than Caddy — `auth` is now a genuinely public service served directly from the VPS. **Gotcha logged:** the NetBird per-client resolver (`100.122.255.254`) caches dnsmasq answers and does **not** clear on `netbird down/up`; clients needed `/etc/hosts` overrides or `resolvectl flush-caches` to pick up the change. Since the service is now fully public, the long-term cleaner option is to drop the override entirely and let it fall through to the IONOS wildcard (which also points at the VPS). +Authentik migrated off LXC 124 onto the VPS (see [investigation](../../../investigations/archive/2026-05-31-authentik-vps-migration.md)). The dnsmasq entry changed from `192.168.8.175` (home Caddy) to `82.165.190.79` (VPS traefik). This is the first LAN entry that intentionally points at the VPS rather than Caddy — `auth` is now a genuinely public service served directly from the VPS. **Gotcha logged:** the NetBird per-client resolver (`100.122.255.254`) caches dnsmasq answers and does **not** clear on `netbird down/up`; clients needed `/etc/hosts` overrides or `resolvectl flush-caches` to pick up the change. Since the service is now fully public, the long-term cleaner option is to drop the override entirely and let it fall through to the IONOS wildcard (which also points at the VPS). ### 2026-05-14 — `nfs-export.hubris.network` added (direct, non-HTTP) NFSv4 export server [nfs-export (102)](../containers/102-nfs-export.md) at `192.168.8.200`. Direct entry, not Caddy-fronted — NFS is L4, no HTTP reverse-proxy meaningful. @@ -119,7 +119,7 @@ NFSv4 export server [nfs-export (102)](../containers/102-nfs-export.md) at `192. New LAN entry for [100-zimaos](../vms/100-zimaos.md) → [caddy (121)](../containers/121-caddy.md) → `192.168.8.195`. Briefly pointed direct-to-VM during install for the initial smoke-test, then re-pointed once a Caddyfile block was added (`reverse_proxy 192.168.8.195` + IONOS DNS-01 TLS). ### 2026-05-13 — `plato.hubris.network` added; `files.hubris.network` removed -New LAN-only entry for [plato (126)](../containers/126-plato.md). Same day, the `files.hubris.network` entry for the just-decommissioned seafile experiment was dropped; queries now fall through to the public IONOS answer (no LAN backend). +New LAN-only entry for [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology). Same day, the `files.hubris.network` entry for the just-decommissioned seafile experiment was dropped; queries now fall through to the public IONOS answer (no LAN backend). ### 2026-05-12 — `files.hubris.network` added (since removed 2026-05-13) Originally added for the seafile (LXC 125) Nextcloud-replacement evaluation. Pointed at 192.168.8.175 (Caddy reverse-proxied to 192.168.8.185:80). Entry removed when the experiment was torn down a day later. diff --git a/knowledge/wiki/infrastructure/index.md b/knowledge/wiki/infrastructure/index.md index bb451f1..5e21264 100644 --- a/knowledge/wiki/infrastructure/index.md +++ b/knowledge/wiki/infrastructure/index.md @@ -36,7 +36,7 @@ are documented in their own pages. Each system below links to its full doc. ## Identity & access -- **[Authentik SSO](../containers/124-authentik.md)** — identity provider. +- **[Authentik SSO](../containers/106-auth-outpost.md)** — identity provider. Core server runs on the VPS; LAN forward-auth outpost at LXC 106. OIDC providers configured for Jellyfin, Jellyseerr, Sabnzbd, qBittorrent, Yuvomi, and more. diff --git a/knowledge/wiki/infrastructure/ingress.md b/knowledge/wiki/infrastructure/ingress.md index 75e6c5f..ebc359d 100644 --- a/knowledge/wiki/infrastructure/ingress.md +++ b/knowledge/wiki/infrastructure/ingress.md @@ -49,7 +49,7 @@ LAN clients resolve via the [Technitium DNS on dns (107)](dns.md) → `192.168.8 ### `auth.hubris.network` — different pattern (local container, not cert-mirror) -Since 2026-05-31 [Authentik runs on the VPS itself](../../../investigations/2026-05-31-authentik-vps-migration.md), so `auth.hubris.network` is served by a **local Docker container**, not proxied to a home backend. It therefore does **not** use the file-provider + cert-mirror pattern above: +Since 2026-05-31 [Authentik runs on the VPS itself](../../../investigations/archive/2026-05-31-authentik-vps-migration.md), so `auth.hubris.network` is served by a **local Docker container**, not proxied to a home backend. It therefore does **not** use the file-provider + cert-mirror pattern above: - Routed via traefik **Docker provider labels** on the `authentik-server` service (`/opt/docker-compose.yml`), not `traefik-dynamic.yaml`. - TLS via traefik's own `letsencrypt` resolver (works here because it's a normal HTTP router, not the HostSNI passthrough). @@ -91,7 +91,7 @@ No cert-mirror entry and no `hubris-public-cert-sync.sh` mapping is needed for ` TRMNL plugins middleware on [trmnl (128)](../containers/128-trmnl.md). File-provider router `trmnl-public` → `192.168.8.211:9851`, `trmnl-ratelimit` (20 rps / 40 burst), cert mirrored as `trmnl.fullchain.crt`/`trmnl.privkey.key`. Verified live from the internet (200 with token / 401 without). It was provisioned during a mesh outage — the `home-lab-network` (192.168.8.0/24) route had no active routing peer because the **mac-mini routing peer's netbird was down** (all home-backed public services 504'd). Bringing netbird up on mac-mini restored the route; no traefik change was needed. ### 2026-05-31 — `auth.hubris.network` now served locally on the VPS -Authentik migrated onto the VPS ([investigation](../../../investigations/2026-05-31-authentik-vps-migration.md)). Unlike the home-backed services above, `auth` is a local container routed via traefik Docker-provider labels with traefik-managed Let's Encrypt — no cert-mirror, no `traefik-dynamic.yaml` router. Admin UI gated by an ipAllowList middleware. Traefik gained a second Docker network (`auth`, `172.30.1.0/24`) to reach it while keeping its DB/Redis isolated from the netbird stack. +Authentik migrated onto the VPS ([investigation](../../../investigations/archive/2026-05-31-authentik-vps-migration.md)). Unlike the home-backed services above, `auth` is a local container routed via traefik Docker-provider labels with traefik-managed Let's Encrypt — no cert-mirror, no `traefik-dynamic.yaml` router. Admin UI gated by an ipAllowList middleware. Traefik gained a second Docker network (`auth`, `172.30.1.0/24`) to reach it while keeping its DB/Redis isolated from the netbird stack. ### 2026-04-28 — wiki entry created Initial documentation. diff --git a/knowledge/wiki/infrastructure/mesh.md b/knowledge/wiki/infrastructure/mesh.md index 52711cc..f39c016 100644 --- a/knowledge/wiki/infrastructure/mesh.md +++ b/knowledge/wiki/infrastructure/mesh.md @@ -109,7 +109,7 @@ Recipe for container-config changes (e.g. adding `extra_hosts`) on Portainer-man ## Related - [DNS split-horizon](dns.md) -- [Authentik (124)](../containers/124-authentik.md) — the IdP that triggers most of these overrides +- [Authentik (124)](../containers/106-auth-outpost.md) — the IdP that triggers most of these overrides - [Nextcloud (114)](../containers/114-nextcloud.md) — example of Technique B - [Gitea (104)](../containers/104-gitea.md) — example of Technique A - [Public ingress (VPS traefik)](ingress.md) — uses the same mesh as transport @@ -117,7 +117,7 @@ Recipe for container-config changes (e.g. adding `extra_hosts`) on Portainer-man ## Changelog ### 2026-05-31 (later) — Authentik moved to the VPS; mesh-dependency for auth eliminated (supersedes the band-aid below) -The earlier same-day fix routed `auth.hubris.network` through VPS Traefik → Caddy → LXC 124 **over the mesh**. That restored service but re-created the original fragility: if the mesh is dark when management restarts, the `192.168.8.175` backend is unreachable and management crash-loops again (the "Bootstrap note" in the entry below). That note is now **obsolete** — Authentik was migrated onto the VPS itself, so OIDC no longer touches the mesh. The `auth-authentik` → `192.168.8.175` route and its `skip-verify` transport were removed from `/opt/traefik-dynamic.yaml`; `auth.hubris.network` is now served by a local `authentik-server` container via Traefik Docker-provider labels, and netbird-mgmt has `depends_on: authentik-server: condition: service_healthy`. The socat / reverse-SSH bootstrap dance is no longer needed. Full detail: [2026-05-31 Authentik VPS migration](../../../investigations/2026-05-31-authentik-vps-migration.md). +The earlier same-day fix routed `auth.hubris.network` through VPS Traefik → Caddy → LXC 124 **over the mesh**. That restored service but re-created the original fragility: if the mesh is dark when management restarts, the `192.168.8.175` backend is unreachable and management crash-loops again (the "Bootstrap note" in the entry below). That note is now **obsolete** — Authentik was migrated onto the VPS itself, so OIDC no longer touches the mesh. The `auth-authentik` → `192.168.8.175` route and its `skip-verify` transport were removed from `/opt/traefik-dynamic.yaml`; `auth.hubris.network` is now served by a local `authentik-server` container via Traefik Docker-provider labels, and netbird-mgmt has `depends_on: authentik-server: condition: service_healthy`. The socat / reverse-SSH bootstrap dance is no longer needed. Full detail: [2026-05-31 Authentik VPS migration](../../../investigations/archive/2026-05-31-authentik-vps-migration.md). ### 2026-05-31 — Netbird mesh recovered; auth.hubris.network exposed via VPS Traefik @@ -139,11 +139,11 @@ The earlier same-day fix routed `auth.hubris.network` through VPS Traefik → Ca ### 2026-05-21 — VPS migrated combined → vanilla netbird stack with external TURN The combined `netbirdio/netbird-server` image was replaced with the canonical multi-container deploy (`netbirdio/management:0.71.3` + `signal:0.71.3` + `relay:0.71.3` + `dashboard:latest` + host coturn) on `/opt/docker-compose.yml`. Driver: combined image silently ignored external `TURNConfig` so symmetric-NAT peers couldn't use TURN. -Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on [LXC 124](../containers/124-authentik.md), upgrading mgmt to 0.71.3. The `store.db` schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old `opt_netbird_data` volume into the new `mgmt_data` volume). Pre-cutover backups at `/root/netbird-*.tgz` on the VPS, ~857 MB, retained for ~7d. +Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on [LXC 124](../containers/106-auth-outpost.md), upgrading mgmt to 0.71.3. The `store.db` schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old `opt_netbird_data` volume into the new `mgmt_data` volume). Pre-cutover backups at `/root/netbird-*.tgz` on the VPS, ~857 MB, retained for ~7d. Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478 inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe). -The new Authentik provider for NetBird is `Client type: Public` (PKCE-only). Confidential would break the dashboard SPA's token exchange. The Device Code grant flow is wired (see [containers/124-authentik.md](../containers/124-authentik.md#device-code-grant--configured-2026-05-21)) so interactive `netbird up` works — `--setup-key` is no longer required for new peers. +The new Authentik provider for NetBird is `Client type: Public` (PKCE-only). Confidential would break the dashboard SPA's token exchange. The Device Code grant flow is wired (see [containers/124-authentik.md](../containers/106-auth-outpost.md#device-code-grant--configured-2026-05-21)) so interactive `netbird up` works — `--setup-key` is no longer required for new peers. **Post-migration JWT-issuer gotcha on existing peers** (cost ~30 min to diagnose 2026-05-21): diff --git a/knowledge/wiki/infrastructure/network.md b/knowledge/wiki/infrastructure/network.md index fb94a5d..99f4c08 100644 --- a/knowledge/wiki/infrastructure/network.md +++ b/knowledge/wiki/infrastructure/network.md @@ -79,10 +79,10 @@ No NAT on Proxmox — traffic flows without double-NAT. ### 2026-06-17 — Fritz!Box DNSv4 server set to Technitium (192.168.8.2) Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs. Configured in Fritz!Box at Internet → Filter → DNS Server → DNSv4 Server → "Use other DNSv4 servers" → Preferred = `192.168.8.2`. No per-device or Netbird setup needed. -Previous pool `.100–.240` overlapped with all static LXCs/VMs (` .101–.239`), creating IP conflict risk (DHCP could hand out an IP that a static service expects). Shrunk pool to `.241–.254` via Technitium API. No services re-IP'd. 11 stale DHCP leases in `.101–.110` will expire naturally. **Open:** ZimaOS (VM 100) holds DHCP lease `.103` but inventory expects `.195` — needs static IP set inside VM. See [plan](../../../plans/2026-06-03-dhcp-pool-exclude-static-ips.md). +Previous pool `.100–.240` overlapped with all static LXCs/VMs (` .101–.239`), creating IP conflict risk (DHCP could hand out an IP that a static service expects). Shrunk pool to `.241–.254` via Technitium API. No services re-IP'd. 11 stale DHCP leases in `.101–.110` will expire naturally. **Open:** ZimaOS (VM 100) holds DHCP lease `.103` but inventory expects `.195` — needs static IP set inside VM. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md). ### 2026-06-02 — Executed migration; Proxmox as subnet router -Fritz!OS 8.x does not support second IP networks on LAN ports, so the final design uses Proxmox as the router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10`; `vmbr0` is a portless internal bridge with `192.168.8.1` alias as the LXC gateway. Technitium DHCP enabled for `192.168.8.100–240`. Caddy service unit was missing and recreated. See [migration plan](../../../plans/2026-06-01-slate-ax-to-sodola-migration.md). +Fritz!OS 8.x does not support second IP networks on LAN ports, so the final design uses Proxmox as the router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10`; `vmbr0` is a portless internal bridge with `192.168.8.1` alias as the LXC gateway. Technitium DHCP enabled for `192.168.8.100–240`. Caddy service unit was missing and recreated. See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md). ### 2026-06-01 — Initial network doc; Slate AX retired; SODOLA switch added -Replaced the GL.iNet Slate AX sub-router with the SODOLA 5-Port 2.5Gbit managed switch. Eliminated double-NAT. See [migration plan](../../../plans/2026-06-01-slate-ax-to-sodola-migration.md). +Replaced the GL.iNet Slate AX sub-router with the SODOLA 5-Port 2.5Gbit managed switch. Eliminated double-NAT. See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md). diff --git a/knowledge/wiki/vms/100-zimaos.md b/knowledge/wiki/vms/100-zimaos.md index ac58ac6..ac68f42 100644 --- a/knowledge/wiki/vms/100-zimaos.md +++ b/knowledge/wiki/vms/100-zimaos.md @@ -46,7 +46,7 @@ The alternative (dedicated virtual data disk on the `library` lvmthin pool, e.g. ## Open items - **DHCP → static IP fixed (2026-06-03).** ZimaOS IP drifted from `.195` (Slate AX) → `.103` (Technitium) after the DHCP migration, causing Caddy 502s. Fixed by injecting a static systemd-networkd config and restarting the VM. IP now pinned at `192.168.8.195`. See [changelog](#2026-06-03--static-ip-set-to-195-dhcp-drift-fixed). -- **No Authentik wiring.** [authentik (124)](../containers/124-authentik.md) isn't enforcing auth in front of ZimaOS yet — ZimaOS handles its own first-run wizard. The Caddyfile block uses bare `reverse_proxy` rather than the `import authentik` pattern used by e.g. artifacto; layer it in once the wizard is complete and a static admin user exists. +- **No Authentik wiring.** [authentik (124)](../containers/106-auth-outpost.md) isn't enforcing auth in front of ZimaOS yet — ZimaOS handles its own first-run wizard. The Caddyfile block uses bare `reverse_proxy` rather than the `import authentik` pattern used by e.g. artifacto; layer it in once the wizard is complete and a static admin user exists. - **No PBS backup.** No Proxmox Backup Server configured on hubris today; this VM is not backed up. - **qemu-guest-agent not installed.** ZimaOS's installer doesn't bundle it, so `qm guest cmd 100 ...` returns "QEMU guest agent is not running". IP discovery during this install was done via console screendump → `qm monitor` → `screendump`. @@ -59,7 +59,7 @@ The alternative (dedicated virtual data disk on the `library` lvmthin pool, e.g. ## Changelog ### 2026-06-03 — Static IP set to `.195`; DHCP drift fixed -ZimaOS had drifted from `.195` (Slate AX DHCP) → `.103` (Technitium DHCP), causing Caddy 502s. Injected `/etc/systemd/network/10-static.network` into overlay (match `en*/eth*`, address `192.168.8.195/24`, gateway `.1`, DNS `.2`). VM restarted; verified reachable at `.195`. Caddy (`zimaos.hubris.network`) now returns 200. See [plan](../../../plans/2026-06-03-dhcp-pool-exclude-static-ips.md). +ZimaOS had drifted from `.195` (Slate AX DHCP) → `.103` (Technitium DHCP), causing Caddy 502s. Injected `/etc/systemd/network/10-static.network` into overlay (match `en*/eth*`, address `192.168.8.195/24`, gateway `.1`, DNS `.2`). VM restarted; verified reachable at `.195`. Caddy (`zimaos.hubris.network`) now returns 200. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md). ### 2026-05-15 — NFS mount relocated to `/media/library` (UI delete fix) @@ -87,4 +87,4 @@ Virtiofs path abandoned — ZimaOS kernel 6.12.25 ships without the virtiofs mod Added `zimaos.hubris.network` site block to `/etc/caddy/Caddyfile` on [caddy (121)](../containers/121-caddy.md): bare `reverse_proxy 192.168.8.195` + IONOS DNS-01 TLS, same pattern as plato/jellyfin. dnsmasq entry repointed from `192.168.8.195` to `192.168.8.175`. Let's Encrypt cert issued on first request. Caddy commit `a219176` pending push to `dtoro/caddy-conf`. ### 2026-05-14 — VM created, ZimaOS 1.6.1 installed (Phase 1) -`qm create 100` with q35/OVMF, no EFI disk, 4 vCPU / 8 GiB / 64 GiB on `local-lvm`. Installed via the official ISO (manual console install). Web UI verified at `http://192.168.8.195`. `onboot=1`, `startup order=20`. dnsmasq entry `zimaos.hubris.network → 192.168.8.195` initially added direct-to-VM on [authentik (124)](../containers/124-authentik.md) (later repointed — see above). `/mnt/library` is **not** yet shared into the VM; Phase 2 (virtiofs) is gated on UI evaluation. +`qm create 100` with q35/OVMF, no EFI disk, 4 vCPU / 8 GiB / 64 GiB on `local-lvm`. Installed via the official ISO (manual console install). Web UI verified at `http://192.168.8.195`. `onboot=1`, `startup order=20`. dnsmasq entry `zimaos.hubris.network → 192.168.8.195` initially added direct-to-VM on [authentik (124)](../containers/106-auth-outpost.md) (later repointed — see above). `/mnt/library` is **not** yet shared into the VM; Phase 2 (virtiofs) is gated on UI evaluation. diff --git a/knowledge/wiki/vms/108-haos.md b/knowledge/wiki/vms/108-haos.md index 6ac7ca6..90c78a7 100644 --- a/knowledge/wiki/vms/108-haos.md +++ b/knowledge/wiki/vms/108-haos.md @@ -7,7 +7,7 @@ Home Assistant OS — the only VM on hubris (HAOS doesn't run cleanly in an LXC, - **HAOS version:** 16.3 (last verified) - **IP:** `192.168.8.101` - **Resources:** 4 GiB RAM, 32 GiB boot disk -- **Public hostname:** [`home.hubris.network`](../infrastructure/dns.md) → [caddy (121)](121-caddy.md) → `192.168.8.101:8123` +- **Public hostname:** [`home.hubris.network`](../infrastructure/dns.md) → [caddy (121)](../containers/121-caddy.md) → `192.168.8.101:8123` ## Auth @@ -18,7 +18,7 @@ Key gotchas: ``` ha dns options --servers "dns://192.168.8.180" --servers "dns://1.1.1.1" ``` - so OIDC discovery resolves internally to [authentik (124)](124-authentik.md). + so OIDC discovery resolves internally to [authentik (124)](../containers/106-auth-outpost.md). - Authentik app slug in the discovery URL is whatever was set in Authentik — confirm via the DB rather than guessing. User set `home-assistant` (with hyphen). - YAML config: - `features.automatic_user_linking: true` — link to existing HA users by `preferred_username` match (otherwise a duplicate is created). @@ -30,8 +30,8 @@ Key gotchas: HA pulls Proxmox metrics via the official Proxmox VE integration. As of 2026-04-21 [claudio-monitor](../infrastructure/monitoring.md) stopped publishing to MQTT/REST (commit `82f0596`) — HA gets metrics from PVE directly; claudio-monitor focuses on alerting. ## Related -- [Authentik (124)](124-authentik.md) -- [Caddy (121)](121-caddy.md) +- [Authentik (124)](../containers/106-auth-outpost.md) +- [Caddy (121)](../containers/121-caddy.md) - [DNS](../infrastructure/dns.md) - [Monitoring](../infrastructure/monitoring.md) diff --git a/operations/agent-enrollment.md b/operations/agent-enrollment.md index dfd41a6..be0b15b 100644 --- a/operations/agent-enrollment.md +++ b/operations/agent-enrollment.md @@ -40,7 +40,7 @@ Bootstrap auto-installs netbird and drives `netbird up` if the mesh isn't alread The new client runs bootstrap straight from a fresh OS. Bootstrap installs netbird (apt/dnf/brew based on the OS), then runs `netbird up --management-url https://netbird.hubris.network --ssh-jwt-cache-ttl 86400`. A device-code URL prints inline. The operator opens it (in a browser logged into Authentik), goes through identification → password → consent, and the CLI returns `Connected`. Bootstrap then proceeds with the rest of preflight. -Pre-condition: the operator must be a registered user in Authentik (typically the lab owner). The first user-login against a netbird account with existing peers is added as `pending_approval=1` and needs an sqlite promotion to `owner` — see [124-authentik.md First-time owner promotion gotcha](../knowledge/wiki/containers/124-authentik.md). Only needed once per account. +Pre-condition: the operator must be a registered user in Authentik (typically the lab owner). The first user-login against a netbird account with existing peers is added as `pending_approval=1` and needs an sqlite promotion to `owner` — see [124-authentik.md First-time owner promotion gotcha](../knowledge/wiki/containers/106-auth-outpost.md). Only needed once per account. **Path A — setup-key (headless/scripted onboarding):** @@ -261,7 +261,7 @@ arguments. If you also want the netbird `--ssh-jwt-cache-ttl` flag rationale to be visible to the classifier (it's not actually durable in 0.71.2, but the -ControlMaster block is — see [runbook-dpkg-interrupted](runbook-dpkg-interrupted.md) +ControlMaster block is — see [runbook-dpkg-interrupted](../.agents/skills/runbook-dpkg-interrupted/SKILL.md) for context), drop a free-text rule into `autoMode.allow` describing the authorization. Optional. @@ -342,7 +342,7 @@ The CLI prints a follow-up checklist that the operator must do manually: | `homelab` CLI doesn't pick up repo updates | Pre-`02db…` bootstrap copied the binary instead of symlinking | One-time migration: `sudo ln -sfn /opt/homelab-context/bin/homelab /usr/local/bin/homelab`. New bootstraps use the symlink, which auto-tracks the synced repo. | | `homelab-context-sync.service` journal shows `fatal: could not read Username for 'https://git.hubris.network'` | Pre-fix bootstrap set the gitea credential helper via `git config --global`, which writes to `/root/.gitconfig` — invisible to the systemd timer's git process (no HOME set). | One-time migration: `sudo git config --system credential.helper "store --file=/etc/homelab-context/git-credentials"`. New bootstraps store the helper in `/etc/gitconfig` instead. | | Chat-mode `!` shell can't `sudo` (`a terminal is required to read the password`) | Claude Code's `!` invocation doesn't allocate a tty, and standard `sudo` won't read its password from stdin or a non-tty pipe. | Run the sudo'd command in a real terminal outside chat. For commands the agent issues repeatedly, configure passwordless sudo for the narrow set (e.g. `/etc/sudoers.d/homelab-self` with ` ALL=(ALL) NOPASSWD: /usr/bin/dnf upgrade -y, /usr/bin/apt-get *`). | -| `netbird status -d` reports `192.168.8.180:53 ... is Unavailable` but DNS actually works | netbird's UDP-53 probe times out over the relay latency (~90ms), but actual queries still flow through systemd-resolved. Cosmetic. | Ignore unless `dig @192.168.8.180 git.hubris.network` also fails — then check dnsmasq on [LXC 124](../knowledge/wiki/containers/124-authentik.md). | +| `netbird status -d` reports `192.168.8.180:53 ... is Unavailable` but DNS actually works | netbird's UDP-53 probe times out over the relay latency (~90ms), but actual queries still flow through systemd-resolved. Cosmetic. | Ignore unless `dig @192.168.8.180 git.hubris.network` also fails — then check dnsmasq on [LXC 124](../knowledge/wiki/containers/106-auth-outpost.md). | | `netbird ssh` rejected with `JWT authentication failed: validate token (expected issuer=https://netbird.hubris.network/oauth2 ...)` | Peer's SSH JWT validator cached the OLD embedded-Dex issuer from before the 2026-05-21 Authentik migration. `systemctl restart netbird` and `netbird down/up` don't clear it — `client/internal/engine_ssh.go` bails out of `updateSSH()` if the SSH server is already running. | Full daemon bounce: `sudo systemctl stop netbird; sleep 3; sudo systemctl start netbird`. Verify with `grep -iE "issuer\|audience" /var/log/netbird/client.log \| tail`. Apply once per peer post-migration. | | `netbird ssh` JWT passes but session closes with `user privilege check failed: user dtoro not found: unknown user dtoro` | netbird-ssh defaults the remote username to the LOCAL one (operator's laptop user). Hubris and LXCs only have `root`. | Always use explicit `root@` prefix manually: `netbird ssh -p 22022 root@proxmox-server.netbird.selfhosted`. `homelab ssh ` does this automatically via `inventory.yaml`'s per-host `ssh.user` field (defaults to `root`). | | `homelab ssh hubris` (or any host on the LAN) fails with `Connection refused` or hangs, despite mesh routing being up | Off-LAN networks (operator on a VPN / coffee shop / symmetric NAT) sometimes can't reach the LAN IP even with the netbird subnet route. | Newer homelab CLIs probe the LAN with a 1.5s TCP connect and transparently fall back to the netbird FQDN. If your `/usr/local/bin/homelab` is a symlink to `/opt/homelab-context/bin/homelab` it'll pick up the fix on the next 5-min context sync. Otherwise pull the latest from gitea. | diff --git a/operations/commands.md b/operations/commands.md index bdc1cd7..797953b 100644 --- a/operations/commands.md +++ b/operations/commands.md @@ -49,7 +49,7 @@ Run from the [hubris host](../knowledge/wiki/hosts/hubris.md) as root. When work - `ras-mc-ctl --errors` — full event log - `cat /sys/devices/system/cpu/cpu0/cpufreq/energy_performance_preference` — should be `balance_power` - `cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor` — should be `powersave` -- `ls /sys/fs/pstore/ /var/lib/systemd/pstore/` — panic traces from a previous crash (empty for pure hardware hangs — see [investigation](../investigations/2026-04-21-hubris-crash-loop.md)) +- `ls /sys/fs/pstore/ /var/lib/systemd/pstore/` — panic traces from a previous crash (empty for pure hardware hangs — see [investigation](../investigations/archive/2026-04-21-hubris-crash-loop.md)) ## Fleet apt operations @@ -88,4 +88,4 @@ Oikos Console (read-mostly dashboard): `oikos.hubris.network` once deployed — - [DNS](../knowledge/wiki/infrastructure/dns.md) - [Monitoring](../knowledge/wiki/infrastructure/monitoring.md) - [Auto-deploy](../knowledge/wiki/infrastructure/auto-deploy.md) -- [Runbook: dpkg-interrupted recovery](runbook-dpkg-interrupted.md) — what to do when apt got killed mid-transaction +- [Runbook: dpkg-interrupted recovery](../.agents/skills/runbook-dpkg-interrupted/SKILL.md) — what to do when apt got killed mid-transaction diff --git a/plans/2026-06-24-trmnl-plugins-lxc.md b/plans/2026-06-24-trmnl-plugins-lxc.md index 294219d..78c88ba 100644 --- a/plans/2026-06-24-trmnl-plugins-lxc.md +++ b/plans/2026-06-24-trmnl-plugins-lxc.md @@ -14,7 +14,7 @@ wiring, same split as Artifacto/Plato. ## Current state - No TRMNL middleware in the lab. Highest LXC id is 127 (see `containers/index.md`). -- Public hostnames terminate at the [VPS netbird traefik](../hosts/netbird-vps.md) → netbird +- Public hostnames terminate at the [VPS netbird traefik](../hosts/netbird-vps.yaml) → netbird mesh → [caddy (121)](../knowledge/wiki/containers/121-caddy.md) → backend LXC. Cert obtained by Caddy (IONOS DNS-01) and mirrored to the VPS by the daily cert-sync timer on the host. - Auto-deploy pipelines are gitea-webhook driven, two shapes (see [auto-deploy](../knowledge/wiki/infrastructure/auto-deploy.md)). @@ -81,7 +81,7 @@ TRMNL cloud --GET 15m, Bearer token--> https://trmnl.hubris.network/munich-hom trmnl.hubris.network { reverse_proxy 192.168.8.<128-ip>:9851 } ``` -7. **Public exposure** on the [VPS](../hosts/netbird-vps.md): add traefik router+service for +7. **Public exposure** on the [VPS](../hosts/netbird-vps.yaml): add traefik router+service for `Host(\`trmnl.hubris.network\`)` → `http://192.168.8.<128-ip>:9851`; add the host to the cert-sync map so the LE cert mirrors over. diff --git a/plans/done/2026-06-01-slate-ax-to-sodola-migration.md b/plans/done/2026-06-01-slate-ax-to-sodola-migration.md index 8f67eae..62fdc67 100644 --- a/plans/done/2026-06-01-slate-ax-to-sodola-migration.md +++ b/plans/done/2026-06-01-slate-ax-to-sodola-migration.md @@ -37,7 +37,7 @@ ISP Fritz!Box takes over `192.168.8.1` — the same gateway IP the Slate AX used. No static IPs or gateway entries change on any LXC or VM. -See [network architecture](../infrastructure/network.md) for the permanent topology reference. +See [network architecture](../../knowledge/wiki/infrastructure/network.md) for the permanent topology reference. ## Pre-flight checklist @@ -60,7 +60,7 @@ See [network architecture](../infrastructure/network.md) for the permanent topol | DHCP range | `192.168.8.100 – 192.168.8.240` | | Assign to | LAN port that connects to SODOLA | | Network isolation | Enabled (blocks main LAN from initiating into homelab) | -| DNS for DHCP clients | `192.168.8.2` (Technitium on [CT 107](../containers/107-dns.md)) | +| DNS for DHCP clients | `192.168.8.2` (Technitium on [CT 107](../../knowledge/wiki/containers/107-dns.md)) | After creating the network, move any port forwards from the Slate AX into Fritz!Box → Internet → Permits (target IPs are now directly reachable on `192.168.8.x`). @@ -92,7 +92,7 @@ pct set --net0 name=eth0,bridge=vmbr0,ip=/24,gw=192.168.8.1 ## DNS after migration -Technitium ([CT 107](../containers/107-dns.md)) at `192.168.8.2` continues to serve split-horizon DNS for `hubris.network`. The Fritz!Box DHCP server for VLAN 10 hands out `192.168.8.2` as the DNS server. This fixes the "update router DHCP DNS from dead .180 → .2" outstanding item in [dns.md](../infrastructure/dns.md). +Technitium ([CT 107](../../knowledge/wiki/containers/107-dns.md)) at `192.168.8.2` continues to serve split-horizon DNS for `hubris.network`. The Fritz!Box DHCP server for VLAN 10 hands out `192.168.8.2` as the DNS server. This fixes the "update router DHCP DNS from dead .180 → .2" outstanding item in [dns.md](../../knowledge/wiki/infrastructure/dns.md). ## Cutover procedure @@ -125,7 +125,7 @@ curl -sk https://auth.hubris.network/if/flow/default-authentication-flow/ | head ## Post-migration -- Update [network.md](../infrastructure/network.md) topology to reflect new state. -- Add changelog entries to [hosts/hubris.md](../hosts/hubris.md) and any affected container pages. -- Update status in [plans/index.md](index.md) to `Done`. +- Update [network.md](../../knowledge/wiki/infrastructure/network.md) topology to reflect new state. +- Add changelog entries to [hosts/hubris.md](../../knowledge/wiki/hosts/hubris.md) and any affected container pages. +- Update status in [plans/index.md](../index.md) to `Done`. - If anything went sideways, open an investigation in `investigations/`. diff --git a/plans/done/2026-06-29-grimmory-migration.md b/plans/done/2026-06-29-grimmory-migration.md index 2f89f23..89526d9 100644 --- a/plans/done/2026-06-29-grimmory-migration.md +++ b/plans/done/2026-06-29-grimmory-migration.md @@ -204,7 +204,7 @@ books.hubris.network { } ``` -Git push → Caddy webhook auto-reloads (see [caddy (121)](../containers/121-caddy.md)). +Git push → Caddy webhook auto-reloads (see [caddy (121)](../../knowledge/wiki/containers/121-caddy.md)). Test: ```bash