6716 lines
518 KiB
YAML
6716 lines
518 KiB
YAML
version: 1
|
||
documents:
|
||
- slug: containers/101-jellyfin
|
||
title: 101 — `jellyfin`
|
||
content: "# 101 — `jellyfin`\n\nMedia server: serves the movies / TV / anime / music libraries from `/mnt/media_local` to\
|
||
\ LAN clients. Hardware transcoding via AMD Radeon 680M + RX 7600 VAAPI. Authentik SSO via OIDC.\n\n## At a glance\n-\
|
||
\ **Host:** **strong** (migrated from hubris 2026-07-05)\n- **IP:** `192.168.8.246`\n- **Privilege:** privileged (recreated\
|
||
\ on strong as priv)\n- **Resources:** 4 cores / 8 GiB RAM / 1 GiB swap / 16 GiB rootfs\n- **GPU:** `/dev/dri/renderD128`\
|
||
\ + `/dev/dri/card0` (AMD Radeon 680M iGPU + RX 7600 dGPU) passed via `dev0` / `dev1` in LXC config\n- **Mounts:** `/mnt/media_local`\
|
||
\ ↔ `/mnt/library`\n- **Public hostname:** [`media.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md)\
|
||
\ → `:8096`\n- **Version:** Jellyfin 10.11.11 (apt package, Ubuntu 24.04 noble repo)\n- **FFmpeg:** jellyfin-ffmpeg7 7.1.4\n\
|
||
\n## Service / port map\n\n| Service | Listen | Notes |\n| -------- | ------ | ----- |\n| jellyfin | `:8096` | HTTP (caddy\
|
||
\ terminates TLS) |\n\n## Hardware acceleration (VAAPI)\n\nGPU is passed through to the LXC via `dev0: /dev/dri/renderD128,gid=993`\
|
||
\ and\n`dev1: /dev/dri/card0,gid=44` in\n`/etc/pve/lxc/101.conf` (strong). The `jellyfin` user is in the `render` (GID\
|
||
\ 993) and\n`video` groups inside the container.\n\n| GPU | Model | Role |\n|-----|-------|------|\n| Radeon 680M | iGPU\
|
||
\ (AMD Ryzen 7 PRO 6850U) | Primary VAAPI encoder/decoder |\n| RX 7600 | dGPU (add-in) | Secondary transcode, HEVC/AV1\
|
||
\ encoding |\n\nEncoding settings (`/etc/jellyfin/encoding.xml`):\n- `HardwareAccelerationType`: `vaapi`\n- `VaapiDevice`:\
|
||
\ `/dev/dri/renderD128`\n- `EnableHardwareEncoding`: `true`\n- `AllowHevcEncoding`: `true`\n- `AllowAv1Encoding`: `true`\n\
|
||
- `EnableTonemapping`: `true`\n- `HardwareDecodingCodecs`: h264, hevc, vc1, vp9, av1\n- `EnableThrottling`: `true`\n-\
|
||
\ `EnableSegmentDeletion`: `true`\n\nTrickplay (`/etc/jellyfin/system.xml`):\n- `EnableHwAcceleration`: `true`\n- `EnableHwEncoding`:\
|
||
\ `true`\n\n## Authentik SSO (OIDC)\n\nJellyfin uses the [SSO-Auth plugin](https://github.com/9p4/jellyfin-plugin-sso)\n\
|
||
v4.0.0.4 for Authentik OIDC login. No Caddy forward-auth gate — the SSO plugin\nhandles auth directly via OIDC redirect\
|
||
\ flow.\n\n### Architecture\n\n```\nUser → media.hubris.network → Caddy (TLS, no forward-auth) → Jellyfin :8096\n \
|
||
\ ↓\n Login\
|
||
\ page with \"Sign in with Authentik\" button\n ↓ (click)\n\
|
||
\ /sso/OID/start/Authentik\n \
|
||
\ ↓ (302 redirect)\n auth.hubris.network OIDC\n\
|
||
\ ↓ (login)\n \
|
||
\ /sso/OID/redirect/Authentik?code=...&state=...\n \
|
||
\ ↓\n Jellyfin SSO plugin validates token → logged in\n```\n\n\
|
||
### Components\n\n1. **SSO-Auth plugin** — installed at `/var/lib/jellyfin/plugins/SSO-Auth_4.0.0.4/`\n - Config: `/var/lib/jellyfin/plugins/configurations/SSO-Auth.xml`\n\
|
||
\ - Provider name: `Authentik`\n - OIDC endpoint: `https://auth.hubris.network/application/o/jellyfin/`\n - `SchemeOverride`:\
|
||
\ `https` (required — without it, plugin generates\n `http://` redirect URIs that Authentik rejects)\n - `EnableAuthorization`:\
|
||
\ `false` (prevents plugin from overwriting admin\n permissions on each SSO login — see\n [jellyfin-sso-plugin](../../../devops/homelab-authentik-admin/references/jellyfin-sso-plugin.md))\n\
|
||
\ - `OidScopes`: `[\"email\"]` (openid+profile added by default by the plugin;\n must be non-null or `OidChallenge()`\
|
||
\ throws `ArgumentNullException`)\n\n2. **Authentik OIDC provider** — `Provider for Jellyfin` (PK 6)\n - Client ID:\
|
||
\ `vt61t5Y2ZVtN6l3QjitkBvwUJjFKvSyl4TDBXcJx`\n - Redirect URI: `https://media.hubris.network/sso/OID/redirect/Authentik`\n\
|
||
\ - Application slug: `jellyfin`\n\n3. **SSO button injection** — `/usr/share/jellyfin/web/sso-inject.js`\n - Injected\
|
||
\ via `<script defer src=\"sso-inject.js?v=3\">` in `index.html`\n - Polls for `.readOnlyContent` div on the login page,\
|
||
\ adds \"Sign in with\n Authentik\" button linking to `/sso/OID/start/Authentik`\n - Cache-busted with `?v=N` parameter\
|
||
\ (bump on changes)\n - **Lost on apt upgrade** — re-inject the script tag and copy the JS file\n after `apt-get\
|
||
\ upgrade jellyfin*`\n\n4. **Caddy config** — no forward-auth gate for media.hubris.network:\n ```caddy\n media.hubris.network\
|
||
\ {\n tls { dns ionos {env.IONOS_AUTH_API_TOKEN} }\n reverse_proxy 192.168.8.206:8096\n }\n ```\n\n###\
|
||
\ Known issues / pitfalls\n\n- **`OidScopes` must be non-null** — if the field is missing from the plugin\n config XML,\
|
||
\ `OidChallenge()` throws `System.ArgumentNullException`. Always\n include `OidScopes` in the provider config (even if\
|
||
\ empty array).\n- **`SchemeOverride: \"https\"` is required** — without it, the plugin generates\n `http://` redirect\
|
||
\ URIs (from the internal HTTP listener). Authentik rejects\n them with \"Redirect URI Error\".\n- **SSO button JS is\
|
||
\ not served by the plugin** — the `__plugin/SSO-Auth.js`\n endpoint returns 404 on Jellyfin 10.11.x when the plugin\
|
||
\ is installed\n manually (not via Jellyfin's plugin manager). The `sso-inject.js` workaround\n in `index.html` is the\
|
||
\ fallback.\n- **No Caddy forward-auth gate** — the SSO plugin's OIDC redirect flow is\n incompatible with Caddy's `import\
|
||
\ authentik` forward-auth. If both are\n enabled, the forward-auth intercepts the OIDC callback and breaks the flow.\n\
|
||
\ Use one or the other, not both. SSO plugin (OIDC redirect) is preferred.\n- **API key for setup** — a temp API key\
|
||
\ can be inserted directly into the\n `ApiKeys` SQLite table for automated configuration:\n ```sql\n INSERT INTO ApiKeys\
|
||
\ VALUES (1, '2026-07-04', '2026-07-04', 'setup', 'jf-setup-key-...');\n ```\n\n## Permissions\nMember of the [media\
|
||
\ GID 10000](../infrastructure/media-permissions.md) standard. Service user `jellyfin` is in the `media` group inside\
|
||
\ the container; `/mnt/media_local` on strong's ludo-lvm is owned `root:media` with mode `2775`.\n\n## Related\n- [Caddy\
|
||
\ reverse proxy](121-caddy.md)\n- [Media permissions](../infrastructure/media-permissions.md)\n- [arriman](122-arriman.md)\
|
||
\ — \\*arr stack writes the libraries jellyfin reads\n- [DNS split-horizon](../infrastructure/dns.md)\n- [Authentik admin](../../../devops/homelab-authentik-admin/SKILL.md)\
|
||
\ — OIDC provider creation, SSO plugin config\n\n## Changelog\n\n### 2026-07-06 — wiki: IP, host, GPU, mount path updated\
|
||
\ for strong migration\nUpdated At-a-glance: IP 206→246, host hubris→strong, mount /mnt/library→/mnt/media_local, GPU\
|
||
\ Radeon 760M→680M+RX7600, privilege unpriv→priv. Permissions section updated. Changelog entry in 122-arriman.md updated\
|
||
\ similarly.\n\n### 2026-07-04 — VAAPI hardware acceleration + Authentik SSO + resource bump\n- Upgraded Jellyfin 10.11.8\
|
||
\ → 10.11.11 (purge + reinstall to fix DB migration bug)\n- Enabled VAAPI hardware acceleration (Radeon 760M): h264/hevc/vc1/vp9/av1\
|
||
\ decode + encode\n- Bumped resources: 2→4 cores, 4→8 GiB RAM, 512→1024 MiB swap\n- Enabled trickplay HW acceleration\
|
||
\ + throttling + segment deletion\n- Installed SSO-Auth plugin v4.0.0.4 with Authentik OIDC provider\n- Configured `SchemeOverride:\
|
||
\ https`, `OidScopes: [\"email\"]`, `EnableAuthorization: false`\n- Injected SSO button via `sso-inject.js` in web `index.html`\n\
|
||
- Removed Caddy forward-auth gate (incompatible with SSO plugin OIDC flow)\n- **Database was wiped** during cache relocation\
|
||
\ attempt — no LVM snapshot\n existed. All watch states, user accounts, and library configs lost.\n Libraries re-added\
|
||
\ via setup wizard.\n\n### 2026-04-28 — wiki entry created\nInitial documentation. No config changes.\n\n### 2026-04-20\
|
||
\ — joined the `media` GID 10000 standard\nIdmap block applied; in-container `media` group at GID 10000 mapped to host\
|
||
\ GID 10000. See [media permissions](../infrastructure/media-permissions.md). Config backup: `/root/101.conf.bak.*`.\n"
|
||
entity_slug: lxc:jellyfin
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
host: strong
|
||
ip: 192.168.8.246
|
||
privileged: privileged
|
||
resources: 4 cores / 8 GiB RAM / 1 GiB swap / 16 GiB rootfs
|
||
gpu: /dev/dri/renderD128 + /dev/dri/card0 (AMD Radeon 680M iGPU + RX 7600 dGPU) passed via dev0 / dev1 in LXC config
|
||
mounts: /mnt/media_local /mnt/library
|
||
public_host: media.hubris.network → caddy → :8096
|
||
version: Jellyfin 10.11.11
|
||
ffmpeg: jellyfin-ffmpeg7 7.1.4
|
||
changelog:
|
||
- date: '2026-07-06'
|
||
title: 'wiki: IP, host, GPU, mount path updated for strong migration'
|
||
body: 'Updated At-a-glance: IP 206→246, host hubris→strong, mount /mnt/library→/mnt/media_local, GPU Radeon 760M→680M+RX7600,
|
||
privilege unpriv→priv. Permissions section updated. Changelog entry in 122-arriman.md updated similarly.'
|
||
- date: '2026-07-04'
|
||
title: VAAPI hardware acceleration + Authentik SSO + resource bump
|
||
body: '- Upgraded Jellyfin 10.11.8 → 10.11.11 (purge + reinstall to fix DB migration bug) - Enabled VAAPI hardware acceleration
|
||
(Radeon 760M): h264/hevc/vc1/vp9/av1 decode + encode - Bumped resources: 2→4 cores, 4→8 GiB RAM, 512→1024 MiB swap -
|
||
Enabled trickplay HW acceleration + throttling + segment deletion - Installed SSO-Auth plugin v4.0.0.4 with Authentik
|
||
OIDC provider - Configured `SchemeOverride: https`, `OidScopes: ["email"]`, `EnableAuthorization: false` - Injected
|
||
SSO button via `sso-inject.js` in web `index.html` - Removed Caddy forward-auth gate (incompatible with SSO plugin OIDC
|
||
flow) - **Database was wiped** during cache relocation attempt — no LVM snapshot existed. All watch states, user accounts,
|
||
and library configs lost. Libraries re-added via setup wizard.'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation. No config changes.
|
||
- date: '2026-04-20'
|
||
title: joined the `media` GID 10000 standard
|
||
body: 'Idmap block applied; in-container `media` group at GID 10000 mapped to host GID 10000. See [media permissions](../infrastructure/media-permissions.md).
|
||
Config backup: `/root/101.conf.bak.*`.'
|
||
- slug: containers/102-nfs-export
|
||
title: 102 — `nfs-export`
|
||
content: "# 102 — `nfs-export`\n\nDedicated, single-purpose LXC that re-exports `/mnt/library` over NFSv4 to clients that\
|
||
\ can't use the host's PVE virtiofs path — currently only [100-zimaos](../vms/100-zimaos.md), which ships a kernel without\
|
||
\ virtiofs support.\n\n## At a glance\n- **Hostname:** `nfs-export`\n- **IP:** `192.168.8.200` (static; LAN-only, no Caddy\
|
||
\ in front because NFS is L4)\n- **LAN DNS:** `nfs-export.hubris.network` → `192.168.8.200` (direct, no Caddy)\n- **Privilege:**\
|
||
\ privileged (`unprivileged: 0`) + `lxc.apparmor.profile: unconfined` — required for `nfs-kernel-server`\n- **Resources:**\
|
||
\ 1 core / 512 MiB RAM / 2 GiB rootfs / 256 MiB swap\n- **Mounts:** host `/mnt/library` ↔ container `/mnt/library` (same\
|
||
\ path on both sides — matches the bind-mount convention used by jellyfin, paperless, arriman, nextcloud, mule-images,\
|
||
\ apps)\n\n## What it does\n\n```\n/mnt/library (host ext4 on nvme1n1)\n │\n ├── bind-mounted into 7 other\
|
||
\ LXCs (jellyfin, paperless, …)\n └── bind-mounted into LXC 102\n │\n └──\
|
||
\ nfs-kernel-server exports /mnt/library\n │\n └── consumed by VM 100\
|
||
\ (ZimaOS)\n```\n\nSame inodes, same page cache. The NFS server is just one more access path on top of a tree that 8 other\
|
||
\ consumers already share — see [media permissions](../infrastructure/media-permissions.md).\n\n## Export config\n\n`/etc/exports`:\n\
|
||
\n```\n/mnt/library 192.168.8.0/24(rw,all_squash,anonuid=33,anongid=10000,no_subtree_check,sec=sys)\n```\n\nInitially\
|
||
\ started as `ro`; promoted to `rw` on 2026-05-14 after the Files-UI evaluation confirmed (a) the library renders correctly\
|
||
\ as a folder under `/DATA`, (b) thumbnails are generated, (c) the squash works — a write from ZimaOS appears on `/mnt/library`\
|
||
\ as `www-data:media` (uid 33, gid 10000), matching the existing tree convention used by Nextcloud and mule-images.\n\n\
|
||
### Guardrails (in order of importance)\n\n1. **`all_squash,anonuid=33,anongid=10000`.** Every write from ZimaOS records\
|
||
\ on disk as `www-data:media` (uid 33, gid 10000), the same identity Nextcloud and mule-images use. Keeps the existing\
|
||
\ tree convention from drifting. See [`project_media_perms`](../infrastructure/media-permissions.md).\n3. **Subnet restriction\
|
||
\ `192.168.8.0/24`.** No public/mesh access; LAN only.\n4. **`no_subtree_check`** + **`sec=sys`** — standard performance/auth\
|
||
\ pair for a homelab.\n5. **No `crossmnt`** because `/mnt/library` is a single ext4 filesystem on the host (no nested\
|
||
\ mounts to traverse).\n\n### What we're *not* doing yet\n\n- No per-subdir export. ZimaOS sees the whole tree; access\
|
||
\ is controlled by filesystem permissions (`drwxr-x---` private dirs like `documents/`, `notes/`, `heaper/` aren't readable\
|
||
\ by `media` group, so ZimaOS-as-squashed-uid won't see them either).\n- No write-back. Until we promote to `rw`, ZimaOS\
|
||
\ can't write — eliminates the lock-domain split concern between NFS clients (NLM/v4) and local LXCs (POSIX flock).\n\
|
||
- No Authentik / forward-auth. NFS doesn't sit behind HTTP, so the standard caddy+authentik path doesn't apply. Subnet\
|
||
\ ACL is the only auth.\n\n## Why this LXC exists (vs serving from host)\n\nWe considered three options before building\
|
||
\ this:\n\n| Option | Outcome |\n|---|---|\n| **NFS on hubris bare-metal host** | Best performance, but adds long-lived\
|
||
\ NFS/RPC daemons to a host with a recent crash episode ([hubris crash 2026-04-21/22](../../sources/investigations/index.md)).\
|
||
\ Rejected. |\n| **SMB on host** | Same host-blast-radius problem, plus 30–50% lower throughput than NFS on Linux↔Linux.\
|
||
\ Rejected. |\n| **NFS in a dedicated LXC** ← this | Within ~2% of host performance (LXC is namespace isolation; IO path\
|
||
\ is unchanged), zero new daemons on hubris, matches the existing fleet pattern. Selected. |\n\nRationale lives in [the\
|
||
\ install plan](https://git.hubris.network/dtoro/oikos) `/root/.claude/plans/i-wannt-you-to-nifty-muffin.md` on\
|
||
\ hubris.\n\n## Operations\n\n- **Reload exports** after editing `/etc/exports`: `pct exec 102 -- exportfs -ra`\n- **List\
|
||
\ active exports:** `pct exec 102 -- exportfs -v`\n- **Watch from outside:** `showmount -e 192.168.8.200`\n- **Service\
|
||
\ health:** `pct exec 102 -- systemctl is-active nfs-server rpcbind`\n- **Restart cleanly:** `pct restart 102` (ZimaOS\
|
||
\ will retry the mount via `nofail`)\n- **Destroy + rebuild:** `pct stop 102 && pct destroy 102 --purge` — reversible\
|
||
\ in seconds; only ZimaOS notices\n\n## Open items\n- Consider tightening the export to subdirs (e.g. `movies`, `tv`,\
|
||
\ `music`, `audiobooks`, `books`, `images`, `podcasts`, `roms`) if you don't want ZimaOS reachable into `documents/`,\
|
||
\ `notes/`, `heaper/`, etc. — though those private subdirs are already invisible to `media`-group perms.\n- ZimaOS architecture\
|
||
\ finding: the Drives panel only enumerates physical/block devices via `GET /v2/local_storage/storages` (read-only API,\
|
||
\ no POST). Network shares **cannot** appear as Drives — they show up as folders in Files. This is intentional in CasaOS's\
|
||
\ design; don't try to work around it. Library-as-folder is the supported model.\n- Consider adding Samba to this LXC\
|
||
\ if a future Mac/iOS client needs SMB on the same tree — same LXC, no host changes.\n- No PBS backup (no PBS configured\
|
||
\ on hubris); the container is fully described in this page + `pct config 102`, rebuild from scratch in <2 min if lost.\n\
|
||
\n## Related\n- [100-zimaos](../vms/100-zimaos.md) — the only consumer today\n- [media permissions](../infrastructure/media-permissions.md)\
|
||
\ — uid 33 / gid 10000 standard\n- [DNS](../infrastructure/dns.md) — `nfs-export.hubris.network` entry (direct, no Caddy)\n\
|
||
\n## Changelog\n\n### 2026-05-14 — Promoted to rw; squash behaviour verified\nAfter ZimaOS Files UI evaluation passed\
|
||
\ (library renders as folder under `/DATA`, thumbnails work, ZimaOS Drives panel ignores NFS by design), flipped export\
|
||
\ to `rw`. Tested: writing `/DATA/library/.zimaos-rw-test` from ZimaOS appears on hubris's `/mnt/library` owned `www-data:media`\
|
||
\ (uid 33, gid 10000), confirming `all_squash,anonuid=33,anongid=10000` works as designed. Also discovered the dead end:\
|
||
\ ZimaOS's `GET /v2/local_storage/storages` is the source of the Drives panel; it returns only physical storage and rejects\
|
||
\ POST/PUT — network shares cannot be promoted to Drives.\n\n### 2026-05-14 — LXC built; NFS export live (read-only)\n\
|
||
Privileged Debian 13 container created with bind-mount `/mnt/library`. `nfs-kernel-server` installed and enabled; export\
|
||
\ `/mnt/library` to `192.168.8.0/24` with `ro,all_squash,anonuid=33,anongid=10000`. Smoke-tested from hubris host (mounted,\
|
||
\ listed library tree, confirmed RO). `nfs-export.hubris.network` added to LXC 124 dnsmasq.\n"
|
||
entity_slug: lxc:nfs-export
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: nfs-export
|
||
ip: 192.168.8.200
|
||
lan_dns: nfs-export.hubris.network → 192.168.8.200
|
||
privileged: 'privileged (unprivileged: 0) + lxc.apparmor.profile: unconfined — required for nfs-kernel-server'
|
||
resources: 1 core / 512 MiB RAM / 2 GiB rootfs / 256 MiB swap
|
||
mounts: host /mnt/library container /mnt/library
|
||
changelog:
|
||
- date: '2026-05-14'
|
||
title: Promoted to rw; squash behaviour verified
|
||
body: 'After ZimaOS Files UI evaluation passed (library renders as folder under `/DATA`, thumbnails work, ZimaOS Drives
|
||
panel ignores NFS by design), flipped export to `rw`. Tested: writing `/DATA/library/.zimaos-rw-test` from ZimaOS appears
|
||
on hubris''s `/mnt/library` owned `www-data:media` (uid 33, gid 10000), confirming `all_squash,anonuid=33,anongid=10000`
|
||
works as designed. Also discovered the dead end: ZimaOS''s `GET /v2/local_storage/storages` is the source of the Drives
|
||
panel; it returns only physical storage and rejects POST/PUT — network shares cannot be promoted to Drives.'
|
||
- date: '2026-05-14'
|
||
title: LXC built; NFS export live (read-only)
|
||
body: Privileged Debian 13 container created with bind-mount `/mnt/library`. `nfs-kernel-server` installed and enabled;
|
||
export `/mnt/library` to `192.168.8.0/24` with `ro,all_squash,anonuid=33,anongid=10000`. Smoke-tested from hubris host
|
||
(mounted, listed library tree, confirmed RO). `nfs-export.hubris.network` added to LXC 124 dnsmasq.
|
||
- slug: containers/103-paperless
|
||
title: 103 — `paperless`
|
||
content: '# 103 — `paperless`
|
||
|
||
|
||
Paperless-ngx for document management. Ingests scans / PDFs from `/mnt/library/documents` and serves OCR''d indexed copies.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `paperless`
|
||
|
||
- **IP:** `192.168.8.130`
|
||
|
||
- **Privilege:** privileged
|
||
|
||
- **Resources:** 2 cores / 3 GiB RAM / 8 GiB rootfs
|
||
|
||
- **Mounts:** `/mnt/library` ↔ `/mnt/library` (consumes `documents/`)
|
||
|
||
- **Public hostname:** [`paperless.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md) → `:8000`
|
||
|
||
|
||
## Service / port map
|
||
|
||
|
||
| Service | Listen | Notes |
|
||
|
||
| ------------------- | ------- | -------------------------------- |
|
||
|
||
| paperless-webserver | `:8000` | HTTP |
|
||
|
||
| paperless-task-queue, paperless-scheduler, paperless-consumer | — | systemd workers |
|
||
|
||
|
||
## Auth
|
||
|
||
Behind [Authentik forward-auth](106-auth-outpost.md). API path `/api/*` bypasses forward-auth (mobile clients can''t follow
|
||
the browser login redirect; bearer token still enforces auth on `/api`). Header propagation: `PAPERLESS_ENABLE_HTTP_REMOTE_USER=true`
|
||
and `PAPERLESS_HTTP_REMOTE_USER_HEADER_NAME=HTTP_X_AUTHENTIK_USERNAME` in `/opt/paperless/paperless.conf`. Django auto-creates
|
||
matching users on first SSO login.
|
||
|
||
|
||
## Storage
|
||
|
||
- Documents at `/mnt/library/documents` (owner `www-data:www-data`, mode 750 — *not* on the `media` group, by design).
|
||
|
||
|
||
## Known issues
|
||
|
||
- ~~Disk usage was 86.9% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
|
||
|
||
|
||
## Related
|
||
|
||
- [Authentik](106-auth-outpost.md)
|
||
|
||
- [Caddy](121-caddy.md)
|
||
|
||
- [DNS](../infrastructure/dns.md)
|
||
|
||
- [Monitoring](../infrastructure/monitoring.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Initial documentation.
|
||
|
||
|
||
### 2026-04-21 — wired into Authentik forward-auth
|
||
|
||
Cookie-domain-level proxy provider, embedded outpost binding. `/api/*` bypass added in Caddy site block to keep mobile
|
||
clients working.
|
||
|
||
'
|
||
entity_slug: lxc:paperless
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: paperless
|
||
ip: 192.168.8.130
|
||
privileged: privileged
|
||
resources: 2 cores / 3 GiB RAM / 8 GiB rootfs
|
||
mounts: /mnt/library /mnt/library
|
||
public_host: paperless.hubris.network → caddy → :8000
|
||
changelog:
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-21'
|
||
title: wired into Authentik forward-auth
|
||
body: Cookie-domain-level proxy provider, embedded outpost binding. `/api/*` bypass added in Caddy site block to keep
|
||
mobile clients working.
|
||
- slug: containers/104-gitea
|
||
title: 104 — `gitea`
|
||
content: "# 104 — `gitea`\n\nSelf-hosted git server. Source of truth for all `dtoro/*` repositories that auto-deploy across\
|
||
\ the lab.\n\n## At a glance\n- **Hostname:** `gitea`\n- **IP:** `192.168.8.121`\n- **Privilege:** privileged\n- **Resources:**\
|
||
\ 1 core / 1 GiB RAM / 8 GiB rootfs\n- **Mounts:** `/mnt/library` ↔ `/mnt/library` (under `/mnt/library/repos`, owner\
|
||
\ UID 102 / GID 105 — don't touch from other containers)\n- **Public hostname:** [`git.hubris.network`](../infrastructure/dns.md)\
|
||
\ → [caddy](121-caddy.md) → `:3000`\n- **Version:** 1.25.4 (last verified)\n\n## Service / port map\n\n| Service | Listen\
|
||
\ | Notes |\n| ------- | -------- | ------------------------------ |\n| gitea | `:3000` \
|
||
\ | HTTP web + git over HTTP |\n| gitea SSH | `:22` | git over SSH |\n| `gitea-deploy-webhook`\
|
||
\ | `127.0.0.1:9797` | Receives push events for the customizations repo |\n\n## Customizations\n`/var/lib/gitea/custom`\
|
||
\ is a checkout of [`dtoro/gitea-customizations`](https://git.hubris.network/dtoro/gitea-customizations).\n- `templates/custom/footer.tmpl`\
|
||
\ injects:\n - **Online3DViewer** v0.18.0 (vendored under `public/assets/o3dv/`) for inline 3D preview of CAD raw-file\
|
||
\ links: 3dm, 3ds, 3mf, amf, bim, brep, dae, fbx, fcstd, glb, gltf, ifc, igs, iges, stp, step, stl, obj, off, ply, wrl.\n\
|
||
\ - **PlantUML client JS** (vanilla, vendored from `gitea.com/davidsvantesson/plantuml-code-highlight`). Browser calls\
|
||
\ `https://git.hubris.network/_plantuml/img/<encoded>` — caddy `handle_path /_plantuml/*` rewrites and proxies to the\
|
||
\ PlantUML server on [apps (105)](105-apps.md) `:8079`.\n- `home.tmpl` is the anonymous landing page (Homesauce hero +\
|
||
\ sign-in/explore buttons).\n- O3DV binary assets are *not* tracked — `deploy.sh` fetches them on first run.\n\n## Auto-deploy\n\
|
||
Push to `dtoro/gitea-customizations` master → gitea webhook → `http://127.0.0.1:9797/deploy` (loopback) → `gitea-deploy-webhook.service`\
|
||
\ validates HMAC against `/etc/gitea-deploy/secret` → runs `/opt/gitea-customizations/deploy.sh`. Then `systemctl restart\
|
||
\ gitea` if templates changed.\n\n`app.ini` `[webhook] ALLOWED_HOST_LIST` includes `127.0.0.1`, `192.168.8.175`, `192.168.8.205`,\
|
||
\ `192.168.8.230`, `192.168.8.136`, `192.168.8.77` — every webhook target across the lab. **Don't strip these when editing\
|
||
\ app.ini.** See [auto-deploy](../infrastructure/auto-deploy.md) for the full list.\n\nGit creds for the deploy service:\
|
||
\ `/etc/gitea-deploy/git-credentials` (mode 600), referenced via `credential.helper = store --file=/etc/gitea-deploy/git-credentials`\
|
||
\ in the repo's `.git/config` (because `ProtectHome=true` blocks `/root`).\n\n## DNS workaround\nLXC has `/etc/hosts`\
|
||
\ override mapping `auth.hubris.network → 192.168.8.175` (router DNS resolves the public IONOS A record, which is a dead\
|
||
\ end for LAN clients — see [DNS](../infrastructure/dns.md)). Belt-and-suspenders systemd unit `hubris-hosts-override.service`\
|
||
\ re-applies it idempotently. Override line is **outside** the `# --- BEGIN PVE ---` markers (Proxmox rewrites inside).\n\
|
||
\n## Related\n- [Caddy](121-caddy.md)\n- [DNS](../infrastructure/dns.md)\n- [Auto-deploy pipelines](../infrastructure/auto-deploy.md)\n\
|
||
- [PlantUML server lives on apps (105)](105-apps.md)\n\n## Changelog\n\n### 2026-06-24 — terminalito deploy webhook (id\
|
||
\ 12)\nPush webhook on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy` ([trmnl (128)](128-trmnl.md)); `app.ini`\
|
||
\ `ALLOWED_HOST_LIST` extended with `192.168.8.211`. See [auto-deploy](../infrastructure/auto-deploy.md).\n\n### 2026-04-28\
|
||
\ — wiki entry created\nInitial documentation.\n\n### 2026-04-22 — `app.ini` ALLOWED_HOST_LIST extended for Artifacto\
|
||
\ webhook\nAdded `192.168.8.205`. See [Artifacto auto-deploy on apps (105)](105-apps.md).\n\n### 2026-04-21 — `/etc/hosts`\
|
||
\ override for `auth.hubris.network` added\nFor OIDC integration with [authentik (124)](106-auth-outpost.md). Outside\
|
||
\ the PVE markers, with a hubris-hosts-override.service for idempotency.\n\n### 2026-04-20 — gitea customizations + auto-deploy\
|
||
\ pipeline shipped\n`dtoro/gitea-customizations` repo created; webhook receiver at loopback `:9797` validates HMAC and\
|
||
\ runs `deploy.sh`. CAD and PlantUML loaders live in `footer.tmpl`.\n"
|
||
entity_slug: lxc:gitea
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: gitea
|
||
ip: 192.168.8.121
|
||
privileged: privileged
|
||
resources: 1 core / 1 GiB RAM / 8 GiB rootfs
|
||
mounts: /mnt/library /mnt/library
|
||
public_host: git.hubris.network → caddy → :3000
|
||
version: 1.25.4
|
||
changelog:
|
||
- date: '2026-06-24'
|
||
title: terminalito deploy webhook (id 12)
|
||
body: Push webhook on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy` ([trmnl (128)](128-trmnl.md)); `app.ini`
|
||
`ALLOWED_HOST_LIST` extended with `192.168.8.211`. See [auto-deploy](../infrastructure/auto-deploy.md).
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-22'
|
||
title: '`app.ini` ALLOWED_HOST_LIST extended for Artifacto webhook'
|
||
body: Added `192.168.8.205`. See [Artifacto auto-deploy on apps (105)](105-apps.md).
|
||
- date: '2026-04-21'
|
||
title: '`/etc/hosts` override for `auth.hubris.network` added'
|
||
body: For OIDC integration with [authentik (124)](106-auth-outpost.md). Outside the PVE markers, with a hubris-hosts-override.service
|
||
for idempotency.
|
||
- date: '2026-04-20'
|
||
title: gitea customizations + auto-deploy pipeline shipped
|
||
body: '`dtoro/gitea-customizations` repo created; webhook receiver at loopback `:9797` validates HMAC and runs `deploy.sh`.
|
||
CAD and PlantUML loaders live in `footer.tmpl`.'
|
||
- slug: containers/105-apps
|
||
title: 105 — `apps`
|
||
content: "# 105 — `apps`\n\nDocker host for everything that doesn't justify its own LXC. Currently runs Artifacto, PlantUML\
|
||
\ server, Portainer (and historically WriteFreely / blog), plus the [homelab-context distribution services](../infrastructure/homelab-context.md)\
|
||
\ (MCP + secrets-issuance) since 2026-05-20. Booklore migrated to [grimmory (130)](130-grimmory.md) on 2026-06-29.\n\n\
|
||
## At a glance\n- **Hostname:** `apps`\n- **IP:** `192.168.8.205`\n- **Privilege:** privileged (so container UID = host\
|
||
\ UID for bind-mount sanity)\n- **Resources:** 2 cores / 4 GiB RAM / 30 GiB rootfs\n- **Mounts:** `/mnt/library` ↔ `/mnt/library`\n\
|
||
- **Public hostnames:** see table below\n\n## Public hostname → service map\n\n| Hostname | Container\
|
||
\ | Backend port | Notes |\n| --------------------------------- | ---------------- | ------------ | ----- |\n|\
|
||
\ `docker.hubris.network` | Portainer | `:9443` | Native OAuth2 via Authentik. Trusted-origins requires\
|
||
\ hostname only (no scheme/port). |\n| `artifacto.hubris.network` | Artifacto | `:3100` | Public `/p/*`,\
|
||
\ `/static/*`, `/healthz` exposed via [VPS traefik](../infrastructure/ingress.md). |\n| `blog.hubris.network` \
|
||
\ | WriteFreely | `:8080` | Native OIDC via `[oauth.generic]`. |\n| `git.hubris.network/_plantuml/*` |\
|
||
\ PlantUML server | `:8079` | Same-origin route from [gitea (104)](104-gitea.md). |\n| `mcp.hubris.network` \
|
||
\ | homelab-mcp | `:9810` | MCP server over SSE. No SSO — agent-facing, gated at source-IP layer. See\
|
||
\ [homelab-context](../infrastructure/homelab-context.md). |\n| `secrets.hubris.network` | secrets-issuance |\
|
||
\ `:9820` | Per-client age-key issuance. Mesh+LAN source-IP gated (`MESH_SUBNETS`). |\n\n## Stacks\n\n### Artifacto\
|
||
\ (`/opt/artifacto/`)\nSelf-hosted HTML-artifact share (Go + SQLite). Repo `dtoro/Artifacto` is the working tree itself\
|
||
\ — `git push origin main` triggers rebuild.\n- Compose: tracked `docker-compose.yml` binds `./data:/data` and `127.0.0.1:3100:3000`.\
|
||
\ `docker-compose.override.yml` uses the `!override` YAML tag on `volumes:` and `ports:` to redirect `data` to `/mnt/library/documents/artifacts`.\n\
|
||
- Data dir: `/mnt/library/documents/artifacts` — owner **UID 10001 : GID 10001** (matches the `app` user baked into the\
|
||
\ Dockerfile), mode 750. Parent `/mnt/library/documents` stays `www-data:www-data 750`; bind targets the child directly\
|
||
\ so no traversal needed.\n- `.env` (untracked) holds `ADMIN_PASSWORD`, `SESSION_SECRET` (`openssl rand -hex 32`), `BASE_URL=https://artifacto.hubris.network`,\
|
||
\ `SSO_GATEWAY_SECRET` (must match `ARTIFACTO_GATEWAY_SECRET` in `/etc/caddy/caddy.env` on [caddy](121-caddy.md)). Don't\
|
||
\ `git checkout .env`.\n- Auth: Authentik forward-auth on the admin surface (`/`, `/login`, `/a/*`, `/api/*`). `X-Authentik-Username`\
|
||
\ is honored only when the `X-Artifacto-Gateway` header matches `SSO_GATEWAY_SECRET` — prevents direct LAN connections\
|
||
\ from spoofing identity.\n- Public exposure: VPS traefik gates path scope to share URLs; admin surface returns 404 at\
|
||
\ the edge. See [ingress](../infrastructure/ingress.md).\n\n#### Auto-deploy pipeline (Artifacto)\nReceiver at `/opt/artifacto-deploy/`\
|
||
\ (outside the app repo): `deploy.sh` + `webhook.py`, secret in `/etc/artifacto-deploy/secret`, systemd unit `artifacto-deploy-webhook.service`.\
|
||
\ Listens on `0.0.0.0:9798`, validates `X-Gitea-Signature` HMAC-SHA256, filters to `refs/heads/main`, runs deploy in a\
|
||
\ daemon thread (returns 202 immediately). Gitea webhook id 7 on `dtoro/Artifacto`.\n\n- Logs: `pct exec 105 -- journalctl\
|
||
\ -u artifacto-deploy-webhook -f`\n- Manual deploy: `pct exec 105 -- /opt/artifacto-deploy/deploy.sh`\n- Health: `pct\
|
||
\ exec 105 -- curl -s http://127.0.0.1:9798/health` → `ok`\n- Slug length bumped 5 → 8 chars (in `internal/store/artifacts.go:41`)\
|
||
\ once the endpoint went public.\n\n### PlantUML (`/opt/plantuml/`)\n`plantuml/plantuml-server:jetty` listening on `192.168.8.205:8079`,\
|
||
\ context path `/plantuml` (`BASE_URL=plantuml`). Reached only via [gitea](104-gitea.md)'s caddy `handle_path /_plantuml/*`\
|
||
\ rewrite — no DNS record of its own. Same-origin from gitea so the embedded markdown previewer can call it without CORS.\n\
|
||
\n### Portainer\nNative OAuth2 (Settings → Authentication → OAuth → Custom). Manual endpoints (no OIDC discovery). Uses\
|
||
\ `portainer-uid` custom-claim scope from Authentik. Container is **not** compose-managed — safe to `docker run` recreate;\
|
||
\ data lives in named volume `portainer_data`. CLI flag: `--trusted-origins docker.hubris.network` (hostname only — `IsTrustedOrigin`\
|
||
\ rejects strings containing `://`).\n\n### homelab-mcp (`/opt/homelab-mcp/`) — DEPRECATED (Go rewrite, Phase 6)\n\n>\
|
||
\ **Status:** This Python MCP server is being replaced by the Go `oikos api` binary\n> running in Docker on mac-mini.\
|
||
\ Cutover pending — see\n> [scripts/cutover-checklist.md](../../scripts/cutover-checklist.md) for the\n> execution plan.\
|
||
\ The Go MCP uses the official MCP Go SDK (Streamable HTTP, not\n> FastMCP) with 15 tools including `get_blast_radius`,\
|
||
\ `request_execution`, and\n> `get_agent_activity`. Source: `internal/mcp/server.go`.\n\n**Current (Python) implementation\
|
||
\ — DO NOT MODIFY, awaiting cutover:**\nFastMCP server (Python venv at `/opt/homelab-mcp/.venv`). Reads from\n`/opt/homelab-context/`\
|
||
\ (this LXC is itself an enrolled\n[homelab-context](../infrastructure/homelab-context.md) client). Source\nis `dtoro/oikos/mcp/server.py`;\
|
||
\ service unit\n`homelab-mcp.service` binds `0.0.0.0:9810`. DNS-rebinding protection\ndisabled at the FastMCP layer because\
|
||
\ mesh+LAN gating is the actual\ntrust boundary.\n\n- Endpoint: `https://mcp.hubris.network/mcp` (Caddy → `:9810`). StreamableHTTP\
|
||
\ transport (POST `/mcp`).\n- 14 tools registered: `get_host`, `list_services`, `find_service`,\n `get_topology`, `search_docs`,\
|
||
\ `get_page`, `get_changelog`, `whoami`,\n `list_my_secrets` (context); `get_service_status`, `tail_log`,\n `list_lxcs`,\
|
||
\ `get_lxc_state`, `ping_service` (read-only management).\n- Management tools shell out via SSH using a dedicated `mcp-reader`\
|
||
\ key\n (planned; not provisioned yet — management tools fail until then).\n\n#### Auto-deploy pipeline (homelab-mcp)\n\
|
||
Receiver at `/opt/homelab-mcp/mcp/deploy/webhook/` (in-tree alongside\nthe server code). Listens on `0.0.0.0:9811`, secret\
|
||
\ in\n`/etc/homelab-mcp-deploy/secret`, systemd unit\n`homelab-mcp-deploy.service`. Gitea webhook id 10 on\n`dtoro/oikos`.\
|
||
\ Deploy script reinstalls the service unit and\nrestarts on push.\n\n### secrets-issuance (`/opt/secrets-issuance/`)\
|
||
\ — DEPRECATED (Go Phase 5)\n\n> **Status:** Replaced by `internal/secrets/` in the Go rewrite. Machine identities\n>\
|
||
\ are now managed via Infisical (`docker compose --profile infisical up`).\nTiny HTTP service that issues per-client age\
|
||
\ keypairs the first time\neach client calls `/issue`. Idempotent: subsequent calls return the\nsame key. Mesh+LAN source-IP\
|
||
\ gated via the `MESH_SUBNETS` env in\n`secrets-issuance.service`. State (per-client keys + denylist) lives at\n`/var/lib/secrets-issuance/`\
|
||
\ (NOT in the repo).\n\n- Endpoint: `https://secrets.hubris.network/issue` (Caddy → `:9820`).\n- `/revoke` is admin-token-gated\
|
||
\ by `/etc/secrets-issuance/admin-token`;\n shreds the local key file and adds the hostname to the denylist.\n Called\
|
||
\ by `homelab client remove`.\n- Trust subnets today: `100.122.0.0/16` (Netbird), `100.64.0.0/10`\n (Tailscale), `192.168.8.0/24`\
|
||
\ (LAN). Tighten if the LAN gets\n untrusted devices.\n\n#### Auto-deploy pipeline (secrets-issuance)\nReceiver at `/opt/secrets-issuance/secrets-issuance/deploy/webhook/`,\n\
|
||
port `9821`, secret in `/etc/secrets-issuance-deploy/secret`, unit\n`secrets-issuance-deploy.service`. Gitea webhook id\
|
||
\ 11.\n\n### WriteFreely (`/opt/writefreely-docker/`)\nNative OIDC via `[oauth.generic]` in `config/config.ini`. `host\
|
||
\ = https://auth.hubris.network`, scope `openid profile email`, `map_user_id = sub`. Redirect URI in Authentik: `https://blog.hubris.network/oauth/callback/generic`.\
|
||
\ Container has `extra_hosts: auth.hubris.network:192.168.8.175`. First-time users with a local WF account must connect\
|
||
\ SSO via Customize/Settings → Linked Accounts.\n\n## Permissions\n- LXC is privileged → in-container UID == host UID.\n\
|
||
- Member of [media GID 10000](../infrastructure/media-permissions.md). Adding `media` to the LXC alone is *not* enough\
|
||
\ for inner Docker containers — each container needs `--group-add 10000` or `user: \"<uid>:10000\"` to write into `/mnt/library`.\n\
|
||
\n## Related\n- [Gitea (104)](104-gitea.md) — uses the PlantUML server\n- [Caddy (121)](121-caddy.md)\n- [Authentik (124)](106-auth-outpost.md)\n\
|
||
- [DNS](../infrastructure/dns.md)\n- [Auto-deploy](../infrastructure/auto-deploy.md)\n- [Public ingress (Artifacto + blog)](../infrastructure/ingress.md)\n\
|
||
\n## Changelog\n\n### 2026-06-29 — Booklore migrated to Grimmory on LXC 130\nBooklore stack removed from Portainer. MariaDB\
|
||
\ dump taken first, then restored into [grimmory (130)](130-grimmory.md)'s fresh MariaDB. `books.hubris.network` Caddy\
|
||
\ backend updated to `192.168.8.213:6060`. Authentik OIDC provider updated to Public client type (PKCE) for Grimmory compatibility.\n\
|
||
\n### 2026-05-20 — homelab-mcp + secrets-issuance live\nTwo new services from the [homelab-context distribution plan](../infrastructure/homelab-context.md):\n\
|
||
`homelab-mcp.service` on `:9810` (MCP read+management surface) and\n`secrets-issuance.service` on `:9820` (per-client\
|
||
\ age-key provisioning).\nCaddy fronts both with Let's Encrypt; new vhosts on\n[caddy](121-caddy.md), split-horizon DNS\
|
||
\ entries on\n[authentik (124)](106-auth-outpost.md). Gitea webhook ids 10 + 11 wire\nauto-deploy. LXC is itself an enrolled\
|
||
\ context client\n(`/opt/homelab-context/`).\n\n### 2026-04-28 — wiki entry created\nInitial documentation.\n\n### 2026-04-22\
|
||
\ — Artifacto deployed + Authentik gateway-secret auto-login wired\nRepo `dtoro/Artifacto`. Compose, data dir, `.env`,\
|
||
\ webhook (id 7), Caddyfile site block, gateway-secret pattern, public-internet share exposure on the VPS — all set up\
|
||
\ that day. Slug length bumped 5 → 8 for brute-force resistance.\n\n### 2026-04-22 — Booklore mariadb data wiped (lesson,\
|
||
\ not a feature)\n`docker compose up -d` from the host shell on a Portainer-managed stack recreated services and re-initialized\
|
||
\ the bind-mounted `./mariadb/config`. Recovered by reseeding from snapshot. Rule recorded: edit Portainer-managed stacks\
|
||
\ via the Portainer UI.\n\n### 2026-04-22 — Portainer wired into Authentik (native OAuth2)\nCustom-claim `portainer_uid`\
|
||
\ scope. CLI added `--trusted-origins docker.hubris.network`.\n\n### 2026-04-22 — `app.ini` host whitelist on gitea extended\n\
|
||
Added `192.168.8.205` so the Artifacto webhook delivery wouldn't be rejected. See [gitea (104)](104-gitea.md).\n\n###\
|
||
\ 2026-04-21 — Booklore + WriteFreely wired into Authentik OIDC\nNative OIDC for both. `extra_hosts` override added inside\
|
||
\ each container so DNS resolves `auth.hubris.network` to the Caddy IP.\n\n### 2026-04-20 — PlantUML server deployed\n\
|
||
Used by [gitea (104)](104-gitea.md)'s footer.tmpl injection. Same-origin via caddy `handle_path /_plantuml/*` rewrite\
|
||
\ — no separate hostname.\n"
|
||
entity_slug: lxc:apps
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: apps
|
||
ip: 192.168.8.205
|
||
privileged: privileged
|
||
resources: 2 cores / 4 GiB RAM / 30 GiB rootfs
|
||
mounts: /mnt/library /mnt/library
|
||
public_hostnames: see table below
|
||
changelog:
|
||
- date: '2026-06-29'
|
||
title: Booklore migrated to Grimmory on LXC 130
|
||
body: Booklore stack removed from Portainer. MariaDB dump taken first, then restored into [grimmory (130)](130-grimmory.md)'s
|
||
fresh MariaDB. `books.hubris.network` Caddy backend updated to `192.168.8.213:6060`. Authentik OIDC provider updated
|
||
to Public client type (PKCE) for Grimmory compatibility.
|
||
- date: '2026-05-20'
|
||
title: homelab-mcp + secrets-issuance live
|
||
body: 'Two new services from the [homelab-context distribution plan](../infrastructure/homelab-context.md): `homelab-mcp.service`
|
||
on `:9810` (MCP read+management surface) and `secrets-issuance.service` on `:9820` (per-client age-key provisioning).
|
||
Caddy fronts both with Let''s Encrypt; new vhosts on [caddy](121-caddy.md), split-horizon DNS entries on [authentik
|
||
(124)](106-auth-outpost.md). Gitea webhook ids 10 + 11 wire auto-deploy. LXC is itself an enrolled context client (`/opt/homelab-context/`).'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-22'
|
||
title: Artifacto deployed + Authentik gateway-secret auto-login wired
|
||
body: Repo `dtoro/Artifacto`. Compose, data dir, `.env`, webhook (id 7), Caddyfile site block, gateway-secret pattern,
|
||
public-internet share exposure on the VPS — all set up that day. Slug length bumped 5 → 8 for brute-force resistance.
|
||
- date: '2026-04-22'
|
||
title: Booklore mariadb data wiped (lesson, not a feature)
|
||
body: '`docker compose up -d` from the host shell on a Portainer-managed stack recreated services and re-initialized the
|
||
bind-mounted `./mariadb/config`. Recovered by reseeding from snapshot. Rule recorded: edit Portainer-managed stacks
|
||
via the Portainer UI.'
|
||
- date: '2026-04-22'
|
||
title: Portainer wired into Authentik (native OAuth2)
|
||
body: Custom-claim `portainer_uid` scope. CLI added `--trusted-origins docker.hubris.network`.
|
||
- date: '2026-04-22'
|
||
title: '`app.ini` host whitelist on gitea extended'
|
||
body: Added `192.168.8.205` so the Artifacto webhook delivery wouldn't be rejected. See [gitea (104)](104-gitea.md).
|
||
- date: '2026-04-21'
|
||
title: Booklore + WriteFreely wired into Authentik OIDC
|
||
body: Native OIDC for both. `extra_hosts` override added inside each container so DNS resolves `auth.hubris.network` to
|
||
the Caddy IP.
|
||
- date: '2026-04-20'
|
||
title: PlantUML server deployed
|
||
body: Used by [gitea (104)](104-gitea.md)'s footer.tmpl injection. Same-origin via caddy `handle_path /_plantuml/*` rewrite
|
||
— no separate hostname.
|
||
- slug: containers/106-auth-outpost
|
||
title: 106 — `auth-outpost`
|
||
content: '# 106 — `auth-outpost`
|
||
|
||
|
||
Authentik **forward-auth outpost** for LAN-gated apps. A stateless proxy that connects outbound to the [VPS Authentik
|
||
core](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md) and serves forward-auth locally, so
|
||
[Caddy (121)](121-caddy.md) never hairpins auth through VPS Traefik.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `auth-outpost`
|
||
|
||
- **IP:** `192.168.8.6` (static)
|
||
|
||
- **Privilege:** privileged (Docker-in-LXC, `features: nesting=1`)
|
||
|
||
- **Resources:** 1 core / 512 MiB / 4 GiB rootfs
|
||
|
||
- **Mounts:** none
|
||
|
||
- **Created:** 2026-06-01, Debian 13, replacing the embedded outpost on [124](106-auth-outpost.md)
|
||
|
||
|
||
## Role
|
||
|
||
|
||
Runs one container — `ghcr.io/goauthentik/proxy` — that opens an outbound websocket to `https://auth.hubris.network` (the
|
||
VPS core), pulls its proxy-provider config, and answers Caddy''s `forward_auth` subrequests on `192.168.8.6:9000` (LAN-only
|
||
bind). Because the call path is **Caddy → outpost (LAN)**, with no Traefik in between, `X-Forwarded-Host` is preserved
|
||
— the failure that 404s when Caddy is pointed at `https://auth.hubris.network` directly (Traefik rewrites the header).
|
||
See the [migration investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md).
|
||
|
||
|
||
## Service / port map
|
||
|
||
| Service | Listen | Notes |
|
||
|
||
| --- | --- | --- |
|
||
|
||
| authentik proxy outpost | `192.168.8.6:9000` | `/outpost.goauthentik.io/*` (ping, auth/caddy) |
|
||
|
||
|
||
## Config paths
|
||
|
||
- `/opt/authentik-outpost/docker-compose.yml` — `ghcr.io/goauthentik/proxy:2026.5.2`, `AUTHENTIK_HOST=https://auth.hubris.network`,
|
||
`AUTHENTIK_INSECURE=false`, port bound `192.168.8.6:9000`.
|
||
|
||
- `/opt/authentik-outpost/.env` (mode 600, **untracked**) — `AUTHENTIK_TOKEN` for outpost `hubris-lan-outpost` (sops-encrypt
|
||
into `secrets/` — TODO Phase 5).
|
||
|
||
|
||
## The outpost object (VPS Authentik)
|
||
|
||
- Outpost **`hubris-lan-outpost`** (type proxy), providers: `hubris-forward-auth` (domain, Paperless + domain-level apps),
|
||
`Provider for Torrent` (qBittorrent), `Provider for Artifacto` (single).
|
||
|
||
- Health: VPS admin → Applications → Outposts → last-seen current; `curl http://192.168.8.6:9000/outpost.goauthentik.io/ping`
|
||
→ `204`; container log `Successfully connected websocket`.
|
||
|
||
|
||
## Caddy wiring
|
||
|
||
[Caddy (121)](121-caddy.md) `(authentik)` snippet `reverse_proxy`/`forward_auth` → `http://192.168.8.6:9000` (was `192.168.8.180:9000`
|
||
on LXC 124). Tracked in `dtoro/caddy-conf`.
|
||
|
||
|
||
## `sso.hubris.network` — the callback domain (critical)
|
||
|
||
Domain-level (`forward_domain`) providers redirect the browser to `{external_host}/outpost.goauthentik.io/callback` after
|
||
login. On LXC 124 that `external_host` was `auth.hubris.network`, which resolved (LAN) to Caddy → the 124 outpost — same
|
||
box, so it worked. Post-migration `auth.hubris.network` points to the **VPS core**, so the callback hit the wrong outpost
|
||
→ **400 / redirect-uri mismatch**.
|
||
|
||
|
||
Fix: the LAN outpost gets its **own** domain.
|
||
|
||
- DNS: `sso.hubris.network → 192.168.8.175` (Caddy). *(in dnsmasq today — must be carried into Technitium in DNS Phase
|
||
2.)*
|
||
|
||
- Caddy: site `sso.hubris.network { tls dns ionos; reverse_proxy 192.168.8.6:9000 }` (in `caddy-conf`).
|
||
|
||
- Authentik: `hubris-forward-auth` and `Provider for Torrent` have `external_host=https://sso.hubris.network`; their `redirect_uris`
|
||
must match (`set_oauth_defaults()` regenerates them from `external_host` — changing `external_host` alone does NOT update
|
||
them).
|
||
|
||
- `Provider for Artifacto` is `forward_single` on its own domain, so it was unaffected.
|
||
|
||
|
||
**Lesson:** when the IdP core and the forward-auth outpost live on different hosts, the outpost needs a dedicated domain
|
||
distinct from the core''s — and proxy-provider `redirect_uris` must be regenerated, not just `external_host`.
|
||
|
||
|
||
## Related
|
||
|
||
- [124 — authentik](106-auth-outpost.md) — old embedded-outpost host (now DNS-only)
|
||
|
||
- [Caddy (121)](121-caddy.md) — forward-auth consumer
|
||
|
||
- [Ingress (VPS traefik)](../infrastructure/ingress.md)
|
||
|
||
- [Authentik VPS migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-06 — Authentik session lifetime extended to 30 days
|
||
|
||
VPS Authentik core `user_login` stage updated: `session_duration` changed from `seconds=0` (session cookie, cleared on
|
||
browser close) to `days=30` (persistent 30-day cookie). Also set `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30` in
|
||
`/opt/authentik.env` on the VPS. See [investigation](../../sources/investigations/2026-06-06-authentik-session-lifetime.md).
|
||
|
||
|
||
### 2026-06-01 — created; forward-auth cut over from LXC 124
|
||
|
||
New dedicated LXC for the LAN forward-auth outpost (Phase 1 of the [architecture migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)).
|
||
Deployed `goauthentik/proxy:2026.5.2` pointed at the VPS core; repointed Caddy `(authentik)` from `192.168.8.180:9000`
|
||
→ `192.168.8.6:9000`. Verified Paperless/qBittorrent/Artifacto return the SSO redirect with **124-Authentik stopped**,
|
||
confirming the frozen instance is out of the path. dnsmasq stays on 124 until [DNS is relocated](106-auth-outpost.md).
|
||
|
||
'
|
||
entity_slug: lxc:auth-outpost
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: auth-outpost
|
||
ip: 192.168.8.6
|
||
privileged: privileged
|
||
resources: 1 core / 512 MiB / 4 GiB rootfs
|
||
mounts: none
|
||
created: 2026-06-01, Debian 13, replacing the embedded outpost on [124]
|
||
changelog:
|
||
- date: '2026-06-06'
|
||
title: Authentik session lifetime extended to 30 days
|
||
body: 'VPS Authentik core `user_login` stage updated: `session_duration` changed from `seconds=0` (session cookie, cleared
|
||
on browser close) to `days=30` (persistent 30-day cookie). Also set `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30`
|
||
in `/opt/authentik.env` on the VPS. See [investigation](../../sources/investigations/2026-06-06-authentik-session-lifetime.md).'
|
||
- date: '2026-06-01'
|
||
title: created; forward-auth cut over from LXC 124
|
||
body: New dedicated LXC for the LAN forward-auth outpost (Phase 1 of the [architecture migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)).
|
||
Deployed `goauthentik/proxy:2026.5.2` pointed at the VPS core; repointed Caddy `(authentik)` from `192.168.8.180:9000`
|
||
→ `192.168.8.6:9000`. Verified Paperless/qBittorrent/Artifacto return the SSO redirect with **124-Authentik stopped**,
|
||
confirming the frozen instance is out of the path. dnsmasq stays on 124 until [DNS is relocated](106-auth-outpost.md).
|
||
- slug: containers/107-dns
|
||
title: 107 — `dns`
|
||
content: '# 107 — `dns`
|
||
|
||
|
||
Homelab DNS server (Technitium). Replaces the dnsmasq that lived on [124 — authentik](106-auth-outpost.md); single-purpose,
|
||
one job.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `dns`
|
||
|
||
- **IP:** `192.168.8.2` (static — stable, decoupled from any app)
|
||
|
||
- **Privilege:** privileged (Docker-in-LXC, `features: nesting=1`)
|
||
|
||
- **Resources:** 1 core / 1 GiB / 8 GiB rootfs
|
||
|
||
- **Created:** 2026-06-01, Debian 13. Its own resolver is `1.1.1.1` (no circular dependency on the DNS it serves).
|
||
|
||
|
||
## Role
|
||
|
||
Authoritative split-horizon DNS for `hubris.network` on the LAN/mesh, plus recursive forwarding (`1.1.1.1`, `8.8.8.8`)
|
||
for everything else. Technitium runs in Docker (`technitium/dns-server:latest`, `network_mode: host`), web console on
|
||
`:5380`.
|
||
|
||
|
||
## The `hubris.network` zone
|
||
|
||
- Specific A overrides: app names → `192.168.8.175` (Caddy), `nfs-export → 192.168.8.200`, `auth/sso/...` as needed.
|
||
|
||
- `auth.hubris.network → 82.165.190.79` (VPS Authentik), `sso.hubris.network → 192.168.8.175` (LAN forward-auth outpost).
|
||
|
||
- **Wildcard** `*.hubris.network → 82.165.190.79` — mirrors the public IONOS wildcard so undefined names (e.g. `netbird`)
|
||
resolve to the VPS, matching public behaviour.
|
||
|
||
- **MX / SPF-TXT / CAA replicated from public** so an authoritative zone doesn''t shadow `hubris.network` email/cert records.
|
||
|
||
|
||
## Config / access
|
||
|
||
- `/opt/technitium/docker-compose.yml`; admin password in `/opt/technitium/admin_password.txt` (mode 600 — sops-encrypt
|
||
in Phase 5).
|
||
|
||
- Console: `http://192.168.8.2:5380` (user `admin`).
|
||
|
||
- API: `http://192.168.8.2:5380/api/...` (token via `/api/user/login`). Zone was built via the API.
|
||
|
||
|
||
## Who points here
|
||
|
||
- **NetBird mesh peers:** resolve via the **NetBird managed DNS zone**, kept in sync *from* this Technitium (see dns-sync
|
||
below). The `home-lab-dns` nameserver group (`→ 192.168.8.2`) is a thin fallback forwarder.
|
||
|
||
- **Homelab DHCP clients:** Technitium''s own DHCP scope hands out `192.168.8.2` as the DNS server for `192.168.8.x` leases
|
||
(see DHCP section below).
|
||
|
||
- **Plain LAN clients (`192.168.178.x`):** Fritz!Box DHCP still hands out Fritz!Box itself (`192.168.178.1`) as DNS —
|
||
no split-horizon for non-mesh clients. Changing this requires a secondary DNS fallback, which Fritz!OS 8.x doesn''t expose
|
||
in a single DHCP field.
|
||
|
||
|
||
## dns-sync (Technitium = authoring source)
|
||
|
||
`/opt/dns-sync/sync.py` (cron `*/10`, logs `/var/log/dns-sync.log`) reconciles this zone''s named A-records → the NetBird
|
||
managed DNS zone via the NetBird API (`/api/dns/zones/{id}/records`). Token at `/opt/dns-sync/netbird-token` (mode 600;
|
||
source of truth in sops `secrets/netbird-pat.yaml`). **Edit DNS only here**; the sync propagates to the mesh. It deletes
|
||
NetBird records absent from Technitium. Tracked: [scripts/dns-sync.py](../../../scripts/dns-sync.py). *Why this exists:*
|
||
NetBird won''t forward to Technitium for mesh peers (self-IP / nameserver-group quirks), so we sync into the managed zone
|
||
instead — see [dns.md](../infrastructure/dns.md).
|
||
|
||
|
||
## DHCP
|
||
|
||
|
||
Technitium also runs a DHCP server for the homelab subnet (enabled 2026-06-02):
|
||
|
||
- **Scope:** `homelab` — `192.168.8.241 – 192.168.8.254`
|
||
|
||
- **Gateway:** `192.168.8.1` (Proxmox `vmbr0` alias)
|
||
|
||
- **DNS:** `192.168.8.2` (self)
|
||
|
||
- **Lease time:** 24 h
|
||
|
||
|
||
Replaces the DHCP that was previously served by the Slate AX router. Static-IP LXCs (`.101–.239`) are excluded from the
|
||
pool. Pool narrowed from `.100–.240` to `.241–.254` on 2026-06-03 to eliminate IP conflict risk.
|
||
|
||
|
||
## Related
|
||
|
||
- [124 — authentik](106-auth-outpost.md) — retired host of the old dnsmasq
|
||
|
||
- [DNS split-horizon](../infrastructure/dns.md)
|
||
|
||
- [Mesh](../infrastructure/mesh.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-24 — A record `trmnl.hubris.network → 192.168.8.175`
|
||
|
||
Added for [trmnl (128)](128-trmnl.md) (LAN path via [Caddy (121)](121-caddy.md)); propagated to the NetBird managed zone
|
||
by `dns-sync`.
|
||
|
||
|
||
### 2026-06-06 — dns-sync cron installed (had been missing since deployment)
|
||
|
||
Although the 2026-06-03 changelog claimed "cron */10", **no crontab was actually configured** on the LXC. The sync was
|
||
running only via ad-hoc manual invocations during incident debugging. Fixed by adding `/etc/cron.d/dns-sync`.
|
||
|
||
|
||
### 2026-06-03 — DHCP pool narrowed to `.241–.254`
|
||
|
||
Previous pool `.100–.240` overlapped with all static LXCs/VMs (`.101–.239`). Shrunk via API (`/api/dhcp/scopes/set`).
|
||
11 stale DHCP leases in `.101–.110` remain until natural expiry (2026-06-04). See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).
|
||
|
||
|
||
### 2026-06-03 — dns-sync added (Technitium → NetBird managed zone)
|
||
|
||
This Technitium became the single DNS authoring source; `/opt/dns-sync/sync.py` (cron */10) reconciles named A-records
|
||
into the NetBird managed zone via the API. Fixed previously-broken mesh names (`sso`, `nfs-export`, `mcp`, `secrets`)
|
||
by adding them to the managed zone; reaped obsolete `files`/`photos-new`. See [dns.md](../infrastructure/dns.md).
|
||
|
||
|
||
### 2026-06-02 — DHCP server enabled; replaces Slate AX DHCP
|
||
|
||
Enabled Technitium''s built-in DHCP server for `192.168.8.0/24` (scope `homelab`, range `.100–.240`, gateway `192.168.8.1`,
|
||
DNS self). Previously the Slate AX sub-router served DHCP for the homelab subnet. With the Slate AX retired and Proxmox
|
||
now the subnet router, Technitium takes over DHCP. Configured via the Technitium API (`/api/dhcp/scopes/set`). DHCP LXCs
|
||
kept their Slate AX leases until expiry, then renewed from Technitium.
|
||
|
||
|
||
### 2026-06-01 — created; replaced dnsmasq on 124
|
||
|
||
Stood up Technitium at `192.168.8.2`, imported the split-horizon zone (specific A + wildcard + MX/SPF/CAA), made it the
|
||
primary nameserver in the NetBird `home-lab-dns` group. Verified all names resolve with dnsmasq/124 stopped; [LXC 124
|
||
retired](106-auth-outpost.md).
|
||
|
||
'
|
||
entity_slug: lxc:dns
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: dns
|
||
ip: 192.168.8.2
|
||
privileged: privileged
|
||
resources: 1 core / 1 GiB / 8 GiB rootfs
|
||
created: 2026-06-01, Debian 13. Its own resolver is 1.1.1.1 (no circular dependency on the DNS it serves).
|
||
changelog:
|
||
- date: '2026-06-24'
|
||
title: A record `trmnl.hubris.network → 192.168.8.175`
|
||
body: Added for [trmnl (128)](128-trmnl.md) (LAN path via [Caddy (121)](121-caddy.md)); propagated to the NetBird managed
|
||
zone by `dns-sync`.
|
||
- date: '2026-06-06'
|
||
title: dns-sync cron installed (had been missing since deployment)
|
||
body: Although the 2026-06-03 changelog claimed "cron */10", **no crontab was actually configured** on the LXC. The sync
|
||
was running only via ad-hoc manual invocations during incident debugging. Fixed by adding `/etc/cron.d/dns-sync`.
|
||
- date: '2026-06-03'
|
||
title: DHCP pool narrowed to `.241–.254`
|
||
body: Previous pool `.100–.240` overlapped with all static LXCs/VMs (`.101–.239`). Shrunk via API (`/api/dhcp/scopes/set`).
|
||
11 stale DHCP leases in `.101–.110` remain until natural expiry (2026-06-04). See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).
|
||
- date: '2026-06-03'
|
||
title: dns-sync added (Technitium → NetBird managed zone)
|
||
body: This Technitium became the single DNS authoring source; `/opt/dns-sync/sync.py` (cron */10) reconciles named A-records
|
||
into the NetBird managed zone via the API. Fixed previously-broken mesh names (`sso`, `nfs-export`, `mcp`, `secrets`)
|
||
by adding them to the managed zone; reaped obsolete `files`/`photos-new`. See [dns.md](../infrastructure/dns.md).
|
||
- date: '2026-06-02'
|
||
title: DHCP server enabled; replaces Slate AX DHCP
|
||
body: Enabled Technitium's built-in DHCP server for `192.168.8.0/24` (scope `homelab`, range `.100–.240`, gateway `192.168.8.1`,
|
||
DNS self). Previously the Slate AX sub-router served DHCP for the homelab subnet. With the Slate AX retired and Proxmox
|
||
now the subnet router, Technitium takes over DHCP. Configured via the Technitium API (`/api/dhcp/scopes/set`). DHCP
|
||
LXCs kept their Slate AX leases until expiry, then renewed from Technitium.
|
||
- date: '2026-06-01'
|
||
title: created; replaced dnsmasq on 124
|
||
body: Stood up Technitium at `192.168.8.2`, imported the split-horizon zone (specific A + wildcard + MX/SPF/CAA), made
|
||
it the primary nameserver in the NetBird `home-lab-dns` group. Verified all names resolve with dnsmasq/124 stopped;
|
||
[LXC 124 retired](106-auth-outpost.md).
|
||
- slug: containers/114-nextcloud
|
||
title: 114 — `nextcloud`
|
||
content: '# 114 — `nextcloud`
|
||
|
||
|
||
Personal cloud / file collaboration. Source-of-truth for the photo libraries surfaced by [mulita (120)](120-mule-images.md).
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `nextcloud`
|
||
|
||
- **IP:** `192.168.8.224`
|
||
|
||
- **Privilege:** privileged
|
||
|
||
- **Resources:** 4 cores / 6 GiB RAM / 25 GiB rootfs
|
||
|
||
- **Mounts:** `/mnt/library` ↔ `/mnt/library` (Nextcloud user files under `/mnt/library/homecloud`)
|
||
|
||
- **Public hostname:** [`cloud.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md)
|
||
|
||
|
||
## Auth
|
||
|
||
Native OIDC via `user_oidc` app. **Username override pattern**: Authentik user `dtoro` maps to local Nextcloud user `admin`
|
||
via the `nc_uid` custom-claim scope. Configured via `occ user_oidc:provider <name> --mapping-uid=nc_uid` and `--scope="openid
|
||
profile email <app>-uid"`. See [Authentik](106-auth-outpost.md#per-app-username-override-pattern-authentik) for the full
|
||
pattern.
|
||
|
||
|
||
Redirect URI: `/index.php/apps/user_oidc/code` (NOT `/apps/...` — pretty URLs aren''t on).
|
||
|
||
|
||
## DNS workaround (load-bearing)
|
||
|
||
Nextcloud''s PHP HTTP client (`OC\Http\Client\DnsPinMiddleware`) calls `dns_get_record()` directly — it bypasses `/etc/hosts`.
|
||
So the Authentik OIDC discovery URL needs **real DNS** answering with `192.168.8.175` for `auth.hubris.network`. Recipe
|
||
applied:
|
||
|
||
- Local `dnsmasq` listening on `127.0.0.1:53`, `/etc/dnsmasq.d/hubris-internal.conf` with `address=/auth.hubris.network/192.168.8.175`,
|
||
`server=192.168.8.1`, `server=1.1.1.1`, `interface=lo`, `bind-interfaces`, `no-hosts`, `no-resolv`.
|
||
|
||
- `pct set 114 --nameserver "127.0.0.1 192.168.8.1 1.1.1.1"` so the LXC starts with the right resolver order.
|
||
|
||
- `/etc/hosts` override kept as belt-and-suspenders.
|
||
|
||
|
||
Also needs `allow_local_remote_servers=true` in `config.php`.
|
||
|
||
|
||
> Once internal DNS at the router level handles `*.hubris.network`, this LXC-local dnsmasq can go away. See [DNS](../infrastructure/dns.md).
|
||
|
||
|
||
## Storage
|
||
|
||
Files at `/mnt/library/homecloud`. Owned by Nextcloud''s own permission model — **deliberately not on the `media` group**,
|
||
NC manages it itself and would complain. See [media permissions](../infrastructure/media-permissions.md).
|
||
|
||
|
||
## Web stack (Apache event MPM + php-fpm)
|
||
|
||
Apache 2.4 with **`mpm_event`** as the MPM and PHP served via **php-fpm 8.4** over the Unix socket `/run/php/php8.4-fpm.sock`
|
||
(mod_php disabled). Routing to FPM is via the distro''s `/etc/apache2/conf-available/php8.4-fpm.conf` (gated by `<IfModule
|
||
!mod_php8.c>`, so just `a2dismod php8.4` flips Apache onto FPM).
|
||
|
||
|
||
Why this matters: with mod_php every Apache worker is forced onto `mpm_prefork` (one heavy process per connection, full
|
||
PHP interpreter loaded in each). On FPM + event, Apache workers are ~5–8 MB and async; PHP work is done by a separate
|
||
FPM pool that keeps opcache hot. Concurrent upload throughput jumps significantly because Apache no longer blocks per-connection
|
||
on a heavy PHP process.
|
||
|
||
|
||
### PHP-FPM pool
|
||
|
||
- Socket: `/run/php/php8.4-fpm.sock` (owner/group `www-data`)
|
||
|
||
- `pm = dynamic`, `pm.max_children = 30`, `pm.start_servers = 6`, `pm.min/max_spare_servers = 4/12`, `pm.max_requests
|
||
= 500`
|
||
|
||
- Steady-state ~6 idle workers at ~80 MB each — bump `max_children` only if peak concurrency exceeds 30.
|
||
|
||
|
||
### PHP limits (FPM + CLI both)
|
||
|
||
Set in `/etc/php/8.4/fpm/php.ini` **and** `/etc/php/8.4/cli/php.ini` (the latter so `occ`/cron see the same ceilings):
|
||
|
||
|
||
| key | value |
|
||
|
||
|---|---|
|
||
|
||
| `upload_max_filesize` | `16G` |
|
||
|
||
| `post_max_size` | `16G` |
|
||
|
||
| `memory_limit` | `512M` |
|
||
|
||
| `max_execution_time` | `3600` |
|
||
|
||
| `max_input_time` | `3600` |
|
||
|
||
| `output_buffering` | `Off` |
|
||
|
||
|
||
After edits: `systemctl reload php8.4-fpm`. Caddy in front (`cloud.hubris.network`) imposes no body-size limit, so these
|
||
are the effective ceiling.
|
||
|
||
|
||
### Opcache + JIT
|
||
|
||
`/etc/php/8.4/mods-available/opcache.ini` — full config maintained, not commented stubs:
|
||
|
||
- `opcache.enable=1`, `opcache.memory_consumption=256`, `opcache.interned_strings_buffer=32`, `opcache.max_accelerated_files=20000`
|
||
|
||
- `opcache.jit=tracing`, `opcache.jit_buffer_size=64M`
|
||
|
||
|
||
### Apache mod_reqtimeout
|
||
|
||
`/etc/apache2/mods-available/reqtimeout.conf` — relaxed body trickle so slow-WAN uploads don''t get killed:
|
||
|
||
- `RequestReadTimeout header=20-40,MinRate=500`
|
||
|
||
- `RequestReadTimeout body=20,MinRate=100` (was `body=10,MinRate=500`)
|
||
|
||
|
||
## MariaDB tuning
|
||
|
||
Overrides in `/etc/mysql/mariadb.conf.d/99-nextcloud-tuning.cnf`:
|
||
|
||
- `innodb_buffer_pool_size = 1G` (default was 128M)
|
||
|
||
- `innodb_log_file_size = 256M`
|
||
|
||
- `innodb_flush_log_at_trx_commit = 2` (group-commit; trades a few ms of durability for write throughput — acceptable
|
||
for NC)
|
||
|
||
- `innodb_flush_method = O_DIRECT`, `innodb_io_capacity = 2000` / `..._max = 4000`
|
||
|
||
- `max_allowed_packet = 256M`, `tmp_table_size = 64M`, `max_heap_table_size = 64M`
|
||
|
||
|
||
Apply with `systemctl restart mariadb` (not reload — `innodb_log_file_size` needs a clean restart).
|
||
|
||
|
||
## Related
|
||
|
||
- [mulita (120)](120-mule-images.md) — reads NC user trees + writes back via WebDAV
|
||
|
||
- [Authentik (124)](106-auth-outpost.md)
|
||
|
||
- [Caddy (121)](121-caddy.md)
|
||
|
||
- [DNS](../infrastructure/dns.md)
|
||
|
||
- [Mesh migration (DNS overrides explained)](../infrastructure/mesh.md)
|
||
|
||
- [Media permissions](../infrastructure/media-permissions.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-05-13 — Apache keepalive bumped (and TCP BBR on host)
|
||
|
||
Follow-up to the FPM/MPM/MariaDB round earlier the same day. After a WiFi-bound Linux client kept seeing ~2 MB/s downloads
|
||
despite a 152 Mbps speed test (so the WiFi medium wasn''t the cap), bumped Apache `KeepAliveTimeout 5 → 60` and `MaxKeepAliveRequests
|
||
100 → 1000` so the mirall desktop client reuses one TLS+TCP connection across many sequential file fetches instead of
|
||
paying handshake + slow-start each time. Paired with TCP BBR + `fq` qdisc enabled on the [hubris host](../hosts/hubris.md#2026-05-13--tcp-bbr--fq-qdisc-enabled),
|
||
which directly addresses single-stream throughput on jittery links. Server-side loopback baseline post-BBR is ~404 MB/s
|
||
for a single 23 MB CR2 download via Caddy. Client-side diagnosis pending (curl single-stream + parallel tests) to determine
|
||
if the remaining gap is mirall config (bandwidth limit / `parallelNetworkRequests`) or path shaping.
|
||
|
||
|
||
### 2026-05-13 — throughput tuning: FPM + event MPM + opcache JIT + MariaDB
|
||
|
||
Investigated upload/download speed. Bottlenecks were inside this LXC — LAN (33.9 Gbit/s caddy→NC iperf3) and `/mnt/library`
|
||
NVMe (714 MB/s sustained) were not the limit. Changes:
|
||
|
||
- **Apache mod_php → php-fpm 8.4** (`a2dismod php8.4`). Flips PHP from in-process to FPM-over-Unix-socket. Apache workers
|
||
now ~5–8 MB instead of ~80 MB each.
|
||
|
||
- **MPM prefork → event** (`a2dismod mpm_prefork && a2enmod mpm_event`). Async I/O, far better concurrency for parallel
|
||
uploads.
|
||
|
||
- **opcache JIT enabled** (`tracing`, 64M buffer), `memory_consumption=256`, `interned_strings_buffer=32`, `max_accelerated_files=20000`.
|
||
|
||
- **MariaDB**: `innodb_buffer_pool_size 128M → 1G`, `innodb_log_file_size 96M → 256M`, `flush_log_at_trx_commit 1 → 2`,
|
||
`O_DIRECT`, higher I/O capacity. New file `99-nextcloud-tuning.cnf`.
|
||
|
||
- **Apache mod_reqtimeout**: body trickle relaxed from `body=10,MinRate=500` → `body=20,MinRate=100` so slow-WAN uploads
|
||
aren''t killed.
|
||
|
||
- **PHP-FPM ini** synced to match what the apache2/cli inis already had (FPM was stuck at 2M/8M because FPM wasn''t serving
|
||
traffic before, but now it does).
|
||
|
||
|
||
Backups: `/etc/php/8.4/fpm/php.ini.bak.20260513`, `/etc/apache2/mods-available/reqtimeout.conf.bak.20260513`, `/etc/php/8.4/mods-available/opcache.ini.bak.20260513`.
|
||
MariaDB tuning is a new file (rollback = `rm`).
|
||
|
||
|
||
Verified: `apache2ctl configtest` Syntax OK, `occ status` clean, `occ setupchecks` shows ✓ on memory_limit / output_buffering
|
||
/ DB indices, `status.php` returns 200 in <40 ms via Caddy with HTTP/2.
|
||
|
||
|
||
### 2026-05-10 — PHP upload limits raised
|
||
|
||
Apache + CLI php.ini bumped from distro defaults (`upload_max_filesize=2M`, `post_max_size=8M`, `memory_limit=128M`, `max_execution_time=30`)
|
||
to NC-recommended (`16G/16G/512M/3600`). Backups at `/etc/php/8.4/{apache2,cli}/php.ini.bak-20260510-*`. Was causing slow/failing
|
||
uploads via web UI.
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Initial documentation.
|
||
|
||
|
||
### 2026-04-26 — Nextcloud-rooted libraries shipped (mulita)
|
||
|
||
[mulita (120)](120-mule-images.md) now reads photos from `/nextcloud-users/<nc_user>/files/...` and writes mutations back
|
||
via WebDAV. Pre-migration DB dump kept at `/root/snapshots/mulita-pre-nc-migration-20260426-075132.dump`.
|
||
|
||
|
||
### 2026-04-21 — wired into Authentik (native OIDC + nc_uid override)
|
||
|
||
Local `dnsmasq` added on this LXC because Guzzle bypasses `/etc/hosts`. Resolver order set to `127.0.0.1 192.168.8.1 1.1.1.1`.
|
||
Username override `dtoro → admin`.
|
||
|
||
'
|
||
entity_slug: lxc:nextcloud
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: nextcloud
|
||
ip: 192.168.8.224
|
||
privileged: privileged
|
||
resources: 4 cores / 6 GiB RAM / 25 GiB rootfs
|
||
mounts: /mnt/library /mnt/library
|
||
public_host: cloud.hubris.network → [caddy]
|
||
changelog:
|
||
- date: '2026-05-13'
|
||
title: Apache keepalive bumped (and TCP BBR on host)
|
||
body: Follow-up to the FPM/MPM/MariaDB round earlier the same day. After a WiFi-bound Linux client kept seeing ~2 MB/s
|
||
downloads despite a 152 Mbps speed test (so the WiFi medium wasn't the cap), bumped Apache `KeepAliveTimeout 5 → 60`
|
||
and `MaxKeepAliveRequests 100 → 1000` so the mirall desktop client reuses one TLS+TCP connection across many sequential
|
||
file fetches instead of paying handshake + slow-start each time. Paired with TCP BBR + `fq` qdisc enabled on the [hubris
|
||
host](../hosts/hubris.md#2026-05-13--tcp-bbr--fq-qdisc-enabled), which directly addresses single-stream throughput on
|
||
jittery links. Server-side loopback baseline post-BBR is ~404 MB/s for a single 23 MB CR2 download via Caddy. Client-side
|
||
diagnosis pending (curl single-stream + parallel tests) to determine if the remaining gap is mirall config (bandwidth
|
||
limit / `parallelNetworkRequests`) or path shaping.
|
||
- date: '2026-05-13'
|
||
title: 'throughput tuning: FPM + event MPM + opcache JIT + MariaDB'
|
||
body: 'Investigated upload/download speed. Bottlenecks were inside this LXC — LAN (33.9 Gbit/s caddy→NC iperf3) and `/mnt/library`
|
||
NVMe (714 MB/s sustained) were not the limit. Changes: - **Apache mod_php → php-fpm 8.4** (`a2dismod php8.4`). Flips
|
||
PHP from in-process to FPM-over-Unix-socket. Apache workers now ~5–8 MB instead of ~80 MB each. - **MPM prefork → event**
|
||
(`a2dismod mpm_prefork && a2enmod mpm_event`). Async I/O, far better concurrency for parallel uploads. - **opcache JIT
|
||
enabled** (`tracing`, 64M buffer), `memory_consumption=256`, `interned_strings_buffer=32`, `max_accelerated_files=20000`.
|
||
- **MariaDB**: `innodb_buffer_pool_size 128M → 1G`, `innodb_log_file_size 96M → 256M`, `flush_log_at_trx_commit 1 →
|
||
2`, `O_DIRECT`, higher I/O capacity. New file `99-nextcloud-tuning.cnf`. - **Apache mod_reqtimeout**: body trickle relaxed
|
||
from `body=10,MinRate=500` → `body=20,MinRate=100` so slow-WAN uploads aren''t killed. - **PHP-FPM ini** synced to match
|
||
what the apache2/cli inis already had (FPM was stuck at 2M/8M because FPM wasn''t serving traffic before, but now it
|
||
does). Backups: `/etc/php/8.4/fpm/php.ini.bak.20260513`, `/etc/apache2/mods-available/reqtimeout.conf.bak.20260513`,
|
||
`/etc/php/8.4/mods-available/opcache.ini.bak.20260513`. MariaDB tuning is a new file (rollback = `rm`). Verified: `apache2ctl
|
||
configtest` Syntax OK, `occ status` clean, `occ setupchecks` shows ✓ on memory_limit / output_buffering / DB indices,
|
||
`status.php` returns 200 in <40 ms via Caddy with HTTP/2.'
|
||
- date: '2026-05-10'
|
||
title: PHP upload limits raised
|
||
body: Apache + CLI php.ini bumped from distro defaults (`upload_max_filesize=2M`, `post_max_size=8M`, `memory_limit=128M`,
|
||
`max_execution_time=30`) to NC-recommended (`16G/16G/512M/3600`). Backups at `/etc/php/8.4/{apache2,cli}/php.ini.bak-20260510-*`.
|
||
Was causing slow/failing uploads via web UI.
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-26'
|
||
title: Nextcloud-rooted libraries shipped (mulita)
|
||
body: '[mulita (120)](120-mule-images.md) now reads photos from `/nextcloud-users/<nc_user>/files/...` and writes mutations
|
||
back via WebDAV. Pre-migration DB dump kept at `/root/snapshots/mulita-pre-nc-migration-20260426-075132.dump`.'
|
||
- date: '2026-04-21'
|
||
title: wired into Authentik (native OIDC + nc_uid override)
|
||
body: Local `dnsmasq` added on this LXC because Guzzle bypasses `/etc/hosts`. Resolver order set to `127.0.0.1 192.168.8.1
|
||
1.1.1.1`. Username override `dtoro → admin`.
|
||
- slug: containers/118-elementsynapse
|
||
title: 118 — `elementsynapse`
|
||
content: '# 118 — `elementsynapse`
|
||
|
||
|
||
Matrix homeserver (Synapse). Backs `@dtoro:avispero`.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `elementsynapse`
|
||
|
||
- **IP:** `192.168.8.242`
|
||
|
||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||
|
||
- **Privilege:** **unprivileged**
|
||
|
||
- **Resources:** 1 core / 2 GiB RAM / **16 GiB rootfs** (grown from 8 GiB on 2026-05-15 after disk-full incident)
|
||
|
||
- **Mounts:** none from `/mnt/library`
|
||
|
||
- **Public hostname:** [`matrix.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md) → `:8008`
|
||
|
||
- **`server_name`:** `avispero` (intentionally different from the public URL).
|
||
|
||
- **Database:** Synapse runs on **SQLite** (`/var/lib/matrix-synapse/homeserver.db`, ~90 MiB). PostgreSQL on `127.0.0.1:5432`
|
||
is dedicated to the mautrix bridge databases (one per bridge), not Synapse itself.
|
||
|
||
|
||
## Service / port map
|
||
|
||
|
||
| Service | Listen | Notes |
|
||
|
||
| -------------- | ------- | -------------------------------- |
|
||
|
||
| synapse | `:8008` | Matrix client-server / federation HTTP. Caddy terminates TLS for `matrix.hubris.network`.
|
||
|
|
||
|
||
| synapse-admin | `:5173` | Web UI (vite dev server) for Synapse admin API. |
|
||
|
||
| mautrix-whatsapp | `:29318` | Docker, `network_mode: host`, compose in `/root/mautrix-whatsapp/`. |
|
||
|
||
| mautrix-meta | `127.0.0.1:29319` | Facebook/Instagram bridge, `/root/mautrix-meta/`. |
|
||
|
||
| mautrix-signal | `127.0.0.1:29328` | `/root/mautrix-signal/`. |
|
||
|
||
| mautrix-slack | `127.0.0.1:29335` | `/root/mautrix-slack/`. |
|
||
|
||
| mautrix-linkedin | `127.0.0.1:29341` | `/root/mautrix-linkedin/`. |
|
||
|
||
|
||
## Bridges (mautrix)
|
||
|
||
|
||
All five bridges run as plain `docker compose` stacks under `/root/mautrix-<name>/` — no systemd unit, `restart: unless-stopped`.
|
||
Their appservice registration YAMLs live in `/etc/matrix-synapse/` and are listed in `homeserver.yaml` under `app_service_config_files`.
|
||
If a bridge daemon is down, Synapse logs `push_bulk … ConnectionRefusedError` and retries every 512 s; restart the bridge
|
||
with `cd /root/mautrix-<name> && docker compose up -d`.
|
||
|
||
|
||
## Operational notes
|
||
|
||
- Rootfs is on `pve` LVM-thin, grow with `pct resize 118 rootfs +<N>G` from the host — resize2fs runs automatically.
|
||
|
||
- `/var/lib/matrix-synapse/media_store` is the dominant grower (~2 GiB at last check). If disk pressure returns, purge
|
||
remote media via the Synapse admin API before resizing further.
|
||
|
||
|
||
## Known issues
|
||
|
||
- ~~Disk usage was 86.8% at last legacy monitor reading on 2026-04-21~~ — resolved by growing rootfs to 16 GiB on 2026-05-15.
|
||
|
||
|
||
## Related
|
||
|
||
- ~~[claudio-bot (123)](archive/123-claudio-bot.md)~~ — decommissioned 2026-06-04, replaced by Hermes Agent
|
||
|
||
- [Caddy](121-caddy.md)
|
||
|
||
- [DNS](../infrastructure/dns.md)
|
||
|
||
- [Monitoring](../infrastructure/monitoring.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-06 — DHCP drift fixed: internal `/etc/network/interfaces` was `dhcp` despite Proxmox static config
|
||
|
||
**Symptom:** Matrix was down. Caddy at `192.168.8.175` couldn''t reach `192.168.8.239:8008` — the LXC was actually at
|
||
`192.168.8.244` because the guest-side dhclient had overridden the PVE-assigned static IP.
|
||
|
||
|
||
**Root cause:** During the 2026-06-02 static-IP migration, `pct set 118 --net0 ... ip=192.168.8.239/24` was applied to
|
||
the Proxmox config, but the internal `/etc/network/interfaces` still had `iface eth0 inet dhcp`. On every DHCP lease renewal,
|
||
dhclient grabbed `.244` from Technitium''s pool.
|
||
|
||
|
||
**Fix:**
|
||
|
||
- Replaced `iface eth0 inet dhcp` with `iface eth0 inet static` + `address 192.168.8.239/24` + `gateway 192.168.8.1`
|
||
|
||
- `ifdown eth0 && ifup eth0` applied the static IP
|
||
|
||
- Killed lingering dhclient process
|
||
|
||
- Verified: `curl http://192.168.8.239:8008` returns 302 from Caddy''s LXC
|
||
|
||
|
||
**Prevention:** The `check-caddy-backends.sh` cron on hubris now runs every 10 minutes, which would have caught this drift
|
||
within 10 minutes of occurrence.
|
||
|
||
|
||
### 2026-05-15 — phantom-notification cleanup for `@admin`
|
||
|
||
After the disk-full incident, the mobile (Element X) badge showed ~125 unread but every room read clean in the UI. Root
|
||
cause: stale rows in `event_push_actions` that were never reaped — Synapse''s read-receipt-driven cleanup didn''t catch
|
||
up. Two contributors:
|
||
|
||
1. **8 of 12 affected rooms** had read receipts past the "unread" stream_ordering — pure stale state, likely from the
|
||
disk-full window stalling rotation/cleanup.
|
||
|
||
2. **The noisiest room** (a Signal-bridged DM) was flooded with `m.room.redaction` events from `@signalbot:avispero` (disappearing
|
||
messages being deleted). Synapse counts redactions as notifications by default, but Element X doesn''t render them — so
|
||
they were invisible "ghosts".
|
||
|
||
|
||
Fix (after stopping Synapse and backing up `homeserver.db`):
|
||
|
||
- `DELETE FROM event_push_actions WHERE stream_ordering <= max(receipt_stream)` per (user, room) — clears receipt-superseded
|
||
rows.
|
||
|
||
- `DELETE FROM event_push_actions WHERE event_id IN (SELECT event_id FROM events WHERE type=''m.room.redaction'')` — removes
|
||
redaction-driven notifications.
|
||
|
||
- Rebuild `event_push_summary` from the remaining rows (preserve `thread_id=''main''` to satisfy the NOT-NULL check constraint).
|
||
|
||
|
||
Result: 125 → 4 (legitimate WhatsApp unreads).
|
||
|
||
|
||
Then, to stop this recurring, an **override push-rule** was installed for `@admin` via the client API:
|
||
|
||
```
|
||
|
||
PUT /_matrix/client/v3/pushrules/global/override/suppress_redactions
|
||
|
||
{ "actions": [], "conditions": [{ "kind": "event_match", "key": "type", "pattern": "m.room.redaction" }] }
|
||
|
||
```
|
||
|
||
With empty `actions`, all `m.room.redaction` events are silently ignored by the push system. Inspect/remove via the same
|
||
endpoint. (Note: rule ids beginning with `.` are reserved for server defaults — custom ids cannot start with a dot.)
|
||
|
||
|
||
### 2026-05-15 — disk-full incident, rootfs grown to 16 GiB, bridges restarted
|
||
|
||
Element/Synapse was unreliable for ~2 days: messages dropping, sync requests stalling, mautrix bridges (slack/signal/meta/linkedin/whatsapp)
|
||
all exited around 2026-05-13. Root cause: rootfs reached 100% (19 MB free of 7.8 G); Synapse logs showed `OSError: [Errno
|
||
28] No space left on device` during flush. Fix: `pct resize 118 rootfs +8G` (host `pve` VG had plenty of headroom), `journalctl
|
||
--vacuum-size=50M` to recover ~200 M, `systemctl restart matrix-synapse`, then `docker compose up -d` in each `/root/mautrix-*`
|
||
directory. Also expanded this wiki page to document the bridges, synapse-admin, and the actual DB backend (Postgres, not
|
||
the SQLite leftover).
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Initial documentation.
|
||
|
||
'
|
||
entity_slug: lxc:elementsynapse
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: elementsynapse
|
||
ip: 192.168.8.242
|
||
host: strong
|
||
privileged: unprivileged
|
||
resources: 1 core / 2 GiB RAM / 16 GiB rootfs
|
||
mounts: none from /mnt/library
|
||
public_host: matrix.hubris.network → caddy → :8008
|
||
'`server_name`': avispero (intentionally different from the public URL).
|
||
database: Synapse runs on SQLite (/var/lib/matrix-synapse/homeserver.db, ~90 MiB). PostgreSQL on 127.0.0.1:5432 is dedicated
|
||
to the mautrix bridge databases (one per bridge), not Synapse itself.
|
||
changelog:
|
||
- date: '2026-06-06'
|
||
title: 'DHCP drift fixed: internal `/etc/network/interfaces` was `dhcp` despite Proxmox static config'
|
||
body: '**Symptom:** Matrix was down. Caddy at `192.168.8.175` couldn''t reach `192.168.8.239:8008` — the LXC was actually
|
||
at `192.168.8.244` because the guest-side dhclient had overridden the PVE-assigned static IP. **Root cause:** During
|
||
the 2026-06-02 static-IP migration, `pct set 118 --net0 ... ip=192.168.8.239/24` was applied to the Proxmox config,
|
||
but the internal `/etc/network/interfaces` still had `iface eth0 inet dhcp`. On every DHCP lease renewal, dhclient grabbed
|
||
`.244` from Technitium''s pool. **Fix:** - Replaced `iface eth0 inet dhcp` with `iface eth0 inet static` + `address
|
||
192.168.8.239/24` + `gateway 192.168.8.1` - `ifdown eth0 && ifup eth0` applied the static IP - Killed lingering dhclient
|
||
process - Verified: `curl http://192.168.8.239:8008` returns 302 from Caddy''s LXC **Prevention:** The `check-caddy-backends.sh`
|
||
cron on hubris now runs every 10 minutes, which would have caught this drift within 10 minutes of occurrence.'
|
||
- date: '2026-05-15'
|
||
title: phantom-notification cleanup for `@admin`
|
||
body: 'After the disk-full incident, the mobile (Element X) badge showed ~125 unread but every room read clean in the
|
||
UI. Root cause: stale rows in `event_push_actions` that were never reaped — Synapse''s read-receipt-driven cleanup didn''t
|
||
catch up. Two contributors: 1. **8 of 12 affected rooms** had read receipts past the "unread" stream_ordering — pure
|
||
stale state, likely from the disk-full window stalling rotation/cleanup. 2. **The noisiest room** (a Signal-bridged
|
||
DM) was flooded with `m.room.redaction` events from `@signalbot:avispero` (disappearing messages being deleted). Synapse
|
||
counts redactions as notifications by default, but Element X doesn''t render them — so they were invisible "ghosts".
|
||
Fix (after stopping Synapse and backing up `homeserver.db`): - `DELETE FROM event_push_actions WHERE stream_ordering
|
||
<= max(receipt_stream)` per (user, room) — clears receipt-superseded rows. - `DELETE FROM event_push_actions WHERE event_id
|
||
IN (SELECT event_id FROM events WHERE type=''m.room.redaction'')` — removes redaction-driven notifications. - Rebuild
|
||
`event_push_summary` from the remaining rows (preserve `thread_id=''main''` to satisfy the NOT-NULL check constraint).
|
||
Result: 125 → 4 (legitimate WhatsApp unreads). Then, to stop this recurring, an **override push-rule** was installed
|
||
for `@admin` via the client API: ``` PUT /_matrix/client/v3/pushrules/global/override/suppress_redactions { "actions":
|
||
[], "conditions": [{ "kind": "event_match", "key": "type", "pattern": "m.room.redaction" }] } ``` With empty `actions`,
|
||
all `m.room.redaction` events are silently ignored by the push system. Inspect/remove via the same endpoint. (Note:
|
||
rule ids beginning with `.` are reserved for server defaults — custom ids cannot start with a dot.)'
|
||
- date: '2026-05-15'
|
||
title: disk-full incident, rootfs grown to 16 GiB, bridges restarted
|
||
body: 'Element/Synapse was unreliable for ~2 days: messages dropping, sync requests stalling, mautrix bridges (slack/signal/meta/linkedin/whatsapp)
|
||
all exited around 2026-05-13. Root cause: rootfs reached 100% (19 MB free of 7.8 G); Synapse logs showed `OSError: [Errno
|
||
28] No space left on device` during flush. Fix: `pct resize 118 rootfs +8G` (host `pve` VG had plenty of headroom),
|
||
`journalctl --vacuum-size=50M` to recover ~200 M, `systemctl restart matrix-synapse`, then `docker compose up -d` in
|
||
each `/root/mautrix-*` directory. Also expanded this wiki page to document the bridges, synapse-admin, and the actual
|
||
DB backend (Postgres, not the SQLite leftover).'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- slug: containers/119-sophia
|
||
title: 119 — `sophia`
|
||
content: '# 119 — `sophia`
|
||
|
||
|
||
Single-container workload "sophia". Reads/writes under `/mnt/library/sophia`.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `sophia`
|
||
|
||
- **IP:** `192.168.8.157`
|
||
|
||
- **Privilege:** privileged
|
||
|
||
- **Resources:** 2 cores / 1 GiB RAM / 10 GiB rootfs
|
||
|
||
- **Mounts:** `/mnt/library` ↔ `/mnt/library`
|
||
|
||
- **Public hostname:** none
|
||
|
||
|
||
## Permissions
|
||
|
||
LXC has `www-data` in the [media GID 10000 group](../infrastructure/media-permissions.md). The `/mnt/library/sophia` subtree
|
||
itself is single-container and not on the setgid shared group.
|
||
|
||
|
||
## Related
|
||
|
||
- [Hubris host](../hosts/hubris.md)
|
||
|
||
- [Media permissions](../infrastructure/media-permissions.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Stub. Document what runs inside next time we work on this LXC.
|
||
|
||
'
|
||
entity_slug: lxc:sophia
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: sophia
|
||
ip: 192.168.8.157
|
||
privileged: privileged
|
||
resources: 2 cores / 1 GiB RAM / 10 GiB rootfs
|
||
mounts: /mnt/library /mnt/library
|
||
public_host: none
|
||
changelog:
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Stub. Document what runs inside next time we work on this LXC.
|
||
- slug: containers/120-mule-images
|
||
title: 120 — `mule-images`
|
||
content: "# 120 — `mule-images`\n\nHosts `mule-image` — the photos app at `photos.hubris.network`. PhotoPrism + Go sidecar\
|
||
\ + SvelteKit, replacing the legacy FastAPI/Celery stack as of 2026-05-22 (see Changelog). Auto-deploys from `dtoro/mule-image`\
|
||
\ on `git push origin main`.\n\n## At a glance\n- **Hostname:** `mule-images`\n- **IP:** `192.168.8.136`\n- **Privilege:**\
|
||
\ privileged\n- **Resources:** 6 cores / 12 GiB RAM / 60 GiB rootfs / 2 GiB swap\n- **Mounts:** `/mnt/library` ↔ `/mnt/library`;\
|
||
\ `/dev/dri/{card0,renderD128}` passed through for VA-API video accel on the AMD Phoenix1 iGPU.\n- **Public hostname:**\
|
||
\ [`photos.hubris.network`](../infrastructure/dns.md) → [caddy](121-caddy.md) → path-routed to PhotoPrism `:2342` / sidecar\
|
||
\ `:8000` / nginx-static `:3000`.\n\n## Stack (`/opt/mule-image`)\n\n`/opt/mule-image` IS the working tree of `dtoro/mule-image`.\
|
||
\ Compose stack: base `docker-compose.yml` + LAN-port-rebind `docker-compose.override.yml` (untracked) + VA-API `docker-compose.gpu.yml`.\
|
||
\ The SvelteKit frontend is built on the host and served as a static bundle by nginx — no `vite dev` in prod.\n\n| Service\
|
||
\ / process | Port (LAN) | Notes |\n| ------------------------ | --------------------- | ----- |\n|\
|
||
\ pp-mariadb | (internal 3306) | MariaDB 11; holds PhotoPrism's `photoprism` DB + sidecar's `mule_sidecar.*`\
|
||
\ |\n| pp-app | `0.0.0.0:2342` | PhotoPrism `:latest`; `PHOTOPRISM_FFMPEG_ENCODER=h264_vaapi`\
|
||
\ |\n| pp-sidecar | `0.0.0.0:8000` | Go service (file rename / folder ops / heap convert / dup scan\
|
||
\ / per-photo marks); also reconciles `USER_BASEPATHS` into PhotoPrism's `auth_users` every 60 s |\n| nginx (host process)\
|
||
\ | `0.0.0.0:3000` | Serves the SvelteKit static bundle from `web/build/` |\n\n**Per-user scoping** is driven\
|
||
\ by `USER_BASEPATHS=admin:admin/files, muli:muli/files` in `.env`. Sidecar applies it to PhotoPrism's `auth_users` table\
|
||
\ on boot + every 60s, `mkdir -p`s each target so PhotoPrism's ACL filter has somewhere to point.\n\n`.env` is **untracked**\
|
||
\ — `git reset --hard` won't touch it, but `git clean -fdx` would. Holds PP_*, MariaDB passwords, SIDECAR_DB_PASSWORD,\
|
||
\ OIDC_* (existing `mule-image` Authentik app, redirect URI now `/api/v1/oidc/redirect`), `USER_BASEPATHS`, and `PP_FFMPEG_ENCODER=vaapi`.\n\
|
||
\n`docker-compose.override.yml` is also untracked — it `!override`s the upstream `127.0.0.1:port` mappings to `0.0.0.0:port`\
|
||
\ so cross-host Caddy on [LXC 121](121-caddy.md) can reach pp-app + sidecar.\n\n## Library access\n\nPhotoPrism reads\
|
||
\ the library **directly off the bind-mounted filesystem** — no Nextcloud webhook integration in the new stack. The base\
|
||
\ path is `/mnt/library/homecloud` and per-user scoping comes from `USER_BASEPATHS` (see Stack section above).\n\n- `dtoro`\
|
||
\ is mapped to NC user `admin` → `/mnt/library/homecloud/admin/files/`\n- `muli` is mapped to NC user `muli` → `/mnt/library/homecloud/muli/files/`\n\
|
||
\nPhotoPrism's container user is `33:10000` (`www-data:media`), matching the host ownership of the NC tree. The legacy\
|
||
\ `oc_webhook_listeners` rows + `NEXTCLOUD_WEBHOOK_SECRET` are gone — when NC writes via WebDAV (still its own primary\
|
||
\ surface), PhotoPrism picks up the new file on its next index pass.\n\n## Authentication\n\nOIDC via Authentik. App slug\
|
||
\ `mule-image`, redirect URI `https://photos.hubris.network/api/v1/oidc/redirect` (PhotoPrism's auto-derived path; PhotoPrism\
|
||
\ builds it from `PHOTOPRISM_SITE_URL`). `OIDC_REGISTER=true` auto-creates a PhotoPrism user at role `user` on first SSO\
|
||
\ login. The sidecar's basepath reconciler then assigns their scoped folder.\n\nThe `mule-image` Authentik app's redirect\
|
||
\ URI was migrated from the legacy FastAPI `/api/v1/auth/oidc/callback` path on 2026-05-22 — same client ID/secret were\
|
||
\ reused. The separate `mule-photos-new` Authentik app was deleted in the same operation.\n\n## Auto-deploy\n\nPush to\
|
||
\ `dtoro/mule-image` `main` → gitea webhook → `http://192.168.8.136:9797/deploy` → `mule-deploy-webhook.service`:\n- Validates\
|
||
\ HMAC against `/etc/mule-deploy/secret`\n- Filters to `refs/heads/main`\n- Runs `/opt/mule-deploy/deploy.sh` in a daemon\
|
||
\ thread (returns 202 immediately — docker builds exceed gitea's request timeout)\n- `git fetch && git reset --hard origin/main`,\
|
||
\ force-recreates **only the sidecar** image (so PhotoPrism's `Session` HMAC key in `pp/storage/config/hub.yml` doesn't\
|
||
\ rotate and invalidate every in-flight OIDC state cookie), reconciles `pp-app` + `mariadb` in place, `cd web && npm ci\
|
||
\ && npm run build`, `systemctl reload nginx`.\n- Conditionally layers `docker-compose.gpu.yml` when `/dev/dri/renderD128`\
|
||
\ exists, and always layers `docker-compose.override.yml` (the LAN-port rebind) when present.\n\nDeploy tooling is **outside**\
|
||
\ the app repo: `/opt/mule-deploy/{deploy.sh,webhook.py}`, secret at `/etc/mule-deploy/secret`. Gitea webhook id 6. `app.ini`\
|
||
\ `ALLOWED_HOST_LIST` on [gitea](104-gitea.md) includes `192.168.8.136`.\n\nLogs: `pct exec 120 -- journalctl -u mule-deploy-webhook\
|
||
\ -f`.\nManual deploy: `pct exec 120 -- /opt/mule-deploy/deploy.sh`.\n\nFor pushes from inside the LXC, gitea creds at\
|
||
\ `/etc/mule-deploy/git-credentials` (mode 600) — same token as `/etc/caddy-deploy/git-credentials` on [caddy](121-caddy.md).\n\
|
||
\n## Related\n- [Nextcloud (114)](114-nextcloud.md) — source of truth for photo libraries\n- [Authentik (124)](106-auth-outpost.md)\n\
|
||
- [Caddy (121)](121-caddy.md)\n- [DNS](../infrastructure/dns.md)\n- [Auto-deploy](../infrastructure/auto-deploy.md)\n\
|
||
- [Gitea (104)](104-gitea.md)\n\n## Changelog\n\n### 2026-05-22 — Cutover to PhotoPrism stack (Mulimage 2.0)\n\nThe `new`\
|
||
\ branch of `dtoro/mule-image` was merged into `main` as commit\n`70dc1b6`. The merge replaces the legacy FastAPI + PostgreSQL\
|
||
\ + Celery\n+ React stack with PhotoPrism + Go sidecar + MariaDB + SvelteKit, lifted\nin from the M0 evaluation on (now-destroyed)\
|
||
\ [LXC 127](#).\n\nCutover on 120 (this session):\n- Bumped to 6 cores / 2 GB swap. `/dev/dri/{card0,renderD128}` already\n\
|
||
\ passed through; added an opt-in `docker-compose.gpu.yml` overlay that\n the deploy script layers in when the device\
|
||
\ is present. PhotoPrism now\n uses `h264_vaapi` instead of `libx264`.\n- Installed host nginx + a small `photos.conf`\
|
||
\ vhost serving the SvelteKit\n static bundle from `/opt/mule-image/web/build/` on `:3000` — no Vite\n dev server in\
|
||
\ prod. SvelteKit's `adapter-static` produces a real SPA\n bundle.\n- Replaced `/opt/mule-deploy/deploy.sh` with the\
|
||
\ 127-style multi-compose\n deploy (force-recreates only sidecar to preserve PhotoPrism's Session\n HMAC key; reconciles\
|
||
\ pp-app + mariadb in place; rebuilds web/ and\n reloads nginx).\n- `/opt/mule-image/.env` was rewritten to the PhotoPrism\
|
||
\ shape. Reused\n passwords from `/root/mule-photos-new-secrets.txt`. `USER_BASEPATHS=\n admin:admin/files, muli:muli/files`\
|
||
\ exposes both libraries.\n- `docker-compose.override.yml` (untracked) `!override`s pp-app +\n sidecar ports to `0.0.0.0`\
|
||
\ so cross-LXC caddy on [121](121-caddy.md)\n can reach them.\n- Authentik: `mule-image` app's redirect URI updated to\
|
||
\ `/api/v1/oidc/redirect`;\n the separate `mule-photos-new` app deleted.\n- Caddyfile (`dtoro/caddy-conf`): `photos.hubris.network`\
|
||
\ switched from\n plain passthrough to path-matchers (PhotoPrism API + sidecar API +\n nginx static catch-all), and\
|
||
\ `photos-new.hubris.network` was removed\n entirely.\n- Cleanups: 4 Nextcloud `oc_webhook_listeners` rows for the legacy\n\
|
||
\ `/api/v1/internal/nc-webhook` endpoint deleted; gitea webhook id 9\n (the `refs/heads/new` deploy hook for 127) deleted;\
|
||
\ `192.168.8.181`\n removed from gitea's `ALLOWED_HOST_LIST`; `photos-new.hubris.network`\n removed from dnsmasq.\n\
|
||
- LXC 127 destroyed via `pct destroy 127` after the new stack passed\n curl verification end-to-end (PP `/api/v1/status`\
|
||
\ returns\n `{\"status\":\"operational\"}` through caddy; sidecar `/api/sidecar/healthz`\n returns `{\"ok\":true,...}`).\n\
|
||
- Rollback safety net: legacy `mulita` postgres dump at\n `/root/backups/mulita-legacy-final-20260521-212036.sql.gz`\
|
||
\ on hubris\n (10 MB, 16,155 photo rows). The `mule-image_mule_db_data` docker\n volume on 120 retains the on-disk postgres\
|
||
\ data for at least 24 h\n before housekeeping removes it.\n\n### 2026-05-14 — Drop AI/vision pipeline, plain Postgres,\
|
||
\ DB↔FS refresh\n\n**AI removal (commits `a27267f` + `423a73a`).** The OpenCLIP-on-ONNX\nbinary classifier (photography\
|
||
\ vs other) and all its scaffolding gone:\n\n- Backend: `app/services/vision/`, `app/tasks/vision.py`,\n `app/services/feature_flags.py`,\
|
||
\ `app/routers/features.py` deleted.\n Admin `/admin/feature-flags`, `/admin/ai/{backfill,rescan}` endpoints\n removed.\
|
||
\ `Photo.needs_review` column + `ix_photos_needs_review`\n index dropped (alembic `0019_drop_ai_remnants`).\n- Frontend:\
|
||
\ AI Settings tab, `useFeaturesQuery`, \"Needs Review\" sidebar\n entry + filter, `needs_review` URL param plumbing,\
|
||
\ FeatureFlag types\n all gone.\n- Infra: `worker-vision` compose service + `models_data` volume\n deleted. `worker-light`\
|
||
\ no longer runs `python -m\n app.services.vision.bootstrap_models` before celery. Backend\n Dockerfile drops the dedicated\
|
||
\ torch RUN layer. `requirements.txt`\n drops `torch`, `torchvision`, `open-clip-torch`, `onnxruntime`.\n\n**Postgres\
|
||
\ image swap.** `pgvector/pgvector:pg16` → `postgres:16`. The\n0019 migration drops the (now-unused) `vector` extension\
|
||
\ first; the\nSQL was pre-applied via `psql` against the still-pgvector container\nand `alembic_version` stamped to `0019`\
|
||
\ so the new backend's\n`bootstrap.py` upgrade-head was a no-op. After the swap surfaced a\ncollation-version mismatch\
|
||
\ (Bookworm glibc 2.36 → Trixie glibc 2.41),\nthe fix was `REINDEX DATABASE mulita` + `ALTER DATABASE … REFRESH\nCOLLATION\
|
||
\ VERSION` on `mulita`, `postgres`, `template1`. The data\nvolume was preserved across the image change.\n\n**One-shot\
|
||
\ DB ↔ filesystem refresh.** New script\n`backend/scripts/full_refresh.py` (run as\n`docker exec -w /app mulita-backend\
|
||
\ python -m scripts.full_refresh\n[--dry-run]`). Phases: `cleanup_data_integrity` → inline scan of every\nactive SourceRoot\
|
||
\ → `prune_missing_photos(dry_run=False)` → new\n`prune_orphan_thumbnails` helper that removes\n`/data/thumbs/{user}/{photo}/`\
|
||
\ dirs for photo_ids that left the DB.\nFirst run: 0 missing photos, 1 stale folder row, **7982 orphan\nthumbnail dirs**\
|
||
\ reaped.\n\n**Manual cleanup post-deploy.** `docker compose up -d --build\n--force-recreate` doesn't reap services removed\
|
||
\ from the YAML, so\n`mulita-worker-vision` had to be `docker rm -f`'d by hand. The\n`mule-image_models_data` volume was\
|
||
\ likewise removed manually.\n\n### 2026-05-11 — Stop duplicating Nextcloud's indexing (Phase 1 + 2)\n\nTwo big architectural\
|
||
\ shifts in one session, both aimed at killing work mule was doing that NC already does for the same source files.\n\n\
|
||
**Phase 1 — thumbnail proxy (commits `576b0c2`, `28738ac`).** Photos table gains `nextcloud_fileid` (migration `0018`).\
|
||
\ `GET /api/v1/photos/{id}/thumb/{size}` proxies NC's `/index.php/core/preview` keyed by that fileid, authenticated with\
|
||
\ the owner's encrypted app password (small=240, medium=640, large=1280). Worker now only writes the `medium` size to\
|
||
\ `/data/thumbs` (the vision worker still loads it from disk); `small` + `large` no longer touch disk. Disk fallback stays\
|
||
\ wired for legacy rows and the rare 404 from NC (iPhone JPEGs mis-extensioned as `.DNG` — verified). Existing 17,591\
|
||
\ photos backfilled via `backend/scripts/backfill_nextcloud_fileid.py`. Tested with real CR2, real DNG, fake DNG, JPEG,\
|
||
\ HEIC — all green.\n\n**Range support for `/original` (commit `18dce33`).** Pre-existing bug surfaced by Phase 1 testing:\
|
||
\ `<video>` couldn't play `.mov` because `FileResponse` was returning 200 with the whole body and no `Accept-Ranges` header,\
|
||
\ so browsers reported \"format not supported.\" Now parses `Range: bytes=START-END`, returns 206 with `Content-Range`,\
|
||
\ streams in 1 MB chunks.\n\n**Fix: backfill_gps was the actual CPU drain, not the watcher (commit `d24c64e`).** `_scan_all_source_roots_async`\
|
||
\ auto-queued `backfill_gps` 30 s after every container boot, which then re-queued `extract_metadata` for every photo\
|
||
\ with `latitude IS NULL` — ~60k pointless tasks per deploy, pinning worker-light at 180+% CPU for ~30 min. Killed the\
|
||
\ auto-trigger; manual `POST /api/v1/library/backfill-gps` still works.\n\n**Phase 2 — webhook receiver replaces watchfiles\
|
||
\ (commits `362fbc6`, `f657e2c`).** New endpoint `POST /api/v1/internal/nc-webhook` (auth via `Authorization: Bearer $NEXTCLOUD_WEBHOOK_SECRET`,\
|
||
\ `hmac.compare_digest`) handles `NodeCreated|NodeWritten|NodeDeleted|NodeRenamed`. Maps NC's `/admin/files/...` path\
|
||
\ to the bind-mounted `/nextcloud-users/admin/files/...` and dispatches the same `scan_folder` / `handle_file_deletion`\
|
||
\ machinery the watcher used. Registered against NC via `backend/scripts/register_nc_webhooks.py` (idempotent: deletes\
|
||
\ existing webhooks targeting the same URL first). `mulita-worker-watcher` container retired; celery `--beat` folded into\
|
||
\ `worker-light` so `discard_missing_photos_beat` still fires. End-to-end verified by uploading a test JPEG through WebDAV\
|
||
\ — both `NodeCreated` and `NodeWritten` fire, mule receives + dispatches, 200 OK.\n\n**NC cron tightened to `*/1` so\
|
||
\ webhook delivery latency drops from up to 5 min to ~60 s** (`crontab -u www-data` on LXC 114). NC dispatches webhook\
|
||
\ calls through its background-job queue; the cron interval = the worst-case latency. `*/5` was the default; `*/1` is\
|
||
\ the upstream recommendation anyway.\n\n**Post-Phase-2 delete-roundtrip patches (commit `9408825`).** End-to-end testing\
|
||
\ of the NC↔mule deletion paths surfaced two real gaps:\n\n1. *Folder delete* — NC fires one `NodeDeletedEvent` for the\
|
||
\ folder, not one per child. The webhook handler bailed with \"unsupported extension\" and photos under the deleted folder\
|
||
\ kept `is_discarded=false` until the 30-min reconcile sweep. Fix: `handle_directory_deletion()` does a single `UPDATE\
|
||
\ photos SET is_discarded=true WHERE filepath LIKE 'dir/%'` when the deleted path has no supported image extension.\n\
|
||
2. *Resurrect on rewrite* — PUT-overwrite of a previously-discarded file fired `NodeWrittenEvent` → `scan_folder`, but\
|
||
\ scan_folder's \"photo exists, skip\" branch left `is_discarded=true`. Fix: when the existing row is discarded, flip\
|
||
\ `is_discarded=false`, clear `discarded_at`, re-queue `extract_metadata`.\n\nVerified end-to-end on `Photos/MuleTestFolder/{test1,test2,nested/test3}.jpg`:\
|
||
\ `DELETE Photos/MuleTestFolder/` discarded all 3 in one shot; `PUT test1.jpg` back resurrected only that one.\n\n**Known\
|
||
\ remaining gap — trashbin restore.** Moving a file out of `/dav/trashbin/.../trash/foo.jpg.dXXXX` back to `/files/...`\
|
||
\ via WebDAV MOVE fires **no event mule subscribes to**. NC's trashbin app emits its own internal event class that isn't\
|
||
\ in the `OCP\\Files\\Events\\Node\\*` set we registered with `webhook_listeners`. Workaround: re-upload via PUT (covered\
|
||
\ by the resurrect-on-rewrite fix above) or wait for the 30-min reconcile sweep.\n\n**Folder rename round-trip (commits\
|
||
\ `f4a03b6` + `f27f3cb`).** Both directions now work:\n\n- *NC-side rename → mule:* webhook's \"renamed\" branch detects\
|
||
\ directory rename (neither path has a supported extension) and calls new `handle_directory_rename(old, new)` in `scan.py`.\
|
||
\ The helper iterates matching rows in Python and prefix-rewrites `Photo.filepath`, `Folder.path`, `SourceRoot.path` in\
|
||
\ one transaction. Cross-source-root case (rare) discards the old subtree and lets scan_folder add fresh rows under the\
|
||
\ new root.\n- *Mule-side rename → NC:* the existing `PATCH /api/v1/folders/{id}` endpoint (`folders.py:202`) already\
|
||
\ does WebDAV MOVE via `nextcloud_dav.move_for_user` and rewrites mule's DB inline. The NodeRenamedEvent that bounces\
|
||
\ back through the webhook hits `handle_directory_rename`, which finds 0 rows under the old prefix and is a no-op — feedback\
|
||
\ loop is idempotent.\n\n**asyncpg gotcha**: the original implementation of `handle_directory_rename` used raw SQL with\
|
||
\ `SUBSTRING(filepath FROM LENGTH(:old_prefix) + 1)` so the offset would be computed server-side. asyncpg's type inference\
|
||
\ miscategorises the `LENGTH()` result and rejects the parameter as \"$2: int (expected str)\". The fix iterates in Python\
|
||
\ (same pattern as the existing PATCH endpoint). Lesson: avoid passing `LENGTH(:x)+1` as an argument to `SUBSTRING(...)`\
|
||
\ via asyncpg + sqlalchemy `text()`.\n\n**Phase 3 — Memories-backed `extract_metadata` (commit `2a5759c`).** Memories\
|
||
\ app re-enabled on NC. `extract_metadata` now tries `GET /index.php/apps/memories/api/image/info/{fileid}` (auth: Basic\
|
||
\ + `OCS-APIRequest: true` header to bypass CSRF) before falling back to ExifTool. Replaces ~80 ms of subprocess with\
|
||
\ ~1-2 ms HTTP for ongoing imports. New helpers: `nextcloud_dav.fetch_memories_info_async()` + `metadata._apply_memories_metadata()`.\
|
||
\ We kept mule's full date-fallback chain (SubSec → DateTimeOriginal → CreateDate → MediaCreateDate → TrackCreateDate\
|
||
\ → filename heuristic → mtime) because 35% of the library (`taken_at_source='path'`) depends on the filename heuristic,\
|
||
\ and Memories alone would silently mis-date those photos to mtime. PhotoInfoPanel reads `exif.Make/Model/ISO/FNumber`\
|
||
\ from `photos.exif_json` — Memories' `exif` blob uses those exact plain key names, so no frontend adapter was needed.\
|
||
\ ExifTool subprocess is still in place as the fallback for brand-new photos racing the NC scan, non-NC photos, and any\
|
||
\ NC HTTP failure.\n\n**Phase 4 — retire `/data/thumbs` (commits `5a67ed7` + `7a1c6b6`).** Vision worker now fetches NC's\
|
||
\ 640px preview via a new sync helper `nextcloud_dav.get_preview_bytes()` instead of reading `/data/thumbs/{id}/medium.webp`.\
|
||
\ `thumbs.WORKER_THUMB_SIZES = set()` — `generate_thumbnails` still computes pHash on the original-res pixels (perceptual\
|
||
\ dedup is mule-only) but stops writing files. All ~22k `medium.webp` purged after verification; `/data/thumbs` shrank\
|
||
\ from 4.1 GB → 94 MB residual.\n\n**SECRET_KEY bug found mid-deploy (commit `7a1c6b6`).** `docker-compose.yml` only set\
|
||
\ `SECRET_KEY` on the `backend` service, not on the workers. Workers' Fernet-based `decrypt(nextcloud_app_password_enc)`\
|
||
\ silently returned empty and `_credentials_for()` raised `NextcloudCredentialsMissing`. This meant Phase 3's `extract_metadata`\
|
||
\ in worker-light had been silently falling back to ExifTool *the entire time*, and Phase 4's vision worker couldn't fetch\
|
||
\ NC previews at all. Fix replicates `SECRET_KEY=${SECRET_KEY:-...}` to all worker services. After any compose edit, run\
|
||
\ `docker exec mulita-<svc> env | grep SECRET_KEY` to confirm propagation.\n\n**Caveats worth knowing for Phase 4:**\n\
|
||
- Vision is `disabled` in production (`redis GET mulita:flags:vision.enabled == \"false\"`). The Phase 4 vision-from-NC\
|
||
\ path is correct but unexercised by live traffic; it'll matter whenever vision is re-enabled.\n- NC's preview generator\
|
||
\ hasn't covered the whole library at the 640px+ tier. A sample of 6 photos: 4 with proper sizes, 1 stuck at 160px (`IMG_4954.DNG`),\
|
||
\ 1 unfetchable (`IMG_0193.DNG`). If vision re-enables and quality matters, run `pct exec 114 -- sudo -u www-data php\
|
||
\ occ preview:generate-all` to backfill.\n- The HTTP `/thumb/{size}` endpoint's inline-regen disk fallback still writes\
|
||
\ one WebP per NC-404 event (mis-extensioned RAW etc). `/data/thumbs` will grow back very slowly from that path.\n\n###\
|
||
\ 2026-05-11 — Duplicates view shows parent folder + full-path tooltip\n\n`GET /api/v1/library/duplicates/groups` now\
|
||
\ includes `filepath` in each member payload. `DuplicatesView` renders a dark bottom-strip on every thumbnail showing\
|
||
\ the last two folder segments — `…/<grandparent>/<parent>` (e.g. `…/files/Phone Photos starting 01-26`) via a `duplicatePathLabel`\
|
||
\ helper. Going up two levels avoids the failure mode where two copies sit under matching parent names (e.g. duplicate\
|
||
\ `2023/` subfolders under different archives) and the label would collapse. Full filepath surfaces through the native\
|
||
\ `title` tooltip on hover. Dimensions chip moves from bottom-left to top-left so the path strip can run edge-to-edge.\
|
||
\ `memberToPhoto` finally stops faking `filepath = filename` — the synthetic Photo it hands to `PhotoThumbnail` now carries\
|
||
\ the real path.\n\n**Infra side-effect during the deploy:** two consecutive `--force-recreate` cycles (deploy.sh and\
|
||
\ a manual `down/up`) raced and left **orphan duplicate worker containers** (`<id>_mulita-worker-light` + the named `mulita-worker-light`,\
|
||
\ same for vision) both pulling from the same Redis queue. Worker concurrency was effectively doubled, host load climbed\
|
||
\ past 120, and `pct exec` into LXC 120 hung for several minutes. Recovery: `docker rm -f` on both orphan IDs. The deploy-gotcha\
|
||
\ entry below covers the pattern; the new contribution from this session is \"look for renamed `<id>_<service>` containers\
|
||
\ and remove them if you ever see load spike + pct hang after a deploy\".\n\n### 2026-05-11 — Drop Title field, add bulk\
|
||
\ Notes editor\n\n- Removed the `Title` (user_title) row from `PhotoInfoPanel.tsx`: the form field, its `titleDraft`/`setTitleDraft`\
|
||
\ state, the `commitTitle` handler, and the `photo?.user_title` watcher in the draft-sync `useEffect`. Column stays on\
|
||
\ `photos` and on the backend model — only the UI affordance went away.\n- New bulk action `set_notes` in `backend/app/routers/photos.py`\
|
||
\ `bulk_action`: validates the value is a string (or null/empty to clear), then sets `user_notes` on every photo in the\
|
||
\ selection in one transaction.\n- Frontend wiring: `bulkSetNotes` in `services/api.ts`, `bulkNotes` mutation in `hooks/useBulkPhotoMutations.ts`\
|
||
\ (optimistic patch with empty-string → null collapse, full rollback on error), surfaced in `RightSidebar.tsx`'s multi-photo\
|
||
\ bulk panel as a 2-row Textarea + Clear + Apply pair below the Tags section. Doesn't auto-fire on every keystroke — clicking\
|
||
\ Apply is the explicit commit (otherwise every keystroke would PATCH N rows).\n- Shared `Photo` interface in `types/photo.ts`\
|
||
\ now includes `user_notes?: string | null` so the optimistic-patch typing accepts the field; previously it only existed\
|
||
\ on `PhotoInfoPanel`'s local `PhotoDetails` shape.\n\n### 2026-05-10 — Right sidebar compact rebuild\n\nRight sidebar\
|
||
\ previously had `<ActiveHeapCard />` + `<Header />` + a scroll region holding two parallel collapsibles (\"Edit\" + \"\
|
||
Metadata\"), with three nested `<Section>` sub-collapsibles (Basic Info / Camera / Location) inside Metadata. A lot of\
|
||
\ chrome for a per-photo form.\n\n- **Removed** `frontend/src/components/heaps/ActiveHeapCard.tsx` entirely (no remaining\
|
||
\ importers; active-heap state stays in the store, Select/Discard buttons still consult it).\n- `RightSidebar.tsx` single-photo\
|
||
\ branch also drops its `<Header />` strip — the new \"METADATA\" collapsible trigger inside `PhotoInfoPanel` is the visible\
|
||
\ title. Multi-photo branch keeps the Header for \"N Photos Selected\".\n- `PhotoInfoPanel.tsx` is now two stacked collapsibles:\n\
|
||
\ - **Metadata** (default expanded): readonly fields (Size / Dimensions grid, Path, GPS inlined when present), a thin\
|
||
\ `<hr>`, then editable form (Filename, Title, Date Taken, Notes, Tags, Rating + Color on one row, Flag).\n - **Camera**\
|
||
\ (default expanded): isolated EXIF block (make+model, lens, 2×2 ISO/Aperture/Shutter/Focal grid).\n- Compact density:\
|
||
\ Notes `rows=3 → 2`, stars/swatches `h-5 → h-4`, `space-y-2.5 → 2`, Flag buttons `text-sm → text-xs`, grid `gap-2 → gap-x-2\
|
||
\ gap-y-1`, empty \"No GPS data\" chip dropped (now hidden when there are no coordinates).\n- Local `Section` helper deleted\
|
||
\ from `PhotoInfoPanel.tsx` (no longer used).\n\n`PreviewView` reuses `RightSidebar` inside its overlay aside, so the\
|
||
\ change applies in both grid and preview.\n\n### 2026-05-10 — Timeline scroll-anchor on section switch\n\nClicking a\
|
||
\ folder in the left sidebar (or any other navigation that changed `currentSection` — All Photos, Rated, Discarded, etc.)\
|
||
\ didn't reset the timeline's scroll position. If the user was deep in All Photos and clicked a folder with fewer rows,\
|
||
\ the new section loaded with the previous scroll offset preserved, often landing on empty space below the last row.\n\
|
||
\n`Timeline.tsx` already had a section-change effect that cleared selection and reset the auto-focus guard, and a separate\
|
||
\ auto-focus effect that selects `photos[0]` once the new query resolves. The latter has an \"ensure visible\" scrollTo,\
|
||
\ but it only fires if the cell is out of view AND only after the next render — by then the user has already seen the\
|
||
\ wrong scroll position. Added `parentRef.current.scrollTop = 0` synchronously inside the section-change effect so the\
|
||
\ first paint of the new section anchors at the top; the auto-focus selectPhoto then highlights photo[0] as before.\n\n\
|
||
### 2026-05-10 — Filter bar 422-on-date-from\n\nUser reported \"all filters broken, no photos shown\" right after the\
|
||
\ metadata-extraction overhaul shipped. Tracing actual API traffic on a fresh session showed `GET /api/v1/photos?per_page=100&date_from=2026-04-10&sort=taken_at&order=desc`\
|
||
\ returning **422 Unprocessable Entity**: pydantic v2's `datetime` parser rejects bare-date strings (`\"2026-04-10\"`)\
|
||
\ for `Optional[datetime]` query params.\n\nThe frontend's `filtersToParams` in `store/filterStore.ts` had been padding\
|
||
\ `date_to` with `T23:59:59` for inclusive end-of-day, but `date_from` went out as a bare `YYYY-MM-DD` — so every date-range\
|
||
\ filter request 422'd, and TanStack Query's failure-state shows an empty grid. From the user's perspective it looked\
|
||
\ like \"filters return nothing across the board\"; from the backend it was a single endpoint signature problem.\n\nFix\
|
||
\ on both sides:\n- **Frontend**: pad `date_from` with `T00:00:00` symmetric to `date_to`'s `T23:59:59`.\n- **Backend**:\
|
||
\ switch `date_from` / `date_to` to `Optional[str]` and parse with `datetime.fromisoformat` inside the handler. `fromisoformat`\
|
||
\ accepts both bare dates (→ midnight) and full ISO strings, so any older client / curl that sends a date-only value still\
|
||
\ works. Tz-aware values get coerced to naive UTC, matching the `taken_at` column shape and the same fix applied to `PATCH\
|
||
\ /photos/{id}` earlier in the session. Bad input now returns 400 with a clear message instead of pydantic's 422.\n\n\
|
||
### 2026-05-10 — Date extraction overhaul\n\nUser reported wrong \"Date Taken\" on stills (JPEG / HEIC / DNG). Tracing\
|
||
\ the pipeline showed the trusted-EXIF list at `services/metadata.py:230-244` ended in `EXIF:ModifyDate`, which is set\
|
||
\ every time a file is re-saved (Lightroom export, batch resize, EXIF strip), so any photo that lost its original capture\
|
||
\ metadata during editing was being labeled `taken_at_source='exif'` with the *edit* timestamp.\n\nPipeline rewrite:\n\
|
||
\n- New trusted-EXIF priority: `SubSecDateTimeOriginal` → `DateTimeOriginal` → `CreateDate` → `QuickTime:MediaCreateDate`\
|
||
\ → `QuickTime:CreateDate`. **`ModifyDate` removed entirely.**\n- Fall back to `guess_date_from_path()` (the same heuristic\
|
||
\ that already powers `has_date_warning`) when no trusted EXIF date is found. New `taken_at_source='path'` value with\
|
||
\ a \"PATH\" badge in the info panel; tooltip explains the date came from filename / folder rather than real EXIF.\n-\
|
||
\ Skip the date-write block entirely when `photo.taken_at_source == 'manual'` so a rescan can't clobber a user correction.\
|
||
\ (Previous behavior overwrote manual edits.)\n- `parse_exif_datetime` accepts tz-aware variants (`%z`), normalizes to\
|
||
\ naive UTC, and rejects the `0000:00:00` placeholder.\n\nBackfill: new `backfill_taken_at` celery task + `POST /api/v1/library/maintenance/backfill-taken-at`\
|
||
\ endpoint. Re-enqueues `extract_metadata` for every non-manual, non-trashed photo so the new rules apply across the existing\
|
||
\ library. Snapshot before the sweep started: 11,223 `exif` + 6,039 `filesystem` + 10 `path`. Will reshape over the next\
|
||
\ ~45 min.\n\nSide note: the default celery queue had ~209k pending tasks at the time we fired the backfill — the watcher's\
|
||
\ 5-minute restart loop (since fixed) had been re-enqueuing scans, and tasks for the 4,158 photos we hard-deleted earlier\
|
||
\ today were still sitting around. Most fail fast (`Photo not found`, ~24ms each); real work runs at ~100ms. Decided to\
|
||
\ let it drain instead of flushing — safer.\n\n### 2026-05-10 — Two cascade-delete + datetime fixes after the hard-remove\
|
||
\ shipped\n\nThe first cut of `delete_nextcloud_source_root` blew up with `folders_parent_id_fkey` violations on the first\
|
||
\ real test (`Taco and Muli - 2024 onward`, 35 folders / 4,158 photos). Two iterations to get it right:\n\n1. **NULL `parent_id`\
|
||
\ before deleting Folders.** Folders have a self-referential `parent_id` FK with no `ON DELETE` rule; postgres checks\
|
||
\ the constraint per row regardless of insertion order, so a flat `DELETE FROM folders WHERE id IN (...)` of the whole\
|
||
\ subtree fails on the parents whose children appear later in the same statement. Fixed with an `UPDATE folders SET parent_id\
|
||
\ = NULL WHERE id IN (folder_ids)` first.\n2. **Widen the NULL UPDATE to cross-source-root children.** A \"Leóns 1st Year\"\
|
||
\ SourceRoot at `.../Taco and Muli - 2024 onward/Leóns 1st Year` had its own folder rows whose `parent_id` pointed into\
|
||
\ the Taco SourceRoot's hierarchy. The first patch only NULLed `parent_id` for folders whose `id` was in the delete set;\
|
||
\ the Leóns folders weren't in that set so they kept their references and the FK still tripped. Fix: `UPDATE folders SET\
|
||
\ parent_id = NULL WHERE parent_id IN (folder_ids)` — kills any incoming reference into the delete set, internal or external.\n\
|
||
\nAfter both fixes: `DELETE /api/v1/nextcloud/source-roots/{id}` for `Taco and Muli - 2024 onward` cleared 4,158 photos\
|
||
\ and 35 folders in a single request and returned 200. UI swaps the trash icon for a `Loader2` spinner while the request\
|
||
\ is in flight (`removeRoot.isPending && removeRoot.variables === r.id`) so the multi-second cascade is visible.\n\nSeparate\
|
||
\ fix in the same session: `PATCH /api/v1/photos/{id}` returned 500 with `can't subtract offset-naive and offset-aware\
|
||
\ datetimes` when the frontend sent a tz-aware ISO string for `taken_at` (the datetime-local input is supposed to be naive\
|
||
\ but real-world locales / paste flows occasionally include `+02:00`). The DB column is `timestamp without time zone`,\
|
||
\ so asyncpg refused to bind. Normalize on the server with `astimezone(timezone.utc).replace(tzinfo=None)` if `tzinfo\
|
||
\ is not None`.\n\nDeploy infra learning: the new `--force-recreate` in `deploy.sh` does NOT reliably recreate containers\
|
||
\ on every push — saw two consecutive deploys leave the backend at the previous `StartedAt` despite a fresh image. Manual\
|
||
\ `docker compose up -d --no-deps --force-recreate <svc>` after deploy still races with whatever compose state the auto-deploy\
|
||
\ left mid-flight, twice landing the stack in a half-broken state (orphaned `<id>_mulita-backend` rename containers, db\
|
||
\ / redis stopped). Recovery: `docker compose down && docker compose up -d`. Open question — `--force-recreate` may need\
|
||
\ to be replaced with something more explicit.\n\n### 2026-05-10 — Hard-remove Nextcloud SourceRoot + reliable delete\
|
||
\ sync\n\n`DELETE /api/v1/nextcloud/source-roots/{id}` was a soft-deactivate (`is_active=false`) — the trash icon in Settings\
|
||
\ only hid the SourceRoot from active queries while leaving every Folder + Photo row in the DB forever. Re-adding the\
|
||
\ same path resurrected ghosts; `prune-missing` reported zero deletes for the soft-removed entry because the cleanup code\
|
||
\ skipped inactive source roots.\n\nEndpoint now hard-deletes the SourceRoot, every Folder under it (chunked at 500),\
|
||
\ and every Photo in those folders. `photo_tags` and `heap_photos` cascade automatically via `ON DELETE CASCADE` on the\
|
||
\ join tables. `FolderShare` uses a stringly-typed `folder_id` (no FK) so cleaned manually for both `folder_type='folder'`\
|
||
\ and `folder_type='source_root'`. Returns `{deleted_photos, deleted_folders}` so the UI can toast a count. Files in Nextcloud\
|
||
\ are untouched.\n\nSync side: the `watch_folders` celery task in `backend/app/tasks/scan.py` already detected filesystem\
|
||
\ deletions and soft-marked photos `is_discarded=true`, but the global `task_soft_time_limit=300` in `app/tasks/celery.py`\
|
||
\ was killing the watcher every five minutes and dropping every FS event during the restart window. The `soft_time_limit=None`\
|
||
\ on the decorator was being interpreted as \"use worker default\" rather than \"unlimited\". Override to `soft_time_limit=0,\
|
||
\ time_limit=0` (Celery convention for unbounded). Backstop: a new `discard_missing_photos()` in `app/services/cleanup.py`\
|
||
\ — soft variant of `prune_missing_photos`, walks every `_sr_state == 'present'` source root and flips `is_discarded=true`\
|
||
\ on Photo rows whose file is gone. Wired as `discard_missing_photos_beat` celery task scheduled every 30 minutes via\
|
||
\ `beat_schedule` on the celery app. Beat runs in-process on `worker-watcher` (`--beat` flag in `docker-compose.yml`)\
|
||
\ — the watcher is already a Redis-locked singleton so no need for a separate beat container.\n\nManual `POST /api/v1/library/maintenance/prune-missing`\
|
||
\ remains the hard-delete path for when the user wants to permanently drop orphan rows; the new beat job only soft-discards\
|
||
\ (file gone → mule-image trash, restorable).\n\n### 2026-05-10 — OIDC auto-redirect on LoginPage\n\n`OIDC_ENABLED=true`\
|
||
\ was already set in `.env`, so the LoginPage rendered a \"Sign in with Authentik\" button next to the password form.\
|
||
\ With a single trusted IdP and a logged-in Authentik session, that extra click was friction without upside.\n\n`LoginPage.tsx`\
|
||
\ now reads `/auth/config` on mount and, if OIDC is on, immediately navigates to the OIDC login URL. Authentik recognizes\
|
||
\ the existing session and bounces back through the callback with no user interaction. Two escape hatches: `?password=1`\
|
||
\ in the URL forces the password form, and a `skipAutoSso` sessionStorage flag (set by `AuthContext.logout` and by the\
|
||
\ OIDC callback's error branch) suppresses the next auto-redirect so logouts actually log out and OIDC failures surface\
|
||
\ their error instead of looping. While the redirect is in flight the page shows \"Signing in with Authentik...\" plus\
|
||
\ a \"Use password instead\" link.\n\n### 2026-05-10 — right sidebar restructure (heap pinned, single scroll, collapsible\
|
||
\ Metadata)\n\nThe right sidepanel had three stacked flex regions: `<ActiveHeapCard />`, `<Header />`, and `<PhotoInfoPanel\
|
||
\ />` — with `PhotoInfoPanel` carrying its own internal scroll. That left the editable fields (filename, title, notes,\
|
||
\ rating, color, flag) stuck above the readonly metadata scroll, effectively two scroll boundaries on one sidebar.\n\n\
|
||
- Moved the scroll boundary up to `RightSidebar.tsx`: only `ActiveHeapCard` and `Header` stay pinned now. Editable + readonly\
|
||
\ content scroll together in one `flex-1 overflow-y-auto` region beneath them.\n- `PhotoInfoPanel.tsx` dropped its `h-full`\
|
||
\ / inner `flex-1 overflow-y-auto`.\n- The four readonly sections (Tags / Basic Info / Camera / Location) are now wrapped\
|
||
\ in a single outer \"Metadata\" `Collapsible`. Default expanded, one click hides the whole block. Sub-sections stay individually\
|
||
\ collapsible.\n- Second pass: the editable form (filename / title / notes / rating / color / flag) got the same treatment\
|
||
\ under an outer \"Edit\" `Collapsible` so the panel is now two equal collapsible groups below the title strip. Dropped\
|
||
\ the X (clear-selection) button from the Header; Esc and grid-empty-area-click still clear.\n- Third pass: split editable\
|
||
\ vs read-only between the two groups consistently. `TagsEditor` and `TakenAtEditor` had been buried inside the readonly\
|
||
\ Metadata sub-sections — `Tags` as its own Section, taken-at wedged into `Basic Info` between size/dims and the filepath.\
|
||
\ Moved both into the Edit collapsible, ordered identification → description → categorization: Filename · Title · Date\
|
||
\ Taken · Notes · Tags · Rating · Color · Flag. Metadata now holds only readonly: `Basic Info` (size, dims, path), `Camera`,\
|
||
\ `Location`.\n\n`PreviewView` reuses `RightSidebar` under an `<aside class=\"overflow-hidden\">`, so the change applies\
|
||
\ in both the grid sidebar and the preview overlay.\n\n### 2026-05-10 — `photos.hubris.network` perf sweep\n\nUser reported\
|
||
\ slow load. Five fixes shipped, in order:\n\n1. **Backend out of dev mode.** `docker-compose.yml` `command:` was running\
|
||
\ `uvicorn … --reload` (single worker, file-watcher). Page loads fan out ~15 parallel API calls; they all serialized on\
|
||
\ one event loop. → `--workers 2 --proxy-headers`. Two uvicorn worker procs now.\n2. **iPhone Apple ProRAW / Linear DNG\
|
||
\ decode fixed.** rawpy 0.26.1 + LibRaw 0.22 rejects Apple Linear DNG (`Photometric Interpretation: Linear Raw`, 8-bit)\
|
||
\ as \"Unsupported file format or not RAW file\". Those files are TIFF containers with developed RGB inside and have no\
|
||
\ embedded preview to fall back to. Added PIL `Image.open(src_path)` as the last fallback in both `_generate_proxy_webp`\
|
||
\ (routers/photos.py) and `process_raw_image` (tasks/thumbs.py). ~1,300 iPhone DNGs were 415-ing on every detail view;\
|
||
\ now decode in <1s via PIL.\n3. **Reconcile DB with disk (renamed-folder case).** `prune_missing_photos` was skipping\
|
||
\ all 4,154 orphaned photo rows under SourceRoot `Taco and Muli - 2024 onward` (renamed in Nextcloud to `Photo Archive\
|
||
\ 2004-2024`) because the leaf path didn't resolve and the code assumed \"drive unmounted → must skip\". Added `_sr_state()`\
|
||
\ to classify as `present` / `renamed` (parent mount fine, leaf gone) / `unmounted` (parent inaccessible). Only `unmounted`\
|
||
\ still skips. Two stale source roots logged with a clearer hint pointing at `POST /api/v1/library/maintenance/prune-missing`.\
|
||
\ User has not yet been asked to run that — endpoint is ready when they are.\n4. **Frontend page size + idle polling.**\
|
||
\ `usePhotosQuery.ts` was first-fetching `per_page=500` (slow paint + 500 thumb requests at once). Split into `PER_PAGE_INITIAL=100`\
|
||
\ for first paint, `PER_PAGE_BACKGROUND=500` for the cursor-chain prefetch. Idle polling for `scan-status` and `worker-status`\
|
||
\ (`useScanActivity.ts`, `ScanProgress.tsx`) bumped from 10s/15s to 30s/30s while idle; active cadence (2s/3s) unchanged.\n\
|
||
5. **Partial index on `photos`.** Default list query `WHERE NOT is_trashed AND NOT is_hidden ORDER BY taken_at DESC NULLS\
|
||
\ LAST, id DESC LIMIT N` was doing a seq-scan + top-N heapsort (~25ms standalone, worse under concurrency). Added migration\
|
||
\ `0017_photos_list_index` creating `ix_photos_list_visible` (partial index on the sort key, restricted to visible rows).\
|
||
\ EXPLAIN now shows an Index Only Scan → **24.7ms → 0.097ms**, ~250× speedup.\n\n**Deploy gotcha — fixed (with caveat).**\
|
||
\ The original `docker compose up -d --build` in `deploy.sh` did not reliably recreate containers when only runtime config\
|
||
\ (`command:`, env-only) or migration files changed; image hash would change but compose treated the existing container\
|
||
\ as \"current enough\". Bit three times this session before `/opt/mule-deploy/deploy.sh` was updated to `docker compose\
|
||
\ up -d --build --force-recreate`. Trade-off accepted: an extra restart cycle on deploys where nothing user-visible changed.\n\
|
||
\n**Caveat:** the first auto-deploy after the flag flip raced with my own earlier manual `docker compose up -d --no-deps\
|
||
\ --force-recreate frontend` and landed the stack in a half-broken state — `mulita-frontend` got stuck under a renamed\
|
||
\ temp container, several services dropped off `mule-image_mulita-network`, frontend nginx restarted in a loop with `host\
|
||
\ not found in upstream \"backend\"`. Fixed by `docker compose down && docker compose up -d`. Don't issue a manual `--force-recreate`\
|
||
\ on a single service while the auto-deploy webhook is also expected to fire — let the deploy own the lifecycle.\n\n**Data\
|
||
\ drift still outstanding.** 4,154 photo rows + 1 unregistered folder (`Photo Archive 2004-2024`) on disk that's not a\
|
||
\ SourceRoot. The reconcile endpoints now work — user decides when to call them. The new folder needs to be added as a\
|
||
\ SourceRoot via the Settings UI before its files will be indexed.\n\n**Proxy cache still empty** (`mule-image_proxies_data`\
|
||
\ volume is 4 KB). Pre-generating ~500–1,500 WebP proxies for non-web-safe formats would make first-open of every RAW/HEIC\
|
||
\ photo instant. Deferred — needs a one-shot script and the disk-space tradeoff isn't worth it until the data-drift reconcile\
|
||
\ happens first.\n\n### 2026-04-28 — wiki entry created\nInitial documentation.\n\n### 2026-04-26 — Nextcloud-rooted libraries\
|
||
\ shipped\nBind `/mnt/library/homecloud` into the workers, reads via filesystem, writes via WebDAV. `users.nextcloud_username`\
|
||
\ override field added; `dtoro → admin` mapping. Surviving SourceRoots cleaned up to NC paths.\n\n### 2026-04-22 — native\
|
||
\ OIDC via Authentik\nAuthlib-based code in `backend/app/auth_oidc.py`. `extra_hosts` override for `auth.hubris.network`\
|
||
\ in compose override (gitignored).\n\n### 2026-04-21 — auto-deploy pipeline shipped\nWebhook receiver at `:9797`, async\
|
||
\ deploy returning 202. Mirrors caddy-conf / gitea-customizations.\n"
|
||
entity_slug: lxc:mule-images
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: mule-images
|
||
ip: 192.168.8.136
|
||
privileged: privileged
|
||
resources: 6 cores / 12 GiB RAM / 60 GiB rootfs / 2 GiB swap
|
||
mounts: /mnt/library /mnt/library; /dev/dri/{card0,renderD128} passed through for VA-API video accel on the AMD Phoenix1
|
||
iGPU.
|
||
public_host: photos.hubris.network → caddy → path-routed to PhotoPrism :2342 / sidecar :8000 / nginx-static :3000.
|
||
changelog:
|
||
- date: '2026-05-22'
|
||
title: Cutover to PhotoPrism stack (Mulimage 2.0)
|
||
body: 'The `new` branch of `dtoro/mule-image` was merged into `main` as commit `70dc1b6`. The merge replaces the legacy
|
||
FastAPI + PostgreSQL + Celery + React stack with PhotoPrism + Go sidecar + MariaDB + SvelteKit, lifted in from the M0
|
||
evaluation on (now-destroyed) [LXC 127](#). Cutover on 120 (this session): - Bumped to 6 cores / 2 GB swap. `/dev/dri/{card0,renderD128}`
|
||
already passed through; added an opt-in `docker-compose.gpu.yml` overlay that the deploy script layers in when the device
|
||
is present. PhotoPrism now uses `h264_vaapi` instead of `libx264`. - Installed host nginx + a small `photos.conf` vhost
|
||
serving the SvelteKit static bundle from `/opt/mule-image/web/build/` on `:3000` — no Vite dev server in prod. SvelteKit''s
|
||
`adapter-static` produces a real SPA bundle. - Replaced `/opt/mule-deploy/deploy.sh` with the 127-style multi-compose
|
||
deploy (force-recreates only sidecar to preserve PhotoPrism''s Session HMAC key; reconciles pp-app + mariadb in place;
|
||
rebuilds web/ and reloads nginx). - `/opt/mule-image/.env` was rewritten to the PhotoPrism shape. Reused passwords from
|
||
`/root/mule-photos-new-secrets.txt`. `USER_BASEPATHS= admin:admin/files, muli:muli/files` exposes both libraries. -
|
||
`docker-compose.override.yml` (untracked) `!override`s pp-app + sidecar ports to `0.0.0.0` so cross-LXC caddy on [121](121-caddy.md)
|
||
can reach them. - Authentik: `mule-image` app''s redirect URI updated to `/api/v1/oidc/redirect`; the separate `mule-photos-new`
|
||
app deleted. - Caddyfile (`dtoro/caddy-conf`): `photos.hubris.network` switched from plain passthrough to path-matchers
|
||
(PhotoPrism API + sidecar API + nginx static catch-all), and `photos-new.hubris.network` was removed entirely. - Cleanups:
|
||
4 Nextcloud `oc_webhook_listeners` rows for the legacy `/api/v1/internal/nc-webhook` endpoint deleted; gitea webhook
|
||
id 9 (the `refs/heads/new` deploy hook for 127) deleted; `192.168.8.181` removed from gitea''s `ALLOWED_HOST_LIST`;
|
||
`photos-new.hubris.network` removed from dnsmasq. - LXC 127 destroyed via `pct destroy 127` after the new stack passed
|
||
curl verification end-to-end (PP `/api/v1/status` returns `{"status":"operational"}` through caddy; sidecar `/api/sidecar/healthz`
|
||
returns `{"ok":true,...}`). - Rollback safety net: legacy `mulita` postgres dump at `/root/backups/mulita-legacy-final-20260521-212036.sql.gz`
|
||
on hubris (10 MB, 16,155 photo rows). The `mule-image_mule_db_data` docker volume on 120 retains the on-disk postgres
|
||
data for at least 24 h before housekeeping removes it.'
|
||
- date: '2026-05-14'
|
||
title: Drop AI/vision pipeline, plain Postgres, DB↔FS refresh
|
||
body: '**AI removal (commits `a27267f` + `423a73a`).** The OpenCLIP-on-ONNX binary classifier (photography vs other) and
|
||
all its scaffolding gone: - Backend: `app/services/vision/`, `app/tasks/vision.py`, `app/services/feature_flags.py`,
|
||
`app/routers/features.py` deleted. Admin `/admin/feature-flags`, `/admin/ai/{backfill,rescan}` endpoints removed. `Photo.needs_review`
|
||
column + `ix_photos_needs_review` index dropped (alembic `0019_drop_ai_remnants`). - Frontend: AI Settings tab, `useFeaturesQuery`,
|
||
"Needs Review" sidebar entry + filter, `needs_review` URL param plumbing, FeatureFlag types all gone. - Infra: `worker-vision`
|
||
compose service + `models_data` volume deleted. `worker-light` no longer runs `python -m app.services.vision.bootstrap_models`
|
||
before celery. Backend Dockerfile drops the dedicated torch RUN layer. `requirements.txt` drops `torch`, `torchvision`,
|
||
`open-clip-torch`, `onnxruntime`. **Postgres image swap.** `pgvector/pgvector:pg16` → `postgres:16`. The 0019 migration
|
||
drops the (now-unused) `vector` extension first; the SQL was pre-applied via `psql` against the still-pgvector container
|
||
and `alembic_version` stamped to `0019` so the new backend''s `bootstrap.py` upgrade-head was a no-op. After the swap
|
||
surfaced a collation-version mismatch (Bookworm glibc 2.36 → Trixie glibc 2.41), the fix was `REINDEX DATABASE mulita`
|
||
+ `ALTER DATABASE … REFRESH COLLATION VERSION` on `mulita`, `postgres`, `template1`. The data volume was preserved across
|
||
the image change. **One-shot DB ↔ filesystem refresh.** New script `backend/scripts/full_refresh.py` (run as `docker
|
||
exec -w /app mulita-backend python -m scripts.full_refresh [--dry-run]`). Phases: `cleanup_data_integrity` → inline
|
||
scan of every active SourceRoot → `prune_missing_photos(dry_run=False)` → new `prune_orphan_thumbnails` helper that
|
||
removes `/data/thumbs/{user}/{photo}/` dirs for photo_ids that left the DB. First run: 0 missing photos, 1 stale folder
|
||
row, **7982 orphan thumbnail dirs** reaped. **Manual cleanup post-deploy.** `docker compose up -d --build --force-recreate`
|
||
doesn''t reap services removed from the YAML, so `mulita-worker-vision` had to be `docker rm -f`''d by hand. The `mule-image_models_data`
|
||
volume was likewise removed manually.'
|
||
- date: '2026-05-11'
|
||
title: Stop duplicating Nextcloud's indexing (Phase 1 + 2)
|
||
body: 'Two big architectural shifts in one session, both aimed at killing work mule was doing that NC already does for
|
||
the same source files. **Phase 1 — thumbnail proxy (commits `576b0c2`, `28738ac`).** Photos table gains `nextcloud_fileid`
|
||
(migration `0018`). `GET /api/v1/photos/{id}/thumb/{size}` proxies NC''s `/index.php/core/preview` keyed by that fileid,
|
||
authenticated with the owner''s encrypted app password (small=240, medium=640, large=1280). Worker now only writes the
|
||
`medium` size to `/data/thumbs` (the vision worker still loads it from disk); `small` + `large` no longer touch disk.
|
||
Disk fallback stays wired for legacy rows and the rare 404 from NC (iPhone JPEGs mis-extensioned as `.DNG` — verified).
|
||
Existing 17,591 photos backfilled via `backend/scripts/backfill_nextcloud_fileid.py`. Tested with real CR2, real DNG,
|
||
fake DNG, JPEG, HEIC — all green. **Range support for `/original` (commit `18dce33`).** Pre-existing bug surfaced by
|
||
Phase 1 testing: `<video>` couldn''t play `.mov` because `FileResponse` was returning 200 with the whole body and no
|
||
`Accept-Ranges` header, so browsers reported "format not supported." Now parses `Range: bytes=START-END`, returns 206
|
||
with `Content-Range`, streams in 1 MB chunks. **Fix: backfill_gps was the actual CPU drain, not the watcher (commit
|
||
`d24c64e`).** `_scan_all_source_roots_async` auto-queued `backfill_gps` 30 s after every container boot, which then
|
||
re-queued `extract_metadata` for every photo with `latitude IS NULL` — ~60k pointless tasks per deploy, pinning worker-light
|
||
at 180+% CPU for ~30 min. Killed the auto-trigger; manual `POST /api/v1/library/backfill-gps` still works. **Phase 2
|
||
— webhook receiver replaces watchfiles (commits `362fbc6`, `f657e2c`).** New endpoint `POST /api/v1/internal/nc-webhook`
|
||
(auth via `Authorization: Bearer $NEXTCLOUD_WEBHOOK_SECRET`, `hmac.compare_digest`) handles `NodeCreated|NodeWritten|NodeDeleted|NodeRenamed`.
|
||
Maps NC''s `/admin/files/...` path to the bind-mounted `/nextcloud-users/admin/files/...` and dispatches the same `scan_folder`
|
||
/ `handle_file_deletion` machinery the watcher used. Registered against NC via `backend/scripts/register_nc_webhooks.py`
|
||
(idempotent: deletes existing webhooks targeting the same URL first). `mulita-worker-watcher` container retired; celery
|
||
`--beat` folded into `worker-light` so `discard_missing_photos_beat` still fires. End-to-end verified by uploading a
|
||
test JPEG through WebDAV — both `NodeCreated` and `NodeWritten` fire, mule receives + dispatches, 200 OK. **NC cron
|
||
tightened to `*/1` so webhook delivery latency drops from up to 5 min to ~60 s** (`crontab -u www-data` on LXC 114).
|
||
NC dispatches webhook calls through its background-job queue; the cron interval = the worst-case latency. `*/5` was
|
||
the default; `*/1` is the upstream recommendation anyway. **Post-Phase-2 delete-roundtrip patches (commit `9408825`).**
|
||
End-to-end testing of the NC↔mule deletion paths surfaced two real gaps: 1. *Folder delete* — NC fires one `NodeDeletedEvent`
|
||
for the folder, not one per child. The webhook handler bailed with "unsupported extension" and photos under the deleted
|
||
folder kept `is_discarded=false` until the 30-min reconcile sweep. Fix: `handle_directory_deletion()` does a single
|
||
`UPDATE photos SET is_discarded=true WHERE filepath LIKE ''dir/%''` when the deleted path has no supported image extension.
|
||
2. *Resurrect on rewrite* — PUT-overwrite of a previously-discarded file fired `NodeWrittenEvent` → `scan_folder`, but
|
||
scan_folder''s "photo exists, skip" branch left `is_discarded=true`. Fix: when the existing row is discarded, flip `is_discarded=false`,
|
||
clear `discarded_at`, re-queue `extract_metadata`. Verified end-to-end on `Photos/MuleTestFolder/{test1,test2,nested/test3}.jpg`:
|
||
`DELETE Photos/MuleTestFolder/` discarded all 3 in one shot; `PUT test1.jpg` back resurrected only that one. **Known
|
||
remaining gap — trashbin restore.** Moving a file out of `/dav/trashbin/.../trash/foo.jpg.dXXXX` back to `/files/...`
|
||
via WebDAV MOVE fires **no event mule subscribes to**. NC''s trashbin app emits its own internal event class that isn''t
|
||
in the `OCP\Files\Events\Node\*` set we registered with `webhook_listeners`. Workaround: re-upload via PUT (covered
|
||
by the resurrect-on-rewrite fix above) or wait for the 30-min reconcile sweep. **Folder rename round-trip (commits `f4a03b6`
|
||
+ `f27f3cb`).** Both directions now work: - *NC-side rename → mule:* webhook''s "renamed" branch detects directory rename
|
||
(neither path has a supported extension) and calls new `handle_directory_rename(old, new)` in `scan.py`. The helper
|
||
iterates matching rows in Python and prefix-rewrites `Photo.filepath`, `Folder.path`, `SourceRoot.path` in one transaction.
|
||
Cross-source-root case (rare) discards the old subtree and lets scan_folder add fresh rows under the new root. - *Mule-side
|
||
rename → NC:* the existing `PATCH /api/v1/folders/{id}` endpoint (`folders.py:202`) already does WebDAV MOVE via `nextcloud_dav.move_for_user`
|
||
and rewrites mule''s DB inline. The NodeRenamedEvent that bounces back through the webhook hits `handle_directory_rename`,
|
||
which finds 0 rows under the old prefix and is a no-op — feedback loop is idempotent. **asyncpg gotcha**: the original
|
||
implementation of `handle_directory_rename` used raw SQL with `SUBSTRING(filepath FROM LENGTH(:old_prefix) + 1)` so
|
||
the offset would be computed server-side. asyncpg''s type inference miscategorises the `LENGTH()` result and rejects
|
||
the parameter as "$2: int (expected str)". The fix iterates in Python (same pattern as the existing PATCH endpoint).
|
||
Lesson: avoid passing `LENGTH(:x)+1` as an argument to `SUBSTRING(...)` via asyncpg + sqlalchemy `text()`. **Phase 3
|
||
— Memories-backed `extract_metadata` (commit `2a5759c`).** Memories app re-enabled on NC. `extract_metadata` now tries
|
||
`GET /index.php/apps/memories/api/image/info/{fileid}` (auth: Basic + `OCS-APIRequest: true` header to bypass CSRF)
|
||
before falling back to ExifTool. Replaces ~80 ms of subprocess with ~1-2 ms HTTP for ongoing imports. New helpers: `nextcloud_dav.fetch_memories_info_async()`
|
||
+ `metadata._apply_memories_metadata()`. We kept mule''s full date-fallback chain (SubSec → DateTimeOriginal → CreateDate
|
||
→ MediaCreateDate → TrackCreateDate → filename heuristic → mtime) because 35% of the library (`taken_at_source=''path''`)
|
||
depends on the filename heuristic, and Memories alone would silently mis-date those photos to mtime. PhotoInfoPanel
|
||
reads `exif.Make/Model/ISO/FNumber` from `photos.exif_json` — Memories'' `exif` blob uses those exact plain key names,
|
||
so no frontend adapter was needed. ExifTool subprocess is still in place as the fallback for brand-new photos racing
|
||
the NC scan, non-NC photos, and any NC HTTP failure. **Phase 4 — retire `/data/thumbs` (commits `5a67ed7` + `7a1c6b6`).**
|
||
Vision worker now fetches NC''s 640px preview via a new sync helper `nextcloud_dav.get_preview_bytes()` instead of reading
|
||
`/data/thumbs/{id}/medium.webp`. `thumbs.WORKER_THUMB_SIZES = set()` — `generate_thumbnails` still computes pHash on
|
||
the original-res pixels (perceptual dedup is mule-only) but stops writing files. All ~22k `medium.webp` purged after
|
||
verification; `/data/thumbs` shrank from 4.1 GB → 94 MB residual. **SECRET_KEY bug found mid-deploy (commit `7a1c6b6`).**
|
||
`docker-compose.yml` only set `SECRET_KEY` on the `backend` service, not on the workers. Workers'' Fernet-based `decrypt(nextcloud_app_password_enc)`
|
||
silently returned empty and `_credentials_for()` raised `NextcloudCredentialsMissing`. This meant Phase 3''s `extract_metadata`
|
||
in worker-light had been silently falling back to ExifTool *the entire time*, and Phase 4''s vision worker couldn''t
|
||
fetch NC previews at all. Fix replicates `SECRET_KEY=${SECRET_KEY:-...}` to all worker services. After any compose edit,
|
||
run `docker exec mulita-<svc> env | grep SECRET_KEY` to confirm propagation. **Caveats worth knowing for Phase 4:**
|
||
- Vision is `disabled` in production (`redis GET mulita:flags:vision.enabled == "false"`). The Phase 4 vision-from-NC
|
||
path is correct but unexercised by live traffic; it''ll matter whenever vision is re-enabled. - NC''s preview generator
|
||
hasn''t covered the whole library at the 640px+ tier. A sample of 6 photos: 4 with proper sizes, 1 stuck at 160px (`IMG_4954.DNG`),
|
||
1 unfetchable (`IMG_0193.DNG`). If vision re-enables and quality matters, run `pct exec 114 -- sudo -u www-data php
|
||
occ preview:generate-all` to backfill. - The HTTP `/thumb/{size}` endpoint''s inline-regen disk fallback still writes
|
||
one WebP per NC-404 event (mis-extensioned RAW etc). `/data/thumbs` will grow back very slowly from that path.'
|
||
- date: '2026-05-11'
|
||
title: Duplicates view shows parent folder + full-path tooltip
|
||
body: '`GET /api/v1/library/duplicates/groups` now includes `filepath` in each member payload. `DuplicatesView` renders
|
||
a dark bottom-strip on every thumbnail showing the last two folder segments — `…/<grandparent>/<parent>` (e.g. `…/files/Phone
|
||
Photos starting 01-26`) via a `duplicatePathLabel` helper. Going up two levels avoids the failure mode where two copies
|
||
sit under matching parent names (e.g. duplicate `2023/` subfolders under different archives) and the label would collapse.
|
||
Full filepath surfaces through the native `title` tooltip on hover. Dimensions chip moves from bottom-left to top-left
|
||
so the path strip can run edge-to-edge. `memberToPhoto` finally stops faking `filepath = filename` — the synthetic Photo
|
||
it hands to `PhotoThumbnail` now carries the real path. **Infra side-effect during the deploy:** two consecutive `--force-recreate`
|
||
cycles (deploy.sh and a manual `down/up`) raced and left **orphan duplicate worker containers** (`<id>_mulita-worker-light`
|
||
+ the named `mulita-worker-light`, same for vision) both pulling from the same Redis queue. Worker concurrency was effectively
|
||
doubled, host load climbed past 120, and `pct exec` into LXC 120 hung for several minutes. Recovery: `docker rm -f`
|
||
on both orphan IDs. The deploy-gotcha entry below covers the pattern; the new contribution from this session is "look
|
||
for renamed `<id>_<service>` containers and remove them if you ever see load spike + pct hang after a deploy".'
|
||
- date: '2026-05-11'
|
||
title: Drop Title field, add bulk Notes editor
|
||
body: '- Removed the `Title` (user_title) row from `PhotoInfoPanel.tsx`: the form field, its `titleDraft`/`setTitleDraft`
|
||
state, the `commitTitle` handler, and the `photo?.user_title` watcher in the draft-sync `useEffect`. Column stays on
|
||
`photos` and on the backend model — only the UI affordance went away. - New bulk action `set_notes` in `backend/app/routers/photos.py`
|
||
`bulk_action`: validates the value is a string (or null/empty to clear), then sets `user_notes` on every photo in the
|
||
selection in one transaction. - Frontend wiring: `bulkSetNotes` in `services/api.ts`, `bulkNotes` mutation in `hooks/useBulkPhotoMutations.ts`
|
||
(optimistic patch with empty-string → null collapse, full rollback on error), surfaced in `RightSidebar.tsx`''s multi-photo
|
||
bulk panel as a 2-row Textarea + Clear + Apply pair below the Tags section. Doesn''t auto-fire on every keystroke —
|
||
clicking Apply is the explicit commit (otherwise every keystroke would PATCH N rows). - Shared `Photo` interface in
|
||
`types/photo.ts` now includes `user_notes?: string | null` so the optimistic-patch typing accepts the field; previously
|
||
it only existed on `PhotoInfoPanel`''s local `PhotoDetails` shape.'
|
||
- date: '2026-05-10'
|
||
title: Right sidebar compact rebuild
|
||
body: 'Right sidebar previously had `<ActiveHeapCard />` + `<Header />` + a scroll region holding two parallel collapsibles
|
||
("Edit" + "Metadata"), with three nested `<Section>` sub-collapsibles (Basic Info / Camera / Location) inside Metadata.
|
||
A lot of chrome for a per-photo form. - **Removed** `frontend/src/components/heaps/ActiveHeapCard.tsx` entirely (no
|
||
remaining importers; active-heap state stays in the store, Select/Discard buttons still consult it). - `RightSidebar.tsx`
|
||
single-photo branch also drops its `<Header />` strip — the new "METADATA" collapsible trigger inside `PhotoInfoPanel`
|
||
is the visible title. Multi-photo branch keeps the Header for "N Photos Selected". - `PhotoInfoPanel.tsx` is now two
|
||
stacked collapsibles: - **Metadata** (default expanded): readonly fields (Size / Dimensions grid, Path, GPS inlined
|
||
when present), a thin `<hr>`, then editable form (Filename, Title, Date Taken, Notes, Tags, Rating + Color on one row,
|
||
Flag). - **Camera** (default expanded): isolated EXIF block (make+model, lens, 2×2 ISO/Aperture/Shutter/Focal grid).
|
||
- Compact density: Notes `rows=3 → 2`, stars/swatches `h-5 → h-4`, `space-y-2.5 → 2`, Flag buttons `text-sm → text-xs`,
|
||
grid `gap-2 → gap-x-2 gap-y-1`, empty "No GPS data" chip dropped (now hidden when there are no coordinates). - Local
|
||
`Section` helper deleted from `PhotoInfoPanel.tsx` (no longer used). `PreviewView` reuses `RightSidebar` inside its
|
||
overlay aside, so the change applies in both grid and preview.'
|
||
- date: '2026-05-10'
|
||
title: Timeline scroll-anchor on section switch
|
||
body: Clicking a folder in the left sidebar (or any other navigation that changed `currentSection` — All Photos, Rated,
|
||
Discarded, etc.) didn't reset the timeline's scroll position. If the user was deep in All Photos and clicked a folder
|
||
with fewer rows, the new section loaded with the previous scroll offset preserved, often landing on empty space below
|
||
the last row. `Timeline.tsx` already had a section-change effect that cleared selection and reset the auto-focus guard,
|
||
and a separate auto-focus effect that selects `photos[0]` once the new query resolves. The latter has an "ensure visible"
|
||
scrollTo, but it only fires if the cell is out of view AND only after the next render — by then the user has already
|
||
seen the wrong scroll position. Added `parentRef.current.scrollTop = 0` synchronously inside the section-change effect
|
||
so the first paint of the new section anchors at the top; the auto-focus selectPhoto then highlights photo[0] as before.
|
||
- date: '2026-05-10'
|
||
title: Filter bar 422-on-date-from
|
||
body: 'User reported "all filters broken, no photos shown" right after the metadata-extraction overhaul shipped. Tracing
|
||
actual API traffic on a fresh session showed `GET /api/v1/photos?per_page=100&date_from=2026-04-10&sort=taken_at&order=desc`
|
||
returning **422 Unprocessable Entity**: pydantic v2''s `datetime` parser rejects bare-date strings (`"2026-04-10"`)
|
||
for `Optional[datetime]` query params. The frontend''s `filtersToParams` in `store/filterStore.ts` had been padding
|
||
`date_to` with `T23:59:59` for inclusive end-of-day, but `date_from` went out as a bare `YYYY-MM-DD` — so every date-range
|
||
filter request 422''d, and TanStack Query''s failure-state shows an empty grid. From the user''s perspective it looked
|
||
like "filters return nothing across the board"; from the backend it was a single endpoint signature problem. Fix on
|
||
both sides: - **Frontend**: pad `date_from` with `T00:00:00` symmetric to `date_to`''s `T23:59:59`. - **Backend**: switch
|
||
`date_from` / `date_to` to `Optional[str]` and parse with `datetime.fromisoformat` inside the handler. `fromisoformat`
|
||
accepts both bare dates (→ midnight) and full ISO strings, so any older client / curl that sends a date-only value still
|
||
works. Tz-aware values get coerced to naive UTC, matching the `taken_at` column shape and the same fix applied to `PATCH
|
||
/photos/{id}` earlier in the session. Bad input now returns 400 with a clear message instead of pydantic''s 422.'
|
||
- date: '2026-05-10'
|
||
title: Date extraction overhaul
|
||
body: 'User reported wrong "Date Taken" on stills (JPEG / HEIC / DNG). Tracing the pipeline showed the trusted-EXIF list
|
||
at `services/metadata.py:230-244` ended in `EXIF:ModifyDate`, which is set every time a file is re-saved (Lightroom
|
||
export, batch resize, EXIF strip), so any photo that lost its original capture metadata during editing was being labeled
|
||
`taken_at_source=''exif''` with the *edit* timestamp. Pipeline rewrite: - New trusted-EXIF priority: `SubSecDateTimeOriginal`
|
||
→ `DateTimeOriginal` → `CreateDate` → `QuickTime:MediaCreateDate` → `QuickTime:CreateDate`. **`ModifyDate` removed entirely.**
|
||
- Fall back to `guess_date_from_path()` (the same heuristic that already powers `has_date_warning`) when no trusted
|
||
EXIF date is found. New `taken_at_source=''path''` value with a "PATH" badge in the info panel; tooltip explains the
|
||
date came from filename / folder rather than real EXIF. - Skip the date-write block entirely when `photo.taken_at_source
|
||
== ''manual''` so a rescan can''t clobber a user correction. (Previous behavior overwrote manual edits.) - `parse_exif_datetime`
|
||
accepts tz-aware variants (`%z`), normalizes to naive UTC, and rejects the `0000:00:00` placeholder. Backfill: new `backfill_taken_at`
|
||
celery task + `POST /api/v1/library/maintenance/backfill-taken-at` endpoint. Re-enqueues `extract_metadata` for every
|
||
non-manual, non-trashed photo so the new rules apply across the existing library. Snapshot before the sweep started:
|
||
11,223 `exif` + 6,039 `filesystem` + 10 `path`. Will reshape over the next ~45 min. Side note: the default celery queue
|
||
had ~209k pending tasks at the time we fired the backfill — the watcher''s 5-minute restart loop (since fixed) had been
|
||
re-enqueuing scans, and tasks for the 4,158 photos we hard-deleted earlier today were still sitting around. Most fail
|
||
fast (`Photo not found`, ~24ms each); real work runs at ~100ms. Decided to let it drain instead of flushing — safer.'
|
||
- date: '2026-05-10'
|
||
title: Two cascade-delete + datetime fixes after the hard-remove shipped
|
||
body: 'The first cut of `delete_nextcloud_source_root` blew up with `folders_parent_id_fkey` violations on the first real
|
||
test (`Taco and Muli - 2024 onward`, 35 folders / 4,158 photos). Two iterations to get it right: 1. **NULL `parent_id`
|
||
before deleting Folders.** Folders have a self-referential `parent_id` FK with no `ON DELETE` rule; postgres checks
|
||
the constraint per row regardless of insertion order, so a flat `DELETE FROM folders WHERE id IN (...)` of the whole
|
||
subtree fails on the parents whose children appear later in the same statement. Fixed with an `UPDATE folders SET parent_id
|
||
= NULL WHERE id IN (folder_ids)` first. 2. **Widen the NULL UPDATE to cross-source-root children.** A "Leóns 1st Year"
|
||
SourceRoot at `.../Taco and Muli - 2024 onward/Leóns 1st Year` had its own folder rows whose `parent_id` pointed into
|
||
the Taco SourceRoot''s hierarchy. The first patch only NULLed `parent_id` for folders whose `id` was in the delete set;
|
||
the Leóns folders weren''t in that set so they kept their references and the FK still tripped. Fix: `UPDATE folders
|
||
SET parent_id = NULL WHERE parent_id IN (folder_ids)` — kills any incoming reference into the delete set, internal or
|
||
external. After both fixes: `DELETE /api/v1/nextcloud/source-roots/{id}` for `Taco and Muli - 2024 onward` cleared 4,158
|
||
photos and 35 folders in a single request and returned 200. UI swaps the trash icon for a `Loader2` spinner while the
|
||
request is in flight (`removeRoot.isPending && removeRoot.variables === r.id`) so the multi-second cascade is visible.
|
||
Separate fix in the same session: `PATCH /api/v1/photos/{id}` returned 500 with `can''t subtract offset-naive and offset-aware
|
||
datetimes` when the frontend sent a tz-aware ISO string for `taken_at` (the datetime-local input is supposed to be naive
|
||
but real-world locales / paste flows occasionally include `+02:00`). The DB column is `timestamp without time zone`,
|
||
so asyncpg refused to bind. Normalize on the server with `astimezone(timezone.utc).replace(tzinfo=None)` if `tzinfo
|
||
is not None`. Deploy infra learning: the new `--force-recreate` in `deploy.sh` does NOT reliably recreate containers
|
||
on every push — saw two consecutive deploys leave the backend at the previous `StartedAt` despite a fresh image. Manual
|
||
`docker compose up -d --no-deps --force-recreate <svc>` after deploy still races with whatever compose state the auto-deploy
|
||
left mid-flight, twice landing the stack in a half-broken state (orphaned `<id>_mulita-backend` rename containers, db
|
||
/ redis stopped). Recovery: `docker compose down && docker compose up -d`. Open question — `--force-recreate` may need
|
||
to be replaced with something more explicit.'
|
||
- date: '2026-05-10'
|
||
title: Hard-remove Nextcloud SourceRoot + reliable delete sync
|
||
body: '`DELETE /api/v1/nextcloud/source-roots/{id}` was a soft-deactivate (`is_active=false`) — the trash icon in Settings
|
||
only hid the SourceRoot from active queries while leaving every Folder + Photo row in the DB forever. Re-adding the
|
||
same path resurrected ghosts; `prune-missing` reported zero deletes for the soft-removed entry because the cleanup code
|
||
skipped inactive source roots. Endpoint now hard-deletes the SourceRoot, every Folder under it (chunked at 500), and
|
||
every Photo in those folders. `photo_tags` and `heap_photos` cascade automatically via `ON DELETE CASCADE` on the join
|
||
tables. `FolderShare` uses a stringly-typed `folder_id` (no FK) so cleaned manually for both `folder_type=''folder''`
|
||
and `folder_type=''source_root''`. Returns `{deleted_photos, deleted_folders}` so the UI can toast a count. Files in
|
||
Nextcloud are untouched. Sync side: the `watch_folders` celery task in `backend/app/tasks/scan.py` already detected
|
||
filesystem deletions and soft-marked photos `is_discarded=true`, but the global `task_soft_time_limit=300` in `app/tasks/celery.py`
|
||
was killing the watcher every five minutes and dropping every FS event during the restart window. The `soft_time_limit=None`
|
||
on the decorator was being interpreted as "use worker default" rather than "unlimited". Override to `soft_time_limit=0,
|
||
time_limit=0` (Celery convention for unbounded). Backstop: a new `discard_missing_photos()` in `app/services/cleanup.py`
|
||
— soft variant of `prune_missing_photos`, walks every `_sr_state == ''present''` source root and flips `is_discarded=true`
|
||
on Photo rows whose file is gone. Wired as `discard_missing_photos_beat` celery task scheduled every 30 minutes via
|
||
`beat_schedule` on the celery app. Beat runs in-process on `worker-watcher` (`--beat` flag in `docker-compose.yml`)
|
||
— the watcher is already a Redis-locked singleton so no need for a separate beat container. Manual `POST /api/v1/library/maintenance/prune-missing`
|
||
remains the hard-delete path for when the user wants to permanently drop orphan rows; the new beat job only soft-discards
|
||
(file gone → mule-image trash, restorable).'
|
||
- date: '2026-05-10'
|
||
title: OIDC auto-redirect on LoginPage
|
||
body: '`OIDC_ENABLED=true` was already set in `.env`, so the LoginPage rendered a "Sign in with Authentik" button next
|
||
to the password form. With a single trusted IdP and a logged-in Authentik session, that extra click was friction without
|
||
upside. `LoginPage.tsx` now reads `/auth/config` on mount and, if OIDC is on, immediately navigates to the OIDC login
|
||
URL. Authentik recognizes the existing session and bounces back through the callback with no user interaction. Two escape
|
||
hatches: `?password=1` in the URL forces the password form, and a `skipAutoSso` sessionStorage flag (set by `AuthContext.logout`
|
||
and by the OIDC callback''s error branch) suppresses the next auto-redirect so logouts actually log out and OIDC failures
|
||
surface their error instead of looping. While the redirect is in flight the page shows "Signing in with Authentik..."
|
||
plus a "Use password instead" link.'
|
||
- date: '2026-05-10'
|
||
title: right sidebar restructure (heap pinned, single scroll, collapsible Metadata)
|
||
body: 'The right sidepanel had three stacked flex regions: `<ActiveHeapCard />`, `<Header />`, and `<PhotoInfoPanel />`
|
||
— with `PhotoInfoPanel` carrying its own internal scroll. That left the editable fields (filename, title, notes, rating,
|
||
color, flag) stuck above the readonly metadata scroll, effectively two scroll boundaries on one sidebar. - Moved the
|
||
scroll boundary up to `RightSidebar.tsx`: only `ActiveHeapCard` and `Header` stay pinned now. Editable + readonly content
|
||
scroll together in one `flex-1 overflow-y-auto` region beneath them. - `PhotoInfoPanel.tsx` dropped its `h-full` / inner
|
||
`flex-1 overflow-y-auto`. - The four readonly sections (Tags / Basic Info / Camera / Location) are now wrapped in a
|
||
single outer "Metadata" `Collapsible`. Default expanded, one click hides the whole block. Sub-sections stay individually
|
||
collapsible. - Second pass: the editable form (filename / title / notes / rating / color / flag) got the same treatment
|
||
under an outer "Edit" `Collapsible` so the panel is now two equal collapsible groups below the title strip. Dropped
|
||
the X (clear-selection) button from the Header; Esc and grid-empty-area-click still clear. - Third pass: split editable
|
||
vs read-only between the two groups consistently. `TagsEditor` and `TakenAtEditor` had been buried inside the readonly
|
||
Metadata sub-sections — `Tags` as its own Section, taken-at wedged into `Basic Info` between size/dims and the filepath.
|
||
Moved both into the Edit collapsible, ordered identification → description → categorization: Filename · Title · Date
|
||
Taken · Notes · Tags · Rating · Color · Flag. Metadata now holds only readonly: `Basic Info` (size, dims, path), `Camera`,
|
||
`Location`. `PreviewView` reuses `RightSidebar` under an `<aside class="overflow-hidden">`, so the change applies in
|
||
both the grid sidebar and the preview overlay.'
|
||
- date: '2026-05-10'
|
||
title: '`photos.hubris.network` perf sweep'
|
||
body: 'User reported slow load. Five fixes shipped, in order: 1. **Backend out of dev mode.** `docker-compose.yml` `command:`
|
||
was running `uvicorn … --reload` (single worker, file-watcher). Page loads fan out ~15 parallel API calls; they all
|
||
serialized on one event loop. → `--workers 2 --proxy-headers`. Two uvicorn worker procs now. 2. **iPhone Apple ProRAW
|
||
/ Linear DNG decode fixed.** rawpy 0.26.1 + LibRaw 0.22 rejects Apple Linear DNG (`Photometric Interpretation: Linear
|
||
Raw`, 8-bit) as "Unsupported file format or not RAW file". Those files are TIFF containers with developed RGB inside
|
||
and have no embedded preview to fall back to. Added PIL `Image.open(src_path)` as the last fallback in both `_generate_proxy_webp`
|
||
(routers/photos.py) and `process_raw_image` (tasks/thumbs.py). ~1,300 iPhone DNGs were 415-ing on every detail view;
|
||
now decode in <1s via PIL. 3. **Reconcile DB with disk (renamed-folder case).** `prune_missing_photos` was skipping
|
||
all 4,154 orphaned photo rows under SourceRoot `Taco and Muli - 2024 onward` (renamed in Nextcloud to `Photo Archive
|
||
2004-2024`) because the leaf path didn''t resolve and the code assumed "drive unmounted → must skip". Added `_sr_state()`
|
||
to classify as `present` / `renamed` (parent mount fine, leaf gone) / `unmounted` (parent inaccessible). Only `unmounted`
|
||
still skips. Two stale source roots logged with a clearer hint pointing at `POST /api/v1/library/maintenance/prune-missing`.
|
||
User has not yet been asked to run that — endpoint is ready when they are. 4. **Frontend page size + idle polling.**
|
||
`usePhotosQuery.ts` was first-fetching `per_page=500` (slow paint + 500 thumb requests at once). Split into `PER_PAGE_INITIAL=100`
|
||
for first paint, `PER_PAGE_BACKGROUND=500` for the cursor-chain prefetch. Idle polling for `scan-status` and `worker-status`
|
||
(`useScanActivity.ts`, `ScanProgress.tsx`) bumped from 10s/15s to 30s/30s while idle; active cadence (2s/3s) unchanged.
|
||
5. **Partial index on `photos`.** Default list query `WHERE NOT is_trashed AND NOT is_hidden ORDER BY taken_at DESC
|
||
NULLS LAST, id DESC LIMIT N` was doing a seq-scan + top-N heapsort (~25ms standalone, worse under concurrency). Added
|
||
migration `0017_photos_list_index` creating `ix_photos_list_visible` (partial index on the sort key, restricted to visible
|
||
rows). EXPLAIN now shows an Index Only Scan → **24.7ms → 0.097ms**, ~250× speedup. **Deploy gotcha — fixed (with caveat).**
|
||
The original `docker compose up -d --build` in `deploy.sh` did not reliably recreate containers when only runtime config
|
||
(`command:`, env-only) or migration files changed; image hash would change but compose treated the existing container
|
||
as "current enough". Bit three times this session before `/opt/mule-deploy/deploy.sh` was updated to `docker compose
|
||
up -d --build --force-recreate`. Trade-off accepted: an extra restart cycle on deploys where nothing user-visible changed.
|
||
**Caveat:** the first auto-deploy after the flag flip raced with my own earlier manual `docker compose up -d --no-deps
|
||
--force-recreate frontend` and landed the stack in a half-broken state — `mulita-frontend` got stuck under a renamed
|
||
temp container, several services dropped off `mule-image_mulita-network`, frontend nginx restarted in a loop with `host
|
||
not found in upstream "backend"`. Fixed by `docker compose down && docker compose up -d`. Don''t issue a manual `--force-recreate`
|
||
on a single service while the auto-deploy webhook is also expected to fire — let the deploy own the lifecycle. **Data
|
||
drift still outstanding.** 4,154 photo rows + 1 unregistered folder (`Photo Archive 2004-2024`) on disk that''s not
|
||
a SourceRoot. The reconcile endpoints now work — user decides when to call them. The new folder needs to be added as
|
||
a SourceRoot via the Settings UI before its files will be indexed. **Proxy cache still empty** (`mule-image_proxies_data`
|
||
volume is 4 KB). Pre-generating ~500–1,500 WebP proxies for non-web-safe formats would make first-open of every RAW/HEIC
|
||
photo instant. Deferred — needs a one-shot script and the disk-space tradeoff isn''t worth it until the data-drift reconcile
|
||
happens first.'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-26'
|
||
title: Nextcloud-rooted libraries shipped
|
||
body: Bind `/mnt/library/homecloud` into the workers, reads via filesystem, writes via WebDAV. `users.nextcloud_username`
|
||
override field added; `dtoro → admin` mapping. Surviving SourceRoots cleaned up to NC paths.
|
||
- date: '2026-04-22'
|
||
title: native OIDC via Authentik
|
||
body: Authlib-based code in `backend/app/auth_oidc.py`. `extra_hosts` override for `auth.hubris.network` in compose override
|
||
(gitignored).
|
||
- date: '2026-04-21'
|
||
title: auto-deploy pipeline shipped
|
||
body: Webhook receiver at `:9797`, async deploy returning 202. Mirrors caddy-conf / gitea-customizations.
|
||
- slug: containers/121-caddy
|
||
title: 121 — `caddy`
|
||
content: '# 121 — `caddy`
|
||
|
||
|
||
The reverse proxy. Terminates TLS for every `*.hubris.network` hostname on the LAN/mesh and forwards to the right backend.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `caddy`
|
||
|
||
- **IP:** `192.168.8.175`
|
||
|
||
- **Privilege:** **unprivileged**
|
||
|
||
- **Resources:** 1 core / 512 MiB RAM / 6 GiB rootfs
|
||
|
||
- **Mounts:** none from `/mnt/library`
|
||
|
||
- **Config:** `/etc/caddy/Caddyfile` is a [git checkout of `dtoro/caddy-conf`](#auto-deploy)
|
||
|
||
- **Cert source:** Let''s Encrypt **DNS-01** via IONOS API (`IONOS_AUTH_API_TOKEN`).
|
||
|
||
|
||
## Sites currently served (live as of 2026-07-06)
|
||
|
||
|
||
- `artifacto.hubris.network` → [apps (105)](105-apps.md) `:3100`
|
||
|
||
- `blog.hubris.network` → [apps (105)](105-apps.md) `:8080`
|
||
|
||
- `books.hubris.network` → [grimmory (130)](130-grimmory.md) `:6060`
|
||
|
||
- `cloud.hubris.network` → [nextcloud (114)](114-nextcloud.md) `:443`
|
||
|
||
- `docker.hubris.network` → [apps (105)](105-apps.md) `:9443`
|
||
|
||
- `git.hubris.network` → [gitea (104)](104-gitea.md) `:3000` (+ `handle_path /_plantuml/*` → apps `:8079`)
|
||
|
||
- `home.hubris.network` → [haos VM (108)](../vms/108-haos.md) `192.168.8.101:8123`
|
||
|
||
- `house.hubris.network` → [house (129)](129-house.md) `:3000`
|
||
|
||
- `jellyseerr.hubris.network` → [arriman (122)](122-arriman.md) `:5056`
|
||
|
||
- `matrix.hubris.network` → [elementsynapse (118)](118-elementsynapse.md) `:8008`
|
||
|
||
- `media.hubris.network` → [jellyfin (101)](101-jellyfin.md) `:8096`
|
||
|
||
- `paperless.hubris.network` → [paperless (103)](103-paperless.md) `:8000`
|
||
|
||
- `photos.hubris.network` → [mule-images (120)](120-mule-images.md) `:3000`
|
||
|
||
- `proxmox.hubris.network` → [hubris host](../hosts/hubris.md) `:8006`
|
||
|
||
- `qbit.hubris.network` → [arriman (122)](122-arriman.md) `:8080`
|
||
|
||
- `roms.hubris.network` → [romm (134)](134-romm.md) `:80`
|
||
|
||
- `sab.hubris.network` → [arriman (122)](122-arriman.md) `:8082` (Authentik forward-auth)
|
||
|
||
- `teddy.hubris.network` → LXC 131 `192.168.8.150:8443`
|
||
|
||
- `trmnl.hubris.network` → [trmnl (128)](128-trmnl.md) `:9851`
|
||
|
||
|
||
> **Reminder:** Caddy alone isn''t enough to make a new subdomain reachable on the LAN. Each one needs an entry in [DNS
|
||
split-horizon](../infrastructure/dns.md) too.
|
||
|
||
|
||
## Snippet: `(authentik)` forward-auth
|
||
|
||
|
||
A snippet at the top of the Caddyfile (used as `import authentik` in any site block) wires forward-auth to the embedded
|
||
Authentik outpost. It points at `http://192.168.8.180:9000` directly (NOT `https://auth.hubris.network`) to avoid Caddy-to-self
|
||
round-tripping that strips `X-Forwarded-Host`. The forward-auth block must explicitly set `header_up X-Forwarded-Host
|
||
{host}`. See [Authentik](106-auth-outpost.md#forward-auth-domain-level-setup).
|
||
|
||
|
||
For apps with mobile clients, `/api/*` (or equivalent) bypasses forward-auth — see the per-app gotchas in [Authentik](106-auth-outpost.md).
|
||
|
||
|
||
## Caddy environment
|
||
|
||
|
||
Untracked secrets in `/etc/caddy/caddy.env`:
|
||
|
||
- `IONOS_AUTH_API_TOKEN` — DNS-01 ACME
|
||
|
||
- `ARTIFACTO_GATEWAY_SECRET` — must match `SSO_GATEWAY_SECRET` on [apps (105)](105-apps.md). Rotate together.
|
||
|
||
|
||
## Auto-deploy
|
||
|
||
|
||
Push to `dtoro/caddy-conf` master → gitea webhook → `http://192.168.8.175:9797/deploy` → `caddy-deploy-webhook.service`
|
||
validates HMAC against `/etc/caddy-deploy/secret` → runs `/etc/caddy/scripts/deploy.sh`:
|
||
|
||
1. `git pull`
|
||
|
||
2. `caddy validate`
|
||
|
||
3. `systemctl reload caddy`
|
||
|
||
|
||
Gitea webhook id 2 on `dtoro/caddy-conf`. Receiver, deploy script, install scripts, systemd unit all live in the repo
|
||
under `scripts/`. The unit has `ReadWritePaths=/etc/caddy` — load-bearing (`ProtectSystem=full` would otherwise block
|
||
`git pull`). Git credentials for the root-run service are at `/etc/caddy-deploy/git-credentials` (mode 600).
|
||
|
||
|
||
> **Systemd unit changes inside the repo do NOT auto-reinstall** — you have to re-run `scripts/webhook/install.sh` manually
|
||
after editing the unit.
|
||
|
||
|
||
## Related
|
||
|
||
- [DNS split-horizon](../infrastructure/dns.md) — must add entry for every new subdomain
|
||
|
||
- [Authentik (124)](106-auth-outpost.md) — forward-auth + IdP
|
||
|
||
- [Auto-deploy](../infrastructure/auto-deploy.md)
|
||
|
||
- [Public ingress (VPS traefik)](../infrastructure/ingress.md) — mirrors Caddy''s certs to the VPS for public exposure
|
||
|
||
- [Gitea (104)](104-gitea.md) — webhook source
|
||
|
||
- All container pages list their Caddy backend.
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-13 — sab.hubris.network gated with Authentik forward-auth; port fixed :8081→:8082
|
||
|
||
`sab.hubris.network` now uses `import authentik` inside a `handle` block. SABnzbd own auth disabled, `local_ranges` set
|
||
for transparent proxy. Port bumped from `:8081` to `:8082` (fix from 2026-06-04) now documented.
|
||
|
||
|
||
### 2026-06-06 — Caddyfile truncated to 43 lines; restored from origin/master + safeguards added
|
||
|
||
**Symptom:** All `*.hubris.network` hosts except `photos` and `auth` (VPS-hosted) returned `tlsv1 alert internal error`
|
||
or timeout. Only 3 site blocks (`photos`, `prism`, `photos2`) remained in the Caddyfile.
|
||
|
||
|
||
**Root cause:** The Caddyfile was manually edited directly on LXC 121 (not via the `dtoro/caddy-conf` git repo), overwriting
|
||
260 lines / 30+ site blocks with 43 lines of photo-only config.
|
||
|
||
|
||
**Fix:**
|
||
|
||
- Restored Caddyfile from `origin/master` (`git checkout --force origin/master -- Caddyfile`)
|
||
|
||
- `systemctl reload caddy`
|
||
|
||
|
||
**Permanent safeguards added to `/etc/caddy/scripts/deploy.sh`:**
|
||
|
||
- **Site-count guard:** refuses to reload if fewer than 20 `*.hubris.network` blocks detected
|
||
|
||
- **Dirty-tree auto-stash:** stashes local changes before `git pull --ff-only` so the webhook doesn''t fail on local edits
|
||
|
||
- **Auto-backup:** saves `Caddyfile.bak.<timestamp>` before any modifications, keeps last 5
|
||
|
||
|
||
Also: [elementsynapse LXC 118](118-elementsynapse.md) found to have DHCP-overridden static IP (actual `.244` vs config
|
||
`.239`) during incident investigation — fixed.
|
||
|
||
|
||
### 2026-06-02 — caddy.service unit missing; recreated
|
||
|
||
After the Slate AX → SODOLA network migration, Caddy was not listening (ports 80/443 dead). Root cause: the custom hubris1
|
||
Debian package (`caddy_1:2.11.3-hubris1_amd64`) does not ship a systemd service unit file. The unit had previously existed
|
||
but was lost (likely on a package reinstall). Recreated at `/lib/systemd/system/caddy.service` with standard Caddy service
|
||
config + `EnvironmentFile=/etc/caddy/caddy.env` (already present in `caddy.service.d/override.conf`). **Risk:** the unit
|
||
will be lost again if the package is reinstalled without the file being tracked. Fix: add the service unit to the `caddy-conf`
|
||
repo or rebuild the hubris1 package to include it.
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Initial documentation. 16 active sites at this date.
|
||
|
||
|
||
### 2026-04-22 — Artifacto + Authentik gateway-secret pattern wired
|
||
|
||
Caddy site `artifacto.hubris.network`: forward-auth on admin paths, `header_up X-Artifacto-Gateway {env.ARTIFACTO_GATEWAY_SECRET}`,
|
||
public matchers `/p/*`, `/static/*`, `/healthz`. `caddy.env` extended with the gateway secret.
|
||
|
||
|
||
### 2026-04-21 — `(authentik)` snippet added
|
||
|
||
Top-of-file snippet used by every gated site. Targets `http://192.168.8.180:9000` directly to dodge hairpin TLS round-trip.
|
||
|
||
|
||
### 2026-04-21 — three new sites added (jellyseerr, qbit, sab → 192.168.8.132)
|
||
|
||
For the [arriman (122)](122-arriman.md) docker stack. Caddy + certs work; needed corresponding [DNS split-horizon](../infrastructure/dns.md)
|
||
entries to be added to land cleanly.
|
||
|
||
|
||
### 2026-04-20 — auto-deploy pipeline shipped
|
||
|
||
Repo `dtoro/caddy-conf`. Webhook receiver at `:9797`, validate + reload on push to master.
|
||
|
||
'
|
||
entity_slug: lxc:caddy
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: caddy
|
||
ip: 192.168.8.175
|
||
privileged: unprivileged
|
||
resources: 1 core / 512 MiB RAM / 6 GiB rootfs
|
||
mounts: none from /mnt/library
|
||
config: /etc/caddy/Caddyfile is a [git checkout of dtoro/caddy-conf]
|
||
cert_source: Let's Encrypt DNS-01 via IONOS API (IONOS_AUTH_API_TOKEN).
|
||
changelog:
|
||
- date: '2026-06-13'
|
||
title: sab.hubris.network gated with Authentik forward-auth; port fixed :8081→:8082
|
||
body: '`sab.hubris.network` now uses `import authentik` inside a `handle` block. SABnzbd own auth disabled, `local_ranges`
|
||
set for transparent proxy. Port bumped from `:8081` to `:8082` (fix from 2026-06-04) now documented.'
|
||
- date: '2026-06-06'
|
||
title: Caddyfile truncated to 43 lines; restored from origin/master + safeguards added
|
||
body: '**Symptom:** All `*.hubris.network` hosts except `photos` and `auth` (VPS-hosted) returned `tlsv1 alert internal
|
||
error` or timeout. Only 3 site blocks (`photos`, `prism`, `photos2`) remained in the Caddyfile. **Root cause:** The
|
||
Caddyfile was manually edited directly on LXC 121 (not via the `dtoro/caddy-conf` git repo), overwriting 260 lines /
|
||
30+ site blocks with 43 lines of photo-only config. **Fix:** - Restored Caddyfile from `origin/master` (`git checkout
|
||
--force origin/master -- Caddyfile`) - `systemctl reload caddy` **Permanent safeguards added to `/etc/caddy/scripts/deploy.sh`:**
|
||
- **Site-count guard:** refuses to reload if fewer than 20 `*.hubris.network` blocks detected - **Dirty-tree auto-stash:**
|
||
stashes local changes before `git pull --ff-only` so the webhook doesn''t fail on local edits - **Auto-backup:** saves
|
||
`Caddyfile.bak.<timestamp>` before any modifications, keeps last 5 Also: [elementsynapse LXC 118](118-elementsynapse.md)
|
||
found to have DHCP-overridden static IP (actual `.244` vs config `.239`) during incident investigation — fixed.'
|
||
- date: '2026-06-02'
|
||
title: caddy.service unit missing; recreated
|
||
body: 'After the Slate AX → SODOLA network migration, Caddy was not listening (ports 80/443 dead). Root cause: the custom
|
||
hubris1 Debian package (`caddy_1:2.11.3-hubris1_amd64`) does not ship a systemd service unit file. The unit had previously
|
||
existed but was lost (likely on a package reinstall). Recreated at `/lib/systemd/system/caddy.service` with standard
|
||
Caddy service config + `EnvironmentFile=/etc/caddy/caddy.env` (already present in `caddy.service.d/override.conf`).
|
||
**Risk:** the unit will be lost again if the package is reinstalled without the file being tracked. Fix: add the service
|
||
unit to the `caddy-conf` repo or rebuild the hubris1 package to include it.'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation. 16 active sites at this date.
|
||
- date: '2026-04-22'
|
||
title: Artifacto + Authentik gateway-secret pattern wired
|
||
body: 'Caddy site `artifacto.hubris.network`: forward-auth on admin paths, `header_up X-Artifacto-Gateway {env.ARTIFACTO_GATEWAY_SECRET}`,
|
||
public matchers `/p/*`, `/static/*`, `/healthz`. `caddy.env` extended with the gateway secret.'
|
||
- date: '2026-04-21'
|
||
title: '`(authentik)` snippet added'
|
||
body: Top-of-file snippet used by every gated site. Targets `http://192.168.8.180:9000` directly to dodge hairpin TLS
|
||
round-trip.
|
||
- date: '2026-04-21'
|
||
title: three new sites added (jellyseerr, qbit, sab → 192.168.8.132)
|
||
body: For the [arriman (122)](122-arriman.md) docker stack. Caddy + certs work; needed corresponding [DNS split-horizon](../infrastructure/dns.md)
|
||
entries to be added to land cleanly.
|
||
- date: '2026-04-20'
|
||
title: auto-deploy pipeline shipped
|
||
body: Repo `dtoro/caddy-conf`. Webhook receiver at `:9797`, validate + reload on push to master.
|
||
- slug: containers/122-arriman
|
||
title: 122 — `arriman`
|
||
content: "# 122 — `arriman`\n\nDocker host running the \\*arr stack via [`ezarr`](https://github.com/ezarr/ezarr) compose.\
|
||
\ Replaced the old yunohost-based [LXC 100](index.md#recently-destroyed-kept-for-archaeology) on 2026-04-21.\n\n## At\
|
||
\ a glance\n- **Hostname:** `arriman`\n- **IP:** `192.168.8.245`\n- **Host:** **strong** (migrated from hubris 2026-07-05)\n\
|
||
- **Privilege:** privileged\n- **Resources:** 4 cores / 8 GiB RAM / 24 GiB rootfs\n- **Mounts:** `/mnt/media_local` ↔\
|
||
\ `/mnt/library`\n- **Public hostnames:** `jellyseerr` / `qbit` / `sab` (see below)\n\n## Compose\n\n`/home/arr/ezarr/docker-compose.yml`.\
|
||
\ Network `ezarr_default` `172.18.0.0/16`. Service users all run with primary GID 10000 (host `media` group) inside their\
|
||
\ containers — see [media permissions](../infrastructure/media-permissions.md). Per-app `/config/<app>-config/` dirs bind\
|
||
\ to `/config` inside.\n\n`.env` contents:\n- `DOWNLOADS_DIR=/mnt/library/downloads/`\n- `MEDIA_DIR=/mnt/library/`\n-\
|
||
\ `MEDIACENTER_GID=10000` ← **must be 10000, not 13000.** s6-setuidgid only honors the primary PGID; `group_add:` doesn't\
|
||
\ propagate. Without this, qBit errors every torrent with \"Permission denied\".\n- `SONARR_UID=13001`, `RADARR_UID=13002`,\
|
||
\ `LIDARR_UID=13003`, `PROWLARR_UID=13006`, `JELLYSEERR_UID=13012`, `BAZARR_UID=13013`, `QBITTORRENT_UID`, `SABNZBD_UID`\n\
|
||
\n## Service / port map\n\nAll services route through gluetun's network namespace. Ports are exposed via\nthe gluetun\
|
||
\ container:\n\n| Service | Host:Container | Public hostname |\n| ------------- | --------------\
|
||
\ | ------------------------------------ |\n| gluetun (VPN) | — | — |\n\
|
||
| sonarr | `8989:8989` | direct only (via gluetun) |\n| radarr | `7878:7878` | direct only\
|
||
\ (via gluetun) |\n| lidarr | `8686:8686` | direct only (via gluetun) |\n| prowlarr \
|
||
\ | `9696:9696` | direct only (via gluetun) |\n| bazarr | `6767:6767` | direct only (via gluetun)\
|
||
\ |\n| jellyseerr | `5056:5055` | [`jellyseerr.hubris.network`](../infrastructure/dns.md) |\n| qbittorrent\
|
||
\ | `8080:8080` | [`qbit.hubris.network`](../infrastructure/dns.md) |\n| sabnzbd | `8082:8082` HTTP,\
|
||
\ `9090:9090` HTTPS | [`sab.hubris.network`](../infrastructure/dns.md) |\n| flaresolverr | `8191:8191` | internal\
|
||
\ only |\n| homarr | `7575:7575` | internal only |\n\nInternal\
|
||
\ *arr ↔ *arr / *arr ↔ qBit/SAB/flaresolverr comms run on `localhost:<port>`\n(services share gluetun's shared network\
|
||
\ namespace). External services reach them\nvia `gluetun:<port>` (e.g. Sonarr → qBittorrent at `localhost:8080` or\n`gluetun:8080`).\n\
|
||
\n## Categories (qBit + SAB + *arr)\n\n| Category | qBit savePath | SAB dir | Consumer \
|
||
\ |\n| -------- | --------------------------- | ------- | --------------------------------- |\n| tv | `/data/torrents/tv`\
|
||
\ | tv | Sonarr |\n| movies | `/data/torrents/movies` | movies | Radarr\
|
||
\ |\n| music | `/data/torrents/music` | music | Lidarr \
|
||
\ |\n| anime | `/data/torrents/anime` | anime | Sonarr (2nd download client, root `/data/media/anime`) |\n\
|
||
| books / audiobooks / comics / podcasts | `/data/torrents/<cat>` | `<cat>` | (reserved, none in use yet) |\n\nPath mapping:\
|
||
\ host `/mnt/library/<cat>` ↔ container `/data/media/<cat>`. Downloads: host `/mnt/library/downloads/<torrents|usenet>/<cat>`\
|
||
\ ↔ container `/data/torrents/<cat>` and `/data/usenet/<cat>`.\n\n## Auth (reverse-proxy + Authentik forward-auth)\n\n\
|
||
### qBit\n\nAuto-login behind forward-auth via IP whitelist. `qBittorrent.conf` lines:\n- `WebUI\\\\AuthSubnetWhitelist=172.18.0.0/16,\
|
||
\ 172.17.0.0/16, 192.168.8.175/32`\n- `WebUI\\\\ReverseProxySupportEnabled=true`\n- `WebUI\\\\TrustedReverseProxiesList=192.168.8.175,\
|
||
\ 172.18.0.0/16`\n\n> **Stop the container before editing `qBittorrent.conf`.** qBit writes its in-memory config on graceful\
|
||
\ shutdown and clobbers any live edits. Recipe: `docker stop qbittorrent && sed -i ... && docker start qbittorrent`.\n\
|
||
\nMobile/desktop clients keep working via `/api/v2/*` path bypass on Caddy.\n\n### SABnzbd\n\nGated with Authentik forward-auth\
|
||
\ (applied 2026-06-13). Caddy `sab.hubris.network` block uses `import authentik` inside a `handle` block. SABnzbd's own\
|
||
\ web auth is disabled:\n\n- `html_login = 0` → no HTML login form\n- `username` / `password` cleared → CherryPy Basic\
|
||
\ Auth not activated\n- `local_ranges = 172.18.0.0/16, 192.168.8.0/24, 127.0.0.0/8` → proxied requests from Caddy (192.168.8.x)\
|
||
\ and Docker-proxy (172.18.x) pass without auth\n\n**API key** (`67ef5a45e4e04157994e977005a33878`) still works for internal\
|
||
\ service-to-service calls (Sonarr/Radarr/Lidarr via Docker internal networking — they talk to SAB at `localhost:8082`,\
|
||
\ not through Caddy).\n\n`host_whitelist`: `sabnzbd, localhost, 127.0.0.1, 192.168.8.132, sab.hubris.network` — extend\
|
||
\ before accessing SAB from a new host.\n\n## Credentials\n\nBackups of all migration-time secrets at `/mnt/library/backup/arr-migration-2026-04-21/state/`:\n\
|
||
- `lxc100-apikey-{sonarr,radarr,prowlarr,lidarr}.txt`\n- `qbit-creds.txt`, `sab-creds.txt`\n- `lxc100-jellyseerr-settings.json`\n\
|
||
\n\\*arr API keys were preserved from the yunohost DB — don't rotate unless needed.\n\nqBittorrent: `admin` / `CDyjxMKvjfGupCKzXFkjsKw4nRrXMRL`.\n\
|
||
\n## Migration gotchas (apply if ever re-importing)\n\n- Imported \\*arr `config.xml`: must have `<BindAddress>*</BindAddress>`,\
|
||
\ `<Port>{ezarr-port}</Port>`, `<UrlBase></UrlBase>`. Otherwise app listens only on 127.0.0.1, wrong port, wrong path.\n\
|
||
- Imported DB paths must be rewritten `/mnt/library/` → `/data/media/` (RootFolders / Series / Movies / Artists tables).\
|
||
\ Without this, scans and imports fail silently.\n- Bazarr `config/config.yaml` (yaml, not ini) — `ip: 0.0.0.0`, `port:\
|
||
\ 6767`, empty top-level `base_url`; sonarr/radarr connection sections normalized to docker service names.\n- Staging\
|
||
\ dir for migrations: `/mnt/library/backup/stage/<app>/` (visible to both old LXC 100 and arriman because both bind-mount\
|
||
\ the same host path).\n\n## Permissions\n\nMember of [media GID 10000](../infrastructure/media-permissions.md). The LXC\
|
||
\ has `www-data, audiobookshelf, radarr, sonarr, lidarr, prowlarr, qbittorrent, bazarr, jellyseerr, mylar, jackett, overseerr,\
|
||
\ plex, arr` in the `media` group. Each docker container also receives GID 10000 as its primary group via `MEDIACENTER_GID=10000`.\n\
|
||
\n## Rollback material (kept until 2026-04-28+)\n\n`/mnt/library/backup/arr-migration-2026-04-21/`:\n- `tars/` — per-subtree\
|
||
\ tars of yunohost app data (granular restore)\n- `snapshots/` — vzdumps of LXC 100 (pre-migration) and LXC 122 (pre-import),\
|
||
\ plus post-stop vzdumps of LXC 100 + 106\n- `stage/<app>-preimport.tar.gz` — original empty ezarr `/config/<app>-config/`\
|
||
\ dirs\n- `state/` — service dumps, listening ports, API keys, jellyseerr settings at migration time\n- `DESTROY_AFTER_2026-04-28.md`\
|
||
\ — destroy checklist + post-destroy restore commands\n\n## Related\n- [Caddy (121)](121-caddy.md)\n- [Authentik (124)](106-auth-outpost.md)\
|
||
\ — forward-auth wiring + per-app `/api/*` bypass\n- [DNS](../infrastructure/dns.md)\n- [Media permissions](../infrastructure/media-permissions.md)\n\
|
||
- [Hubris host](../hosts/hubris.md)\n\n## Changelog\n\n### 2026-06-13 — SABnzbd gated with Authentik forward-auth\nSABnzbd\
|
||
\ now uses Authentik forward-auth (same `import authentik` Caddy pattern as qBit). SABnzbd's own web auth disabled: `html_login=0`,\
|
||
\ credentials cleared, `local_ranges` extended to cover Docker bridge + homelab LAN. API key still works for internal\
|
||
\ *arr service calls. See [Auth section](#auth-reverse-proxy--authentik-forward-auth) above.\n\n### 2026-06-04 — all arr\
|
||
\ services moved behind gluetun VPN; SAB port conflict fixed\n\n- All services (sonarr, radarr, lidarr, bazarr, prowlarr,\
|
||
\ jellyseerr, homarr,\n flaresolverr) now use `network_mode: service:gluetun` — whole stack routes\n through the VPN\n\
|
||
- Port mappings moved from individual services to gluetun container\n- **Fixed SABnzbd port conflict**: was crashing in\
|
||
\ a restart loop because\n qBittorrent held port 8080 inside the shared gluetun namespace. Changed\n SAB internal port\
|
||
\ to 8082 (config at `/config/sabnzbd-config/sabnzbd.ini`)\n- Caddy `sab.hubris.network` updated to point to `:8082`\n\
|
||
- Jellyseerr's `extra_hosts` (auth.hubris.network) moved to gluetun since\n `extra_hosts` conflicts with `network_mode`\n\
|
||
\n### 2026-06-02 — ProtonVPN added (gluetun); LXC IP set static\n- Added `gluetun` container to compose as a WireGuard\
|
||
\ VPN sidecar (ProtonVPN, server AL#57, located in Tirana, Albania)\n- **qbittorrent** and **sabnzbd** now use `network_mode:\
|
||
\ service:gluetun` — all traffic routes through the VPN\n- Ports 8080 (qBit WebUI), 6881 tcp/udp (qBit BT), 8081 (SAB\
|
||
\ WebUI) exposed through gluetun\n- gluetun config at `gluetun-config/wireguard/wg0.conf` (read-only mount)\n- Healthcheck\
|
||
\ on gluetun; qBit/SAB wait for `service_healthy` before starting\n- LXC IP changed from DHCP to static (`192.168.8.132`)\
|
||
\ via `pct set` + `/etc/network/interfaces`\n- **After first start:** Sonarr/Radarr/Lidarr download client host needs\
|
||
\ updating from `qbittorrent` → `gluetun` (SAB similarly `sabnzbd` → `gluetun`)\n- **Also fixed:** 7 other DHCP LXCs (101\
|
||
\ jellyfin, 103 paperless, 104 gitea, 105 apps, 114 nextcloud, 118 elementsynapse, 120 mule-images, 121 caddy) set to\
|
||
\ static IPs to prevent floating on reboot. See infrastructure/dns.md.\n\n### 2026-04-28 — wiki entry created\nInitial\
|
||
\ documentation.\n\n### 2026-04-26 — `MEDIACENTER_GID` flipped 13000 → 10000\nqBit was erroring every torrent with \"\
|
||
Permission denied\" because s6-setuidgid only honors the primary PGID and `group_add:` doesn't propagate. Setting the\
|
||
\ primary GID to 10000 (the host `media` group) fixed it.\n\n### 2026-04-21 — \\*arr stack migrated yunohost → docker-compose\n\
|
||
Old LXC 100 stopped; ezarr stack stood up here; data and DBs imported with path rewrites. flaresolverr folded in from\
|
||
\ old LXC 106. Three Caddy sites added (jellyseerr, qbit, sab) → see [caddy](121-caddy.md). Rollback artifacts under `/mnt/library/backup/arr-migration-2026-04-21/`.\n"
|
||
entity_slug: lxc:arriman
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: arriman
|
||
ip: 192.168.8.245
|
||
host: strong
|
||
privileged: privileged
|
||
resources: 4 cores / 8 GiB RAM / 24 GiB rootfs
|
||
mounts: /mnt/media_local /mnt/library
|
||
public_hostnames: jellyseerr / qbit / sab
|
||
changelog:
|
||
- date: '2026-06-13'
|
||
title: SABnzbd gated with Authentik forward-auth
|
||
body: 'SABnzbd now uses Authentik forward-auth (same `import authentik` Caddy pattern as qBit). SABnzbd''s own web auth
|
||
disabled: `html_login=0`, credentials cleared, `local_ranges` extended to cover Docker bridge + homelab LAN. API key
|
||
still works for internal *arr service calls. See [Auth section](#auth-reverse-proxy--authentik-forward-auth) above.'
|
||
- date: '2026-06-04'
|
||
title: all arr services moved behind gluetun VPN; SAB port conflict fixed
|
||
body: '- All services (sonarr, radarr, lidarr, bazarr, prowlarr, jellyseerr, homarr, flaresolverr) now use `network_mode:
|
||
service:gluetun` — whole stack routes through the VPN - Port mappings moved from individual services to gluetun container
|
||
- **Fixed SABnzbd port conflict**: was crashing in a restart loop because qBittorrent held port 8080 inside the shared
|
||
gluetun namespace. Changed SAB internal port to 8082 (config at `/config/sabnzbd-config/sabnzbd.ini`) - Caddy `sab.hubris.network`
|
||
updated to point to `:8082` - Jellyseerr''s `extra_hosts` (auth.hubris.network) moved to gluetun since `extra_hosts`
|
||
conflicts with `network_mode`'
|
||
- date: '2026-06-02'
|
||
title: ProtonVPN added (gluetun); LXC IP set static
|
||
body: '- Added `gluetun` container to compose as a WireGuard VPN sidecar (ProtonVPN, server AL#57, located in Tirana,
|
||
Albania) - **qbittorrent** and **sabnzbd** now use `network_mode: service:gluetun` — all traffic routes through the
|
||
VPN - Ports 8080 (qBit WebUI), 6881 tcp/udp (qBit BT), 8081 (SAB WebUI) exposed through gluetun - gluetun config at
|
||
`gluetun-config/wireguard/wg0.conf` (read-only mount) - Healthcheck on gluetun; qBit/SAB wait for `service_healthy`
|
||
before starting - LXC IP changed from DHCP to static (`192.168.8.132`) via `pct set` + `/etc/network/interfaces` - **After
|
||
first start:** Sonarr/Radarr/Lidarr download client host needs updating from `qbittorrent` → `gluetun` (SAB similarly
|
||
`sabnzbd` → `gluetun`) - **Also fixed:** 7 other DHCP LXCs (101 jellyfin, 103 paperless, 104 gitea, 105 apps, 114 nextcloud,
|
||
118 elementsynapse, 120 mule-images, 121 caddy) set to static IPs to prevent floating on reboot. See infrastructure/dns.md.'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-26'
|
||
title: '`MEDIACENTER_GID` flipped 13000 → 10000'
|
||
body: qBit was erroring every torrent with "Permission denied" because s6-setuidgid only honors the primary PGID and `group_add:`
|
||
doesn't propagate. Setting the primary GID to 10000 (the host `media` group) fixed it.
|
||
- date: '2026-04-21'
|
||
title: \*arr stack migrated yunohost → docker-compose
|
||
body: Old LXC 100 stopped; ezarr stack stood up here; data and DBs imported with path rewrites. flaresolverr folded in
|
||
from old LXC 106. Three Caddy sites added (jellyseerr, qbit, sab) → see [caddy](121-caddy.md). Rollback artifacts under
|
||
`/mnt/library/backup/arr-migration-2026-04-21/`.
|
||
- slug: containers/128-trmnl
|
||
title: 128 — `trmnl`
|
||
content: '# 128 — `trmnl`
|
||
|
||
|
||
Self-hosted middleware for TRMNL e-ink plugins. TRMNL cloud polls it; it fetches/shapes live data into JSON the plugin''s
|
||
Liquid template renders.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `trmnl`
|
||
|
||
- **IP:** `192.168.8.211`
|
||
|
||
- **Privilege:** unprivileged
|
||
|
||
- **Resources:** 1 core / 768 MiB RAM / 8 GiB rootfs (Debian 13)
|
||
|
||
- **Mounts:** none
|
||
|
||
- **Public hostname:** `trmnl.hubris.network` (via [VPS ingress](../infrastructure/ingress.md))
|
||
|
||
|
||
## Role
|
||
|
||
Runs one FastAPI aggregator (`server.app:app`, port 9851) that mounts a router per plugin from the `dtoro/terminalito`
|
||
repo. First consumer: `munich-home` (`/munich-home/dashboard`) — weather (Open-Meteo), MVG transit, Google Calendar, plus
|
||
server-side Kita/quote logic. Talks out to the public internet for those APIs; TRMNL cloud polls it inbound every 15 min.
|
||
Bearer-token gated (`TRMNL_POLL_TOKEN`); `/health` is open.
|
||
|
||
|
||
## Service / port map
|
||
|
||
| Service | Listen | Notes |
|
||
|
||
|---------|--------|-------|
|
||
|
||
| `trmnl-plugins` | `0.0.0.0:9851` | uvicorn aggregator; `EnvironmentFile=/etc/trmnl-plugins/env` |
|
||
|
||
|
||
## Storage / config paths
|
||
|
||
- `/opt/terminalito` — git checkout (origin = internal gitea `http://192.168.8.121:3000/dtoro/terminalito.git`)
|
||
|
||
- `/opt/terminalito/server/.venv` — venv
|
||
|
||
- `/etc/trmnl-plugins/env` — `TRMNL_POLL_TOKEN` (+ Google/MVG creds once enrolled)
|
||
|
||
- `/etc/systemd/system/trmnl-plugins.service`
|
||
|
||
|
||
## Auto-deploy
|
||
|
||
Wired — [auto-deploy](../infrastructure/auto-deploy.md) Shape B, webhook id 12 on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy`
|
||
(`terminalito-deploy.service`). Push to `main` → `server/deploy/deploy.sh` (`git pull` + pip + reinstall units + restart
|
||
`trmnl-plugins`). Secret `/etc/terminalito-deploy/secret`; git creds `/etc/terminalito-deploy/git-credentials` wired as
|
||
a repo-local `credential.helper`. Manual: `pct exec 128 -- /opt/terminalito/server/deploy/deploy.sh`.
|
||
|
||
|
||
## Secrets
|
||
|
||
Not yet SOPS-enrolled. The poll token is set directly in `/etc/trmnl-plugins/env`. Google Calendar + MVG creds are pending:
|
||
enroll via `homelab client add trmnl` + bootstrap, add `secrets/trmnl-oauth.yaml`, then `server/deploy/render-env.sh`
|
||
builds the env from `homelab secret trmnl-oauth`. Until then calendar/transit cards degrade to empty; weather works.
|
||
|
||
|
||
## Related
|
||
|
||
- [Caddy (121)](121-caddy.md) — LAN reverse proxy (`trmnl.hubris.network → 192.168.8.211:9851`)
|
||
|
||
- [VPS ingress](../infrastructure/ingress.md) — public edge (cert mirror + traefik router)
|
||
|
||
- [DNS (107)](107-dns.md) — Technitium A record `trmnl → 192.168.8.175` (LAN path via Caddy)
|
||
|
||
- [Gitea (104)](104-gitea.md) — source repo `dtoro/terminalito`
|
||
|
||
- [Plan: 2026-06-24 TRMNL plugins LXC](../../../plans/2026-06-24-trmnl-plugins-lxc.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
### 2026-06-24 — auto-deploy + LAN DNS wired
|
||
|
||
Gitea Shape-B deploy pipeline (webhook id 12, `:9797`) — push to `dtoro/terminalito` redeploys; verified end-to-end. Technitium
|
||
A record `trmnl.hubris.network → 192.168.8.175` added on [dns (107)](107-dns.md) (propagated to the NetBird managed zone
|
||
via dns-sync), so LAN clients take the short path through [Caddy (121)](121-caddy.md). See [auto-deploy](../infrastructure/auto-deploy.md).
|
||
|
||
|
||
### 2026-06-24 — public path live
|
||
|
||
Verified end-to-end from the internet: `https://trmnl.hubris.network/munich-home/dashboard` → 200 with token, 401 without;
|
||
`/health` 200. The provision-time outage was the netbird `home-lab-network` (192.168.8.0/24) route having no active routing
|
||
peer — the **mac-mini routing peer''s netbird daemon was down** (artifacto/blog were 504 too). Bringing netbird up on
|
||
mac-mini restored the route; the edge recovered with no config change. See [ingress](../infrastructure/ingress.md) / [mesh](../infrastructure/mesh.md).
|
||
|
||
|
||
### 2026-06-24 — provisioned
|
||
|
||
LXC 128 created (Debian 13, unprivileged, `192.168.8.211`). Deployed `trmnl-plugins.service` on :9851 from `dtoro/terminalito`.
|
||
Caddy block added (`dtoro/caddy-conf`) + LE cert via IONOS DNS-01; verified `/health` 200 and `/munich-home/dashboard`
|
||
(live weather) through Caddy. Cert mirrored to VPS (`trmnl.fullchain.crt`/`trmnl.privkey.key`) + traefik router `trmnl-public`
|
||
→ `192.168.8.211:9851` added to `/opt/traefik-dynamic.yaml`. **Public path pending**: VPS↔home netbird route was down
|
||
at provision time (`No networks available`, 3/6 peers — artifacto/blog also 504); resolves when the mesh route recovers.
|
||
**LAN pending**: Technitium A record not yet added. Not SOPS-enrolled; Google/MVG creds pending.
|
||
|
||
'
|
||
entity_slug: lxc:trmnl
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: trmnl
|
||
ip: 192.168.8.211
|
||
privileged: unprivileged
|
||
resources: 1 core / 768 MiB RAM / 8 GiB rootfs
|
||
mounts: none
|
||
public_host: trmnl.hubris.network (via VPS ingress)
|
||
changelog:
|
||
- date: '2026-06-24'
|
||
title: auto-deploy + LAN DNS wired
|
||
body: Gitea Shape-B deploy pipeline (webhook id 12, `:9797`) — push to `dtoro/terminalito` redeploys; verified end-to-end.
|
||
Technitium A record `trmnl.hubris.network → 192.168.8.175` added on [dns (107)](107-dns.md) (propagated to the NetBird
|
||
managed zone via dns-sync), so LAN clients take the short path through [Caddy (121)](121-caddy.md). See [auto-deploy](../infrastructure/auto-deploy.md).
|
||
- date: '2026-06-24'
|
||
title: public path live
|
||
body: 'Verified end-to-end from the internet: `https://trmnl.hubris.network/munich-home/dashboard` → 200 with token, 401
|
||
without; `/health` 200. The provision-time outage was the netbird `home-lab-network` (192.168.8.0/24) route having no
|
||
active routing peer — the **mac-mini routing peer''s netbird daemon was down** (artifacto/blog were 504 too). Bringing
|
||
netbird up on mac-mini restored the route; the edge recovered with no config change. See [ingress](../infrastructure/ingress.md)
|
||
/ [mesh](../infrastructure/mesh.md).'
|
||
- date: '2026-06-24'
|
||
title: provisioned
|
||
body: 'LXC 128 created (Debian 13, unprivileged, `192.168.8.211`). Deployed `trmnl-plugins.service` on :9851 from `dtoro/terminalito`.
|
||
Caddy block added (`dtoro/caddy-conf`) + LE cert via IONOS DNS-01; verified `/health` 200 and `/munich-home/dashboard`
|
||
(live weather) through Caddy. Cert mirrored to VPS (`trmnl.fullchain.crt`/`trmnl.privkey.key`) + traefik router `trmnl-public`
|
||
→ `192.168.8.211:9851` added to `/opt/traefik-dynamic.yaml`. **Public path pending**: VPS↔home netbird route was down
|
||
at provision time (`No networks available`, 3/6 peers — artifacto/blog also 504); resolves when the mesh route recovers.
|
||
**LAN pending**: Technitium A record not yet added. Not SOPS-enrolled; Google/MVG creds pending.'
|
||
- slug: containers/129-house
|
||
title: 129 — `house`
|
||
content: '# 129 — `house`
|
||
|
||
|
||
Yuvomi family planner (formerly Oikos). Self-hosted family planner with 14 modules: calendar, tasks, meals, groceries,
|
||
budget, documents, notes, contacts, birthdays, housekeeping, recipes, reminders.
|
||
|
||
|
||
## At a glance
|
||
|
||
|
||
- **Hostname:** `house`
|
||
|
||
- **IP:** `192.168.8.244`
|
||
|
||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||
|
||
- **Privilege:** unprivileged
|
||
|
||
- **Resources:** 1 core / 1344 MiB RAM / 8 GiB rootfs (Debian 13)
|
||
|
||
- **Mounts:** none
|
||
|
||
- **Public hostname:** [`house.hubris.network`](../infrastructure/ingress.md) → VPS traefik → Caddy
|
||
|
||
|
||
## Service / port map
|
||
|
||
|
||
| Service | Listen | Notes |
|
||
|
||
|---------|--------|-------|
|
||
|
||
| `oikos` (Yuvomi) | `0.0.0.0:3000` | Docker Compose at `/opt/yuvomi/`, image `ghcr.io/ulsklyc/yuvomi` |
|
||
|
||
|
||
## Integrations
|
||
|
||
|
||
- **Authentik SSO (OIDC):** Provider `Provider for Yuvomi` (PK 31) in Authentik on VPS. Env vars in `/opt/yuvomi/.env`:
|
||
`OIDC_ISSUER`, `OIDC_CLIENT_ID`, `OIDC_CLIENT_SECRET`. Redirect URI: `https://house.hubris.network/auth/oidc/callback`.
|
||
|
||
- **Paperless DMS connector (native):** Yuvomi connects directly to Paperless-ngx API at `http://192.168.8.130:8000/`.
|
||
API token stored in SQLite `dms_accounts` table. Search, link, and upload documents from Yuvomi to Paperless via Settings
|
||
→ Documents → DMS.
|
||
|
||
- **Weather widget:** Open-Meteo (free, no API key). Munich coordinates set.
|
||
|
||
- **Google Calendar:** OAuth configured via env vars (`GOOGLE_CLIENT_ID`, `GOOGLE_CLIENT_SECRET`, `GOOGLE_REDIRECT_URI`).
|
||
Redirect URI: `https://house.hubris.network/api/v1/calendar/google/callback`. Authorize in Settings → Calendar → Connect
|
||
Google Calendar.
|
||
|
||
|
||
## Config paths
|
||
|
||
|
||
- `/opt/yuvomi/docker-compose.yml` — downloaded from upstream
|
||
|
||
- `/opt/yuvomi/.env` — config including secrets (untracked)
|
||
|
||
- `/opt/yuvomi/data/` — SQLCipher SQLite DB (`oikos.db`)
|
||
|
||
- `/opt/yuvomi/backups/` — auto backups
|
||
|
||
- `/opt/yuvomi/modules/` — Yuvomi modules (empty for now)
|
||
|
||
|
||
## Related
|
||
|
||
|
||
- [Caddy (121)](121-caddy.md) — LAN reverse proxy (`house.hubris.network → 192.168.8.212:3000`)
|
||
|
||
- [VPS ingress](../infrastructure/ingress.md) — public edge (cert mirror + traefik router)
|
||
|
||
- [DNS (107)](107-dns.md) — Technitium A record `house → 192.168.8.175` (LAN path via Caddy)
|
||
|
||
- [Paperless (103)](103-paperless.md) — native DMS connector (API at `:8000`)
|
||
|
||
- [TRMNL (128)](128-trmnl.md) — Google Calendar tokens source
|
||
|
||
- [Deployment plan](../../../plans/done/2026-06-25-yuvomi-deployment.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-27 — Google Calendar OAuth env vars configured
|
||
|
||
`GOOGLE_CLIENT_ID`, `GOOGLE_CLIENT_SECRET`, `GOOGLE_REDIRECT_URI` set in `.env`. New OAuth client ID (`-bho4iq..`).
|
||
|
||
|
||
### 2026-06-26 — provisioned
|
||
|
||
|
||
LXC 129 created (Debian 13, unprivileged, `192.168.8.212`). Docker installed. Yuvomi container running on `:3000` from
|
||
`ghcr.io/ulsklyc/yuvomi:latest`. Caddy block + DNS A record + VPS traefik router `house-public` for public access. Authentik
|
||
OIDC provider created (PK 31). WebDAV document bridge on paperless LXC (103) at `:8088` for Paperless auto-import.'
|
||
entity_slug: lxc:house
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: house
|
||
ip: 192.168.8.244
|
||
host: strong
|
||
privileged: unprivileged
|
||
resources: 1 core / 1344 MiB RAM / 8 GiB rootfs
|
||
mounts: none
|
||
public_host: house.hubris.network → VPS traefik → Caddy
|
||
changelog:
|
||
- date: '2026-06-27'
|
||
title: Google Calendar OAuth env vars configured
|
||
body: '`GOOGLE_CLIENT_ID`, `GOOGLE_CLIENT_SECRET`, `GOOGLE_REDIRECT_URI` set in `.env`. New OAuth client ID (`-bho4iq..`).'
|
||
- date: '2026-06-26'
|
||
title: provisioned
|
||
body: LXC 129 created (Debian 13, unprivileged, `192.168.8.212`). Docker installed. Yuvomi container running on `:3000`
|
||
from `ghcr.io/ulsklyc/yuvomi:latest`. Caddy block + DNS A record + VPS traefik router `house-public` for public access.
|
||
Authentik OIDC provider created (PK 31). WebDAV document bridge on paperless LXC (103) at `:8088` for Paperless auto-import.
|
||
- slug: containers/130-grimmory
|
||
title: 130 — `grimmory`
|
||
content: '# 130 — `grimmory`
|
||
|
||
|
||
Self-hosted digital library (eBooks, comics, audiobooks). Community fork/successor of Booklore, with smart shelves, metadata
|
||
enrichment, Kobo/KOReader sync, OPDS, and a built-in EPUB/PDF reader. Migrated from [apps (105)](105-apps.md) on 2026-06-29.
|
||
|
||
|
||
## At a glance
|
||
|
||
|
||
- **Hostname:** `grimmory`
|
||
|
||
- **IP:** `192.168.8.247`
|
||
|
||
- **Host:** **strong** (migrated from hubris 2026-07-05)
|
||
|
||
- **Privilege:** privileged (UID = host UID for `/mnt/library` media GID)
|
||
|
||
- **Resources:** 1 core / 2 GiB RAM / 16 GiB rootfs (Debian 13)
|
||
|
||
- **Mounts:** `/mnt/media_local` ↔ `/mnt/library`
|
||
|
||
- **Public hostname:** `books.hubris.network`
|
||
|
||
|
||
## Service / port map
|
||
|
||
|
||
| Service | Listen | Notes |
|
||
|
||
|---------|--------|-------|
|
||
|
||
| Grimmory | `192.168.8.247:6060` | Docker Compose at `/opt/grimmory/` |
|
||
|
||
| MariaDB | internal only | Sidecar in the same compose stack |
|
||
|
||
|
||
## Compose
|
||
|
||
|
||
Located at `/opt/grimmory/docker-compose.yml`. Key points:
|
||
|
||
|
||
- Image: `ghcr.io/grimmory-tools/grimmory:latest`
|
||
|
||
- Books library: `/mnt/library/books` → `/books` (read-write; media GID 10000 via `GROUP_ID=10000`)
|
||
|
||
- Bookdrop (watched folder for auto-import): `/opt/grimmory/bookdrop` → `/bookdrop`
|
||
|
||
- App data (covers, DB config): `/opt/grimmory/data` → `/app/data`
|
||
|
||
- MariaDB config: `/opt/grimmory/mariadb/config` → `/config` (linuxserver/mariadb image)
|
||
|
||
- `extra_hosts: auth.hubris.network:192.168.8.175` — routes Authentik OIDC discovery to Caddy from inside the container
|
||
|
||
- `FORCE_DISABLE_OIDC=false` — OIDC stays enabled; provider configured via Grimmory admin UI
|
||
|
||
|
||
Credentials live in `/opt/grimmory/.env` (untracked):
|
||
|
||
- `DATABASE_PASSWORD` / `MYSQL_PASSWORD` — MariaDB Grimmory user password
|
||
|
||
- `MYSQL_ROOT_PASSWORD` — MariaDB root password
|
||
|
||
|
||
## Authentik OIDC
|
||
|
||
|
||
Uses Confidential client (client secret stored in Grimmory''s DB — migrated from Booklore). The OIDC config carried over
|
||
in the database dump; no manual re-entry needed.
|
||
|
||
|
||
- **Authentik provider:** `Provider for Grimmory` (renamed from `Provider for Booklore` on migration)
|
||
|
||
- **Client ID:** `L1u0eFsNhbKgiIvvFeIr2mvZdbtFyzidCq2h6thL`
|
||
|
||
- **Client type:** Confidential (client secret in `oidc_provider_details` in MariaDB `app_settings`)
|
||
|
||
- **Redirect URI:** `https://books.hubris.network/oauth2-callback`
|
||
|
||
- **Scopes:** openid, profile, email, offline_access
|
||
|
||
- **Back-channel logout:** `http://192.168.8.247:6060/api/v1/auth/oidc/backchannel-logout`
|
||
|
||
- **Application slug:** `booklore` → Issuer URI: `https://auth.hubris.network/application/o/booklore/`
|
||
|
||
|
||
## Media permissions
|
||
|
||
|
||
LXC is privileged → in-container UID = host UID. Docker container gets media GID via `GROUP_ID=10000` env var (Grimmory/linuxserver
|
||
pattern). The `/mnt/library/books` subtree is owned `:media` mode `2775` (setgid). See [media-permissions](../infrastructure/media-permissions.md).
|
||
|
||
|
||
## Related
|
||
|
||
|
||
- [apps (105)](105-apps.md) — previous host (Booklore)
|
||
|
||
- [Caddy (121)](121-caddy.md) — `books.hubris.network → 192.168.8.247:6060`
|
||
|
||
- [Authentik (124)](106-auth-outpost.md) — OIDC provider `Grimmory`
|
||
|
||
- [DNS (107)](107-dns.md) — `books.hubris.network → 192.168.8.175` (unchanged from Booklore)
|
||
|
||
- [Media permissions](../infrastructure/media-permissions.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-29 — provisioned; Booklore migrated
|
||
|
||
|
||
LXC 130 created (Debian 13, privileged, `192.168.8.213`). Docker installed. Grimmory compose deployed at `/opt/grimmory/`.
|
||
MariaDB dump from Booklore (LXC 105) restored — schema-compatible since Grimmory is a direct fork. Caddy `books.hubris.network`
|
||
backend updated from `192.168.8.205:6060` to `192.168.8.213:6060`. Authentik provider updated: Booklore → Grimmory, Confidential
|
||
→ Public (PKCE). Booklore stack removed from Portainer on LXC 105.
|
||
|
||
'
|
||
entity_slug: lxc:grimmory
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: grimmory
|
||
ip: 192.168.8.247
|
||
host: strong
|
||
privileged: privileged
|
||
resources: 1 core / 2 GiB RAM / 16 GiB rootfs
|
||
mounts: /mnt/media_local /mnt/library
|
||
public_host: books.hubris.network
|
||
changelog:
|
||
- date: '2026-06-29'
|
||
title: provisioned; Booklore migrated
|
||
body: 'LXC 130 created (Debian 13, privileged, `192.168.8.213`). Docker installed. Grimmory compose deployed at `/opt/grimmory/`.
|
||
MariaDB dump from Booklore (LXC 105) restored — schema-compatible since Grimmory is a direct fork. Caddy `books.hubris.network`
|
||
backend updated from `192.168.8.205:6060` to `192.168.8.213:6060`. Authentik provider updated: Booklore → Grimmory,
|
||
Confidential → Public (PKCE). Booklore stack removed from Portainer on LXC 105.'
|
||
- slug: containers/131-teddycloud
|
||
title: 131 — `teddycloud`
|
||
content: "# 131 — `teddycloud`\n\nSelf-hosted [TeddyCloud](https://github.com/toniebox-reverse-engineering/teddycloud),\
|
||
\ a\nreimplementation of the Toniebox cloud backend — lets Tonie figurines play custom/ripped\naudio content against a\
|
||
\ local server instead of the official cloud.\n\nPredates the client-enrollment convention entirely; nobody wrote it down.\
|
||
\ Found and\ndocumented on 2026-07-06 after Oikos's drift detector (`oikos/drift.py`) flagged\n`pve_id 131` as live on\
|
||
\ hubris (via `pct list`) with no `inventory.yaml` entry — see\n[OIKOS.md](../../../.agents/OIKOS.md)'s Week 3 build-status\
|
||
\ note. `containers/132-rclone.md` had already\nmentioned it in passing (\"LXC 131 was already taken by an undocumented\
|
||
\ `teddycloud`\ncontainer\"), and `hosts/strong.md`'s 2026-07-05 migration changelog fixed a DHCP conflict\nfor it — but\
|
||
\ it never got its own inventory entry or doc page until now.\n\n## At a glance\n\n- **Hostname:** `teddycloud`\n- **Host:**\
|
||
\ hubris (confirmed via `pct config 131` run directly on hubris — the original\n drift finding's `pct list` source)\n\
|
||
- **IP:** `192.168.8.150` (static; was briefly `192.168.8.243` via DHCP until the\n 2026-07-05 strong-migration work\
|
||
\ assigned it a fixed address — see hosts/strong.md)\n- **Privilege:** unconfirmed (not checked — read-only investigation\
|
||
\ didn't need it)\n- **Resources:** 1 core / 1 GiB RAM / 512 MiB swap / 16 GiB rootfs (`local-lvm`), Debian 12\n (bookworm)\n\
|
||
- **Mounts:** `/mnt/library` (`mp0`)\n- **Public hostname:** `teddy.hubris.network` → `192.168.8.150:8443` (see\n [caddy\
|
||
\ (121)](121-caddy.md))\n- **Enrollment:** none — no `age_pubkey`, not a `homelab` CLI client. It's a plain\n docker-compose\
|
||
\ app container, not a fleet-managed host. No action needed unless it starts\n needing secrets.\n\n## Service\n\nRuns\
|
||
\ via `docker compose` at `/opt/teddycloud` — container `teddycloud-teddycloud-1`,\nimage `ghcr.io/toniebox-reverse-engineering/teddycloud:latest`,\
|
||
\ publishing `80`, `443`,\n`8080`, and `8443`. Caddy routes `teddy.hubris.network` to the `:8443` port.\n\n## Risk notes\n\
|
||
\n**No Caddy forward-auth gate** — unlike `sab.hubris.network` on the same Caddyfile (which\nis explicitly annotated `(Authentik\
|
||
\ forward-auth)`), `teddy.hubris.network` has no auth\nannotation. It's reachable to anyone on the LAN/mesh who can resolve\
|
||
\ the hostname. Not\naddressed as part of this doc pass — flagging it here since it's now visible in one place\nfor the\
|
||
\ first time.\n\n## Related\n\n- [Hubris host](../hosts/hubris.md)\n- [Caddy (121)](121-caddy.md) — terminates `teddy.hubris.network`\n\
|
||
- [rclone (132)](132-rclone.md) — landed on pve_id 132 specifically because 131 was already\n taken by this container\n\
|
||
- [Containers index](index.md)\n- [OIKOS.md](../../../.agents/OIKOS.md) — drift detector that caught this\n\n## Changelog\n\
|
||
\n### 2026-07-06 — documented for the first time (drift-caught)\n\nAdded to `inventory.yaml` and given this page. Verified\
|
||
\ live via read-only `pct config 131`\n+ `pct exec 131 -- ...` on hubris: hostname, IP, resources, and that it runs via\n\
|
||
`docker compose` (not a raw binary or systemd unit). No changes made to the running\ncontainer — this is pure documentation\
|
||
\ catch-up.\n"
|
||
entity_slug: lxc:teddycloud
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: teddycloud
|
||
host: hubris (confirmed via pct config 131 run directly on hubris — the original
|
||
ip: 192.168.8.150 (static; was briefly 192.168.8.243 via DHCP until the
|
||
privileged: unconfirmed
|
||
resources: 1 core / 1 GiB RAM / 512 MiB swap / 16 GiB rootfs (local-lvm), Debian 12
|
||
mounts: /mnt/library
|
||
public_host: teddy.hubris.network → 192.168.8.150:8443 (see
|
||
enrollment: none — no age_pubkey, not a homelab CLI client. It's a plain
|
||
changelog:
|
||
- date: '2026-07-06'
|
||
title: documented for the first time (drift-caught)
|
||
body: 'Added to `inventory.yaml` and given this page. Verified live via read-only `pct config 131` + `pct exec 131 --
|
||
...` on hubris: hostname, IP, resources, and that it runs via `docker compose` (not a raw binary or systemd unit). No
|
||
changes made to the running container — this is pure documentation catch-up.'
|
||
- slug: containers/132-rclone
|
||
title: 132 — `rclone`
|
||
content: "# 132 — `rclone`\n\nOff-host backup appliance. Mirrors selected `/mnt/library` folders to **Proton Drive**\nwith\
|
||
\ a plain `rclone sync` (monthly), and serves rclone's Web GUI on the LAN for browsing\nand ad-hoc runs. **Replaces**\
|
||
\ the disabled restic-on-USB job — see [backups](../infrastructure/backups.md).\n\nProvisioned 2026-07-01. (LXC 131 was\
|
||
\ already taken by an undocumented `teddycloud` container,\nso this landed on **132**.)\n\n## At a glance\n\n- **Hostname:**\
|
||
\ `rclone`\n- **IP:** `192.168.8.214` (static, set in PVE `net0` config — same pattern as grimmory/authentik)\n- **Privilege:**\
|
||
\ privileged (root in-container = host root → reads every `/mnt/library` subtree,\n incl. `homecloud/` and `documents/`,\
|
||
\ regardless of owner)\n- **Resources:** 1 core / 2 GiB RAM / 8 GiB rootfs (Debian 13) — bumped from 1 GiB on 2026-07-03\n\
|
||
\ after `rclone-rcd.service` was OOM-killed under real load (see \"Known issue\" below)\n- **Mounts:** `/mnt/library`\
|
||
\ **read-only** (`mp0: /mnt/library,mp=/mnt/library,ro=1`) — a backup\n job must never be able to write into the library\n\
|
||
- **Public hostname:** none — the UI is **LAN-only, no auth** (by design)\n\n## Service / port map\n\n| Service | Listen\
|
||
\ | Notes |\n|---------|--------|-------|\n| rclone Web GUI (`rcd`) | `192.168.8.214:5572` | `rclone-rcd.service`, **`--rc-no-auth`**,\
|
||
\ LAN-only. Browse `/mnt/library` + `proton:`, run ad-hoc syncs, live job status |\n| monthly mirror | — | `rclone-backup.service`\
|
||
\ + `.timer` (`OnCalendar=*-*-01 03:00`) |\n\n## Backup design\n\n- **Mode:** plain mirror — `rclone sync` (Proton mirrors\
|
||
\ local; deletions propagate; **no versioning**).\n- **Encryption:** Proton Drive's built-in E2E only (no rclone `crypt`\
|
||
\ overlay → files stay\n browsable in Proton's web UI).\n- **Selected set:** `/etc/rclone-backup/folders.list` — one\
|
||
\ absolute source path per line\n (`#`/blank ignored). This file *is* the picked set the monthly timer mirrors. Extensible\
|
||
\ to other\n disks once bind-mounted into this LXC.\n- **Path mapping:** source `S` → `proton:library-backup/<S without\
|
||
\ leading slash>`\n (e.g. `/mnt/library/notes` → `proton:library-backup/mnt/library/notes`).\n- **Runner:** `/usr/local/sbin/rclone-backup.sh\
|
||
\ [folder ...]` (Python, despite the `.sh` name — kept\n the path stable) — no arg = every enabled line. Submits each\
|
||
\ folder as an **async job through the\n rclone rc API** served by `rclone-rcd.service` (the same daemon backing the\
|
||
\ Web GUI on `:5572`),\n so scheduled/ad-hoc runs show up live in the GUI's **Jobs panel**, not just in logs. Gentle\
|
||
\ on\n Proton's rate limits (`Transfers=4, TPSLimit=8, FastList=true` via the rc `_config` payload). The\n rc API here\
|
||
\ requires **POST for every call** including `job/status` and `core/stats` — GET with\n query params 404s.\n- **Logs\
|
||
\ / \"past runs\":** per-run logs in `/var/log/rclone-backup/<safe>-<ts>.log`; one-line\n JSON summary per run appended\
|
||
\ to `/var/log/rclone-backup/runs.jsonl`.\n- **Failure notify:** `OnFailure=rclone-backup-notify@%n.service` → logs to\
|
||
\ journal today;\n **TODO** wire to Hermes `send_message` (Matrix) per [backups](../infrastructure/backups.md).\n\n##\
|
||
\ rclone + Proton Drive\n\n- **rclone** installed from the official binary (not apt) so the `protondrive` backend is present\n\
|
||
\ (`rclone v1.74.3`).\n- Remote **`proton:`** (type `protondrive`). Config at `/root/.config/rclone/rclone.conf`, mode\
|
||
\ 600.\n **This file is a secret** (holds the obscured Proton password + TOTP secret + session) — **never\n commit it.**\
|
||
\ Escrow the Proton account creds in the password manager.\n- **Config gotchas** (from rclone docs/forum):\n - Log into\
|
||
\ Proton via a **browser at least once** first, or key generation fails.\n - For unattended runs, store the **TOTP _secret_**\
|
||
\ (not a 6-digit code) so rclone self-generates\n codes; obscure with `rclone obscure`.\n - Passwords with **extended-ASCII**\
|
||
\ characters are known to break auth.\n - Proton's API is rate-limited → keep `--transfers`/`--tpslimit` conservative\
|
||
\ (baked into the runner).\n- **DR escrow (pending):** store the Proton creds as sops secret `secrets/protondrive.yaml`,\
|
||
\ granted\n to this LXC's age key, so the remote can be rebuilt after a re-provision.\n\n## The UI (rclone Web GUI)\n\
|
||
\n`rclone rcd --rc-web-gui --rc-no-auth --rc-addr 0.0.0.0:5572` (assets auto-downloaded on first\nstart). Reach it at\
|
||
\ **http://192.168.8.214:5572** on the LAN.\n\n> **Security note:** `--rc-no-auth` exposes *full* rclone control — including\
|
||
\ deleting remote data —\n> to anyone on the LAN (accepted per the design choice). The container has only a LAN NIC, so\
|
||
\ it is\n> not publicly reachable. Harden later by adding `--rc-user/--rc-pass` or fronting it with Authentik.\n\n## Tracked\
|
||
\ config (deferred)\n\n**Not yet tracked.** The runner, systemd units, and `folders.list` currently live as plain files\n\
|
||
directly on the LXC — fully functional, just not version-controlled or auto-deployed. A\n`dtoro/rclone` gitea repo package\
|
||
\ (runner, units, `install.sh`, webhook receiver) is pre-built and\nstaged at `/root/rclone-repo` on the LXC for whenever\
|
||
\ this gets tracked (Shape A, like\n[caddy](121-caddy.md)). Gitea `ALLOWED_HOST_LIST` already includes `192.168.8.214`\
|
||
\ in anticipation.\nSee [auto-deploy](../infrastructure/auto-deploy.md).\n\n**Selected folders (live in `/etc/rclone-backup/folders.list`):**\
|
||
\ `/mnt/library/cloud` (287G),\n`/mnt/library/documents` (249M), `/mnt/library/repos` (83M). `/mnt/library/notes` was\
|
||
\ synced once as\na connectivity test (not in the recurring set). Proton quota checked: 2 TiB plan, ~1.65 TiB free\nafter\
|
||
\ this set.\n\n## Enrollment gotcha: `pct exec` PATH\n\n`pct exec` (lxc-attach) does **not** source `/etc/environment`\
|
||
\ or run a login shell, so\n`/usr/local/bin` (where bootstrap installs `sops`) isn't on `$PATH` by default — bootstrap's\
|
||
\ own\n`command -v sops` post-install check failed under `pct exec` even though the binary installed fine.\nFixed by symlinking\
|
||
\ `/usr/local/bin/{sops,homelab}` into `/usr/bin` (always on the minimal PATH),\nrather than relying on `/etc/environment`.\
|
||
\ Same category as the documented [`pct exec` no-initgroups\ngotcha](../infrastructure/media-permissions.md#gotchas) —\
|
||
\ worth adding to\n[agent-enrollment.md troubleshooting](../../../.agents/operations/agent-enrollment.md#troubleshooting)\
|
||
\ if it recurs\non future LXC bootstraps.\n\n## Known issue: `rclone-rcd.service` OOM-killed under 1 GiB RAM (root cause,\
|
||
\ resolved)\n\nWhat looked like repeated \"protondrive silently stalls\" was actually **`rclone-rcd.service` (the rc\n\
|
||
API daemon backing the Web GUI and, since the rc-API redesign, all actual sync work) getting\nOOM-killed** under the original\
|
||
\ 1 GiB RAM allocation — `journalctl` confirms\n`A process of this unit has been killed by the OOM killer` at the exact\
|
||
\ moment a transfer had\n\"frozen.\" systemd's own `Restart=on-failure` (5s) auto-respawns it, but every in-flight job's\
|
||
\ state\nis lost on the kill, which looked identical to a silent backend hang from the outside (frozen\n`core/stats`,\
|
||
\ no new log lines). **Fix: bumped the LXC's memory to 2 GiB** (`pct set 132 -memory\n2048` — applies live via the host\
|
||
\ cgroup, confirmed via `cat /sys/fs/cgroup/lxc/132/memory.max` on\nhubris, no container reboot needed). After the bump,\
|
||
\ the full folder set (`cloud` 287G, `documents`\n249M, `repos` 83M) completed cleanly with no further kills.\n\n**`rclone-backup-watchdog.timer`**\
|
||
\ (every 5 min) → `rclone-backup-watchdog.sh`: if\n`rclone-backup.service` is active but total transferred bytes (global\
|
||
\ `core/stats` on the rc API)\nhaven't moved for 15 minutes, it restarts both `rclone-rcd.service` (clears any stuck/orphaned\
|
||
\ job —\nthis is the actual daemon holding the work, not the thin wrapper) and then `rclone-backup.service`\n(`--no-block`,\
|
||
\ load-bearing — see below). Kept as a safety net even after the RAM fix, in case\nmemory pressure returns under a larger\
|
||
\ folder set later. State kept in\n`/var/lib/rclone-backup/watchdog-state.json`, cleared whenever the service isn't running.\n\
|
||
\n**Two watchdog design bugs found and fixed while chasing this (2026-07-03):**\n1. **Wrong stats-group key.** Per-job\
|
||
\ progress polling queried `core/stats` under `job/<jobid>`,\n but rclone tracks stats under whatever `_group` name\
|
||
\ the job was submitted with. Made a perfectly\n healthy sync look stalled at 0 bytes for 22+ hours in its own log.\
|
||
\ Fixed by using the same\n `group` variable consistently. **Lesson: distrust the per-run log's \"progress bytes=\"\
|
||
\ line during\n an incident; cross-check with unfiltered `core/stats` first.**\n2. **Watchdog restarted only the thin\
|
||
\ wrapper, and blocked doing it.** The actual `rclone sync` work\n runs inside `rclone-rcd.service`, not `rclone-backup.service`\
|
||
\ — restarting the wrapper alone left\n any stuck job orphaned inside `rcd` while a new wrapper submitted a duplicate\
|
||
\ job on top. Worse,\n `systemctl restart rclone-backup.service` (no `--no-block`) blocks until the *new* invocation's\n\
|
||
\ long-running `ExecStart` exits — which could be hours — so the watchdog's own oneshot service\n never logged \"\
|
||
Finished,\" and `OnUnitActiveSec` (which schedules relative to the previous run\n *finishing*) never fired again. The\
|
||
\ watchdog silently disabled itself after exactly one use.\n Fixed: restart `rclone-rcd.service` first, then `rclone-backup.service`\
|
||
\ with `--no-block`.\n\n## Related\n\n- [Backups](../infrastructure/backups.md) — this job supersedes the disabled restic-on-USB\
|
||
\ backup\n- [Hubris host](../hosts/hubris.md) — owns `/mnt/library`\n- [Media permissions](../infrastructure/media-permissions.md)\
|
||
\ — read-only consumer of `/mnt/library`\n- [Containers index](index.md)\n\n## Changelog\n\n### 2026-07-03 — root cause\
|
||
\ found (OOM, not Proton); RAM bumped to 2 GiB; full folder set completed\n\nWhat looked like repeated silent \"protondrive\
|
||
\ stalls\" turned out to be **`rclone-rcd.service`\ngetting OOM-killed** under the original 1 GiB RAM allocation — confirmed\
|
||
\ via\n`journalctl -u rclone-rcd.service` showing `killed by the OOM killer` at the exact freeze point.\nBumped the LXC\
|
||
\ to 2 GiB RAM (live, no reboot). After the bump: `cloud` (287G) completed cleanly\n(exit 0), `documents` (249M) completed\
|
||
\ with 1 minor error (259.7 MB transferred), `repos` (83M)\ncompleting as of this entry. Also fixed two real watchdog\
|
||
\ bugs found while chasing this (wrong\nstats-group key making a healthy sync look frozen; watchdog restarting only the\
|
||
\ wrapper with a\nblocking `systemctl restart`, causing it to silently disable itself after one use) — see \"Known\nissue\"\
|
||
\ above for full detail. The watchdog is kept as a safety net going forward even though the RAM\nbump addresses the actual\
|
||
\ root cause.\n\n### 2026-07-02 — runner rewritten to submit jobs via the rc API (GUI job visibility)\n\nThe original\
|
||
\ runner (`rclone sync` invoked as a standalone CLI subprocess) was invisible to the Web\nGUI's Jobs panel — the GUI only\
|
||
\ tracks work submitted through its own `rcd` process. Rewrote\n`/usr/local/sbin/rclone-backup.sh` in Python, submitting\
|
||
\ each folder via `POST /sync/sync` with\n`_async: true` against `http://127.0.0.1:5572` (the running `rclone-rcd.service`),\
|
||
\ then polling\n`POST /job/status` + `POST /core/stats` (both **must be POST** — GET-with-querystring 404s on this\nrc\
|
||
\ API) until finished, logging periodic progress snapshots and the same `runs.jsonl` summary line\nas before. Verified\
|
||
\ live: submitted job visible in `POST /job/list`'s `runningIds` while running,\ncompleted cleanly (`success: true`) once\
|
||
\ done. Deployed via atomic rename (write-then-`mv`) rather\nthan truncating in place, specifically so it wouldn't risk\
|
||
\ corrupting the still-running original\n`cloud`+`documents`+`repos` sync mid-flight (verified after the fact: that sync's\
|
||
\ bash process was\nunaffected, kept running to completion under the old in-memory script content). The already-running\n\
|
||
scheduled sync from before this change is a standalone process and won't retroactively appear in the\nGUI; every run after\
|
||
\ this point will.\n\n### 2026-07-02 — Proton Drive auth fixed; real folder set enabled; first live sync\n\nInitial `rclone\
|
||
\ config` failed 2FA (`422 ... auth/v4/2fa`) because a live 6-digit TOTP code was\nentered instead of the TOTP secret\
|
||
\ — reconfigured with the secret, auth now works\n(`rclone lsd proton:` lists the Drive). Verified end-to-end with a real\
|
||
\ sync of `/mnt/library/notes`\n(219 objects, 5.964 MiB, exit 0) — confirmed files land as plain, browsable objects on\
|
||
\ Proton (not\nan opaque archive), matching the plain-mirror + Proton-E2E design. Checked Proton quota (2 TiB\nplan, 1.945\
|
||
\ TiB free) before enabling a large folder. `folders.list` set to the real selection:\n`cloud` (287G), `documents` (249M),\
|
||
\ `repos` (83M); a full sync of that set was kicked off via the\nactual `rclone-backup.service` unit (not an ad-hoc call)\
|
||
\ to validate the real monthly path early\nrather than waiting for the Aug 1 timer. Tracked-repo step (`dtoro/rclone`\
|
||
\ on gitea) deferred by\nchoice — runner/units/`folders.list` remain plain files on the LXC for now; the repo package\
|
||
\ stays\nstaged at `/root/rclone-repo` for later.\n\n### 2026-07-01 — provisioned; enrolled\n\nLXC 132 created (Debian\
|
||
\ 13, privileged, `192.168.8.214`, `/mnt/library` read-only). rclone v1.74.3\ninstalled from the official binary (`protondrive`\
|
||
\ backend present). Runner + monthly timer +\n`folders.list` deployed; rclone Web GUI (`rcd`, LAN-only no-auth) live on\
|
||
\ `:5572`. Enrolled into\nhomelab-context (`--no-mesh`, LAN-only issuance): age key issued, inventory finalized, shared\n\
|
||
secrets granted, `homelab whoami` + `homelab secret hello` verified. Gitea `ALLOWED_HOST_LIST`\nupdated to include `192.168.8.214`.\
|
||
\ Hit and fixed a `pct exec` PATH gotcha (see below). Proton Drive\nremote, `dtoro/rclone` tracked repo + webhook, and\
|
||
\ the `secrets/protondrive.yaml` escrow remain\noperator-run follow-ups (credentialed steps — Proton password/2FA, repo\
|
||
\ creation). Restic-on-USB\nbackup deprecated in the same change.\n"
|
||
entity_slug: lxc:rclone
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: rclone
|
||
ip: 192.168.8.214
|
||
privileged: privileged (root in-container = host root → reads every /mnt/library subtree,
|
||
resources: 1 core / 2 GiB RAM / 8 GiB rootfs (Debian 13) — bumped from 1 GiB on 2026-07-03
|
||
mounts: '/mnt/library read-only (mp0: /mnt/library,mp=/mnt/library,ro=1) — a backup'
|
||
public_host: none — the UI is LAN-only, no auth
|
||
changelog:
|
||
- date: '2026-07-03'
|
||
title: root cause found (OOM, not Proton); RAM bumped to 2 GiB; full folder set completed
|
||
body: 'What looked like repeated silent "protondrive stalls" turned out to be **`rclone-rcd.service` getting OOM-killed**
|
||
under the original 1 GiB RAM allocation — confirmed via `journalctl -u rclone-rcd.service` showing `killed by the OOM
|
||
killer` at the exact freeze point. Bumped the LXC to 2 GiB RAM (live, no reboot). After the bump: `cloud` (287G) completed
|
||
cleanly (exit 0), `documents` (249M) completed with 1 minor error (259.7 MB transferred), `repos` (83M) completing as
|
||
of this entry. Also fixed two real watchdog bugs found while chasing this (wrong stats-group key making a healthy sync
|
||
look frozen; watchdog restarting only the wrapper with a blocking `systemctl restart`, causing it to silently disable
|
||
itself after one use) — see "Known issue" above for full detail. The watchdog is kept as a safety net going forward
|
||
even though the RAM bump addresses the actual root cause.'
|
||
- date: '2026-07-02'
|
||
title: runner rewritten to submit jobs via the rc API (GUI job visibility)
|
||
body: 'The original runner (`rclone sync` invoked as a standalone CLI subprocess) was invisible to the Web GUI''s Jobs
|
||
panel — the GUI only tracks work submitted through its own `rcd` process. Rewrote `/usr/local/sbin/rclone-backup.sh`
|
||
in Python, submitting each folder via `POST /sync/sync` with `_async: true` against `http://127.0.0.1:5572` (the running
|
||
`rclone-rcd.service`), then polling `POST /job/status` + `POST /core/stats` (both **must be POST** — GET-with-querystring
|
||
404s on this rc API) until finished, logging periodic progress snapshots and the same `runs.jsonl` summary line as before.
|
||
Verified live: submitted job visible in `POST /job/list`''s `runningIds` while running, completed cleanly (`success:
|
||
true`) once done. Deployed via atomic rename (write-then-`mv`) rather than truncating in place, specifically so it wouldn''t
|
||
risk corrupting the still-running original `cloud`+`documents`+`repos` sync mid-flight (verified after the fact: that
|
||
sync''s bash process was unaffected, kept running to completion under the old in-memory script content). The already-running
|
||
scheduled sync from before this change is a standalone process and won''t retroactively appear in the GUI; every run
|
||
after this point will.'
|
||
- date: '2026-07-02'
|
||
title: Proton Drive auth fixed; real folder set enabled; first live sync
|
||
body: 'Initial `rclone config` failed 2FA (`422 ... auth/v4/2fa`) because a live 6-digit TOTP code was entered instead
|
||
of the TOTP secret — reconfigured with the secret, auth now works (`rclone lsd proton:` lists the Drive). Verified end-to-end
|
||
with a real sync of `/mnt/library/notes` (219 objects, 5.964 MiB, exit 0) — confirmed files land as plain, browsable
|
||
objects on Proton (not an opaque archive), matching the plain-mirror + Proton-E2E design. Checked Proton quota (2 TiB
|
||
plan, 1.945 TiB free) before enabling a large folder. `folders.list` set to the real selection: `cloud` (287G), `documents`
|
||
(249M), `repos` (83M); a full sync of that set was kicked off via the actual `rclone-backup.service` unit (not an ad-hoc
|
||
call) to validate the real monthly path early rather than waiting for the Aug 1 timer. Tracked-repo step (`dtoro/rclone`
|
||
on gitea) deferred by choice — runner/units/`folders.list` remain plain files on the LXC for now; the repo package stays
|
||
staged at `/root/rclone-repo` for later.'
|
||
- date: '2026-07-01'
|
||
title: provisioned; enrolled
|
||
body: 'LXC 132 created (Debian 13, privileged, `192.168.8.214`, `/mnt/library` read-only). rclone v1.74.3 installed from
|
||
the official binary (`protondrive` backend present). Runner + monthly timer + `folders.list` deployed; rclone Web GUI
|
||
(`rcd`, LAN-only no-auth) live on `:5572`. Enrolled into homelab-context (`--no-mesh`, LAN-only issuance): age key issued,
|
||
inventory finalized, shared secrets granted, `homelab whoami` + `homelab secret hello` verified. Gitea `ALLOWED_HOST_LIST`
|
||
updated to include `192.168.8.214`. Hit and fixed a `pct exec` PATH gotcha (see below). Proton Drive remote, `dtoro/rclone`
|
||
tracked repo + webhook, and the `secrets/protondrive.yaml` escrow remain operator-run follow-ups (credentialed steps
|
||
— Proton password/2FA, repo creation). Restic-on-USB backup deprecated in the same change.'
|
||
- slug: containers/133-seanime
|
||
title: seanime (LXC 133) — Seanime anime media server
|
||
content: "# seanime (LXC 133) — Seanime anime media server\n\n## Summary\n\n| Field | Value |\n|-------|-------|\n| VMID\
|
||
\ | 133 |\n| Host | strong |\n| Role | Anime media server (online streaming + torrent client) |\n| LAN IP | 192.168.8.248/28\
|
||
\ (vmbr1) |\n| Public URL | https://seanime.hubris.network |\n| Backend | Seanime v3.9.0, systemd service |\n| Port |\
|
||
\ 43211 |\n\n## Mounts\n\n- `/anime` → `/mnt/media_local/anime` (ludo-lvm, bind mount) — existing anime collection\n\n\
|
||
## Services\n\n### Seanime (native binary)\n\n- Binary: `/opt/seanime/bin/seanime`\n- Data: `/opt/seanime/data/`\n- Config:\
|
||
\ `/opt/seanime/data/config.toml`\n- Service: `systemctl status seanime`\n- Flags: `--host 0.0.0.0 --port 43211 --datadir\
|
||
\ /opt/seanime/data --disable-password`\n- Config highlights:\n - `secureMode = 'lax'` (allows non-local Caddy proxy)\n\
|
||
\ - `trustedProxies = ['192.168.8.0/24']`\n\n### Torrent client (qBittorrent on arriman)\n\n- Host: `192.168.8.245:8080`\n\
|
||
- Auth: subnet whitelist (no password needed from homelab LAN)\n- qBittorrent config updated: `AuthSubnetWhitelist=192.168.8.0/24`\n\
|
||
\n### Extensions installed\n\n**Online streaming (8):**\nHiAnime, AniWatch, KickAssAnime, Anicrush, Animo, AniNeko, Senshi,\
|
||
\ Sudatchi\n\n## Caddy\n\n- Config: `seanime.hubris.network` block in `/etc/caddy/Caddyfile` (dtoro/caddy-conf repo)\n\
|
||
- TLS: Let's Encrypt DNS-01 via IONOS API\n- Backend: `192.168.8.248:43211`\n\n## DNS\n\n- `seanime.hubris.network` A\
|
||
\ record → `192.168.8.175` (Caddy)\n\n## Changelog\n\n### 2026-07-05 — initial creation\n- LXC 133 created on strong (Debian\
|
||
\ 13, 2 cores / 2 GiB)\n- Seanime v3.9.0 installed as native binary + systemd service\n- /anime bind mount from ludo-lvm\n\
|
||
- qBittorrent on arriman configured as torrent client\n- Caddy reverse proxy + DNS records set up\n- Online streaming\
|
||
\ extensions installed"
|
||
entity_slug: lxc:seanime
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
host: strong
|
||
ip: 192.168.8.248/28
|
||
resources: 2 cores / 2 GiB RAM
|
||
mounts: /anime → /mnt/media_local/anime
|
||
public_host: seanime.hubris.network
|
||
backend: Seanime v3.9.0
|
||
port: '43211'
|
||
changelog:
|
||
- date: '2026-07-05'
|
||
title: initial creation
|
||
body: '- LXC 133 created on strong (Debian 13, 2 cores / 2 GiB) - Seanime v3.9.0 installed as native binary + systemd
|
||
service - /anime bind mount from ludo-lvm - qBittorrent on arriman configured as torrent client - Caddy reverse proxy
|
||
+ DNS records set up - Online streaming extensions installed'
|
||
- slug: containers/134-romm
|
||
title: 134 — `romm`
|
||
content: "# 134 — `romm`\n\nSelf-hosted ROM manager ([RomM](https://romm.app)). Browse, search, and play\nyour retro game\
|
||
\ library from the browser. Runs on Docker Compose with a\nMariaDB sidecar.\n\n## At a glance\n\n- **Hostname:** `romm`\n\
|
||
- **IP:** `192.168.8.249/28` (static, vmbr1 on strong)\n- **Privilege:** privileged\n- **Resources:** 1 core / 2 GiB RAM\
|
||
\ / 16 GiB rootfs (Debian 13, ludo-lvm)\n- **Mounts:** `/mnt/media_local` → `/mnt/library`\n- **Public hostname:** `roms.hubris.network`\n\
|
||
\n## Service / port map\n\n| Service | Listen | Notes |\n|---------|--------|-------|\n| RomM | `192.168.8.249:80`\
|
||
\ | HTTP (Caddy terminates TLS) |\n| MariaDB | internal only | Sidecar in the same compose stack |\n\n## Compose\n\
|
||
\nLocated at `/opt/romm/docker-compose.yml`. Key points:\n\n- Image: `rommapp/romm:latest`\n- DB sidecar: `mariadb:latest`\
|
||
\ with healthcheck\n- ROM library: `/mnt/library/roms` → `/romm/library` (writable)\n- Resources (covers, etc.): Docker\
|
||
\ named volume `romm_resources` → `/romm/resources`\n- Saves/states: `/opt/romm/assets` → `/romm/assets`\n- Config: `/opt/romm/config`\
|
||
\ → `/romm/config`\n- Auth key: auto-generated, stored in `/opt/romm/.env`\n\nEnvironment (`/opt/romm/.env`):\n- `DB_ROOT_PASSWD`\
|
||
\ / `DB_PASSWD` — MariaDB credentials\n- `DB_USER=romm-user` / `DB_NAME=romm`\n- `ROMM_AUTH_SECRET_KEY` — auto-generated\n\
|
||
\n## ROM library structure\n\nRomM expects `/mnt/library/roms/<platform>/<game>/<rom>`. Create platform\ndirectories as\
|
||
\ needed:\n\n```\n/mnt/media_local/roms/\n├── gba/\n│ └── Pokemon - Emerald/\n│ └── Pokemon Emerald.gba\n├── snes/\n\
|
||
│ └── Super Mario World/\n│ └── Super Mario World.sfc\n└── psx/\n └── Final Fantasy VII/\n └── Final\
|
||
\ Fantasy VII.bin\n```\n\n## Media permissions\n\nThe `/mnt/media_local/roms` directory is owned `root:media` with mode\
|
||
\ `2775`\n(setgid). New files inherit the `media` GID (10000). The LXC is privileged so\nno idmap block is needed — in-container\
|
||
\ UID/GID matches the host. Docker\ncontainers within the LXC run as-is (read-only mount).\n\n## Related\n\n- [Strong\
|
||
\ host](../hosts/strong.md)\n- [Caddy (121)](121-caddy.md) — `roms.hubris.network → 192.168.8.249:80`\n- [DNS (107)](107-dns.md)\
|
||
\ — `roms.hubris.network A 192.168.8.175`\n- [Media permissions](../infrastructure/media-permissions.md)\n- [RomM docs](https://docs.romm.app)\n\
|
||
\n## Changelog\n\n### 2026-07-05 — provisioned\n\nLXC 134 created on strong (Debian 13, privileged, `192.168.8.249/28`).\n\
|
||
Docker + Compose installed. RomM stack deployed at `/opt/romm/`.\nCreated `/mnt/media_local/roms` with setgid `media:GID=10000`\
|
||
\ mode `2775`.\nCaddy `roms.hubris.network` → `192.168.8.249:80`.\nDNS `roms A 192.168.8.175` added to Technitium.\nHubris\
|
||
\ /32 route for `.249` added to `50-strong-route`."
|
||
entity_slug: lxc:romm
|
||
tags:
|
||
- container
|
||
at_glance:
|
||
hostname: romm
|
||
ip: 192.168.8.249/28
|
||
privileged: privileged
|
||
resources: 1 core / 2 GiB RAM / 16 GiB rootfs
|
||
mounts: /mnt/media_local → /mnt/library
|
||
public_host: roms.hubris.network
|
||
changelog:
|
||
- date: '2026-07-05'
|
||
title: provisioned
|
||
body: LXC 134 created on strong (Debian 13, privileged, `192.168.8.249/28`). Docker + Compose installed. RomM stack deployed
|
||
at `/opt/romm/`. Created `/mnt/media_local/roms` with setgid `media:GID=10000` mode `2775`. Caddy `roms.hubris.network`
|
||
→ `192.168.8.249:80`. DNS `roms A 192.168.8.175` added to Technitium. Hubris /32 route for `.249` added to `50-strong-route`.
|
||
- slug: hosts/hubris
|
||
title: '`hubris` — Proxmox host'
|
||
content: "# `hubris` — Proxmox host\n\nProxmox VE host running 1 VM and 13 LXC containers — the whole homelab's\nworkloads\
|
||
\ still live here. As of 2026-07-01, hubris is node 1 of the 2-node\n`Homelab` cluster (see [Cluster](#cluster)); the\
|
||
\ second node is\n[strong](strong.md), which hosts nothing yet.\n\n## At a glance\n- **Role:** Proxmox VE 9.1.2 hypervisor\
|
||
\ (kernel `6.14.11-4-pve`)\n- **Hardware:** GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS (Phoenix APU), 12 vCPU / ~28 GiB\
|
||
\ RAM, 2× Samsung 990 EVO Plus NVMe (one SSD primary, one for `library` LVM). 2× Realtek RTL8125 NICs (`r8169`).\n- **BIOS:**\
|
||
\ 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See [investigations](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\n\
|
||
- **Uplink:** `vmbr1` (slave: `eno1`) → SODOLA switch → Fritz!Box 7590. DHCP-reserved `192.168.178.10/24`, gateway `192.168.178.1`.\n\
|
||
- **Homelab bridge:** `vmbr0` — portless internal bridge, `192.168.8.77/24` + `192.168.8.1/24` alias (LXC default gateway).\
|
||
\ All 16 LXCs and the HAOS VM are on `vmbr0`. Proxmox routes between `vmbr0` and `vmbr1`; Fritz!Box has a static route\
|
||
\ `192.168.8.0/24 → 192.168.178.10`.\n- **WiFi:** disabled 2026-06-02 — `wlp3s0` removed from `/etc/network/interfaces`,\
|
||
\ wpa config deleted. Was used as a failover to the now-retired Slate AX AP.\n- **Mesh:** Netbird `wt0` `100.122.38.109/16`.\
|
||
\ Resolver: `100.122.38.109` (the local netbird daemon, which forwards to LAN/upstream and learns `*.hubris.network` answers\
|
||
\ via that path). See [mesh](../infrastructure/mesh.md).\n- **UI:** `https://proxmox.hubris.network` (via [caddy](../containers/121-caddy.md))\
|
||
\ or `https://192.168.8.77:8006`.\n\n## Storage\n\n| Pool | Type | Size | Use \
|
||
\ |\n| ------------ | --------- | ----- | -------------------------------------- |\n| `local` | dir \
|
||
\ | ~95G | ISOs, templates, `/etc`, configs |\n| `local-lvm` | lvmthin | 856G | LXC/VM rootfs \
|
||
\ |\n| `library` | lvmthin | 3.7T | Backs `/mnt/library` ext4 (mounted as `/dev/mapper/library-library`)\
|
||
\ |\n\n`/mnt/library` holds the shared media + data pool: `anime`, `audiobooks`, `books`, `comics`, `documents`, `downloads`,\
|
||
\ `heaper`, `homecloud`, `images`, `marimo`, `movies`, `music`, `notes`, `podcasts`, `repos`, `roms`, `sophia`. Bind-mounted\
|
||
\ into every container that needs it. Permissions standard: [media GID 10000](../infrastructure/media-permissions.md).\n\
|
||
\n## Cluster\n\nMember of `Homelab`, a 2-node Proxmox cluster with [strong](strong.md)\n(cluster/OS hostname `strong`),\
|
||
\ formed 2026-07-01.\n\n- **Corosync ring0:** hubris's internal `192.168.8.77` (the `vmbr0` address).\n strong reaches\
|
||
\ it via the existing Fritz!Box static route\n (`192.168.8.0/24 → 192.168.178.10`) — no dedicated corosync link, just\
|
||
\ the\n household LAN. Fine for a home cluster; not latency-isolated.\n- **Quorum:** 2 nodes, 1 vote each, no QDevice\
|
||
\ tiebreaker. Quorum needs both\n votes — if either node is down (reboot, maintenance, network hiccup), the\n survivor's\
|
||
\ running guests keep working but `/etc/pve` goes read-only:\n no start/stop/create/edit until quorum returns. Decided\
|
||
\ to skip a QDevice\n for now; revisit if hubris's periodic reboots (BIOS/thermal work, see\n Quirks below) make this\
|
||
\ painful in practice.\n- **Storage:** `local` / `local-lvm` are the standard per-node default IDs\n (every node has\
|
||
\ its own, not actually shared). The `library` lvmthin pool\n is explicitly restricted to `nodes hubris` in `/etc/pve/storage.cfg`\
|
||
\ since\n it's a physical thinpool that only exists on this host's hardware.\n- strong currently hosts no LXCs/VMs —\
|
||
\ it exists solely as a cluster\n member so far. See [strong.md](strong.md) and the [library-SSD\n migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)\n\
|
||
\ for what comes next (physical drive move, service migration — not started).\n\n## Tenants\n\n### VMs\n- [108 — `haos-16.3`](../vms/108-haos.md)\
|
||
\ — Home Assistant OS, 4 GiB / 32 GiB\n\n### LXC containers\nSee [containers/index](../containers/index.md). 10 active\
|
||
\ on hubris (101, 118, 122, 129, 130 migrated to [strong](strong.md) 2026-07-05).\n\n## Boot-time tuning (load-bearing)\n\
|
||
\n- **TCP BBR + `fq` qdisc** in `/etc/sysctl.d/99-bbr.conf`: `net.ipv4.tcp_congestion_control=bbr`, `net.core.default_qdisc=fq`,\
|
||
\ socket buffers up to 64 MiB (`rmem_max`/`wmem_max`/`tcp_rmem`/`tcp_wmem`). `tcp_bbr` module pinned to load at boot via\
|
||
\ `/etc/modules-load.d/bbr.conf`. Improves single-TCP-stream throughput on jittery / lossy links (notably WiFi clients\
|
||
\ downloading from [114 — nextcloud](../containers/114-nextcloud.md)). Container traffic crosses this kernel, so every\
|
||
\ guest inherits the change.\n- **`cpu-epp.service`** (enabled) sets `scaling_governor=powersave` + EPP=`balance_power`\
|
||
\ at boot — drops idle Tctl from ~95 °C → ~60 °C, clocks from 4.4 GHz pinned → 1.1–2 GHz idle. In `amd-pstate=active`\
|
||
\ mode the governor MUST be `powersave`, not `performance`, for EPP to apply. Unit ordering: `After=sysinit.target` +\
|
||
\ `Before=pve-guests.service` so it runs before the guest fleet starts (the original `After=multi-user.target` left the\
|
||
\ hottest boot window on `performance`). Fixed 2026-04-22.\n- **Crash capture:** `/etc/sysctl.d/60-crash-capture.conf`\
|
||
\ (panic on oops/hardlockup/softlockup/rcu, auto-reboot 10 s), `/etc/modprobe.d/softdog.conf` (`soft_panic=1 soft_margin=60`),\
|
||
\ `/etc/systemd/system.conf.d/watchdog.conf` (`RuntimeWatchdogSec=15s`). Pstore traces collected to `/var/lib/systemd/pstore/`\
|
||
\ by `systemd-pstore.service`. **Caveat:** silent CPU lockups leave pstore empty.\n- **`rasdaemon`** (Debian pkg) collects\
|
||
\ MCE / memory / PCIe AER / thermal events to `/var/lib/rasdaemon/ras-mc_event.db`. Query with `ras-mc-ctl --summary`\
|
||
\ / `--errors`. (mcelog is retired in Debian 13 — don't go looking for it.)\n\n## Network performance baseline (2026-05-14)\n\
|
||
\n| Path | Throughput | Notes |\n| ---- | ---------- | ----- |\n| WAN (Ookla, from host) | ↓ 113.5 Mbit/s · ↑ 19.9 Mbit/s\
|
||
\ · 24.6 ms | Asymmetric residential link. 20 Mbit upload is the cap for public ingress (artifacto, blog via VPS traefik).\
|
||
\ |\n| `eno1` link | 1000 Mbit/s, Full duplex | Negotiated correctly to router. |\n| Host → LXC over `vmbr0` | **34.7\
|
||
\ Gbit/s** | Single TCP stream, iperf3 10 s, 0 retransmits. |\n| LXC → LXC over `vmbr0` | **34.8 Gbit/s** | Single TCP\
|
||
\ stream, iperf3 5 s, 0 retransmits. |\n\nIntra-host throughput is memory-bandwidth bound — not a bottleneck for any current\
|
||
\ workload.\n\n**eno1 error counters** (uptime 14 d 6 h at measurement time): `tx_packets` 99.99 M / `rx_packets` 96.08\
|
||
\ M; `rx_errors` 1,619,203 (~1.7 % of RX); `align_errors` 10,919; `tx_errors` 20. The 1.6 M rx_errors in 14 days is the\
|
||
\ noteworthy datapoint — likely cable / switch-port. Re-check this counter in ~1 week and confirm it is not still climbing\
|
||
\ at the same rate; if it is, swap the patch cable first.\n\n## Host services owned by external repos\n\n| What \
|
||
\ | Repo | Path on host |\n| ---------------------------------\
|
||
\ | -------------------------- | ----------------------------- |\n| `claudio-monitor` (5-min watchdog) | `dtoro/claudio-monitor`\
|
||
\ | `/opt/claudio-monitor` |\n| `backup-library` (restic) | `dtoro/backup-library` | `/opt/backup-library`\
|
||
\ *(disabled)* |\n| `hubris-public-cert-sync` | (script + systemd unit) | `/usr/local/bin/hubris-public-cert-sync.sh`,\
|
||
\ daily timer |\n\nSee [monitoring](../infrastructure/monitoring.md), [backups](../infrastructure/backups.md), [ingress](../infrastructure/ingress.md).\n\
|
||
\n## Quirks\n\n- `/etc/pve` is fuse — the Proxmox cluster filesystem, now genuinely cluster-synced (2-node) rather than\
|
||
\ the single-node-but-still-fuse case this note used to describe.\n- ZFS is **not** in use; storage is LVM-thin + ext4.\n\
|
||
- Two Realtek 8125 NICs use the in-tree `r8169` driver, not the OOT `r8125`.\n- Hardware is thermally marginal. NVMe sensors\
|
||
\ live near warn temp under load. Thermal pads installed on the SSDs 2026-04-23; host relocated to a better-ventilated\
|
||
\ spot 2026-04-29.\n- Non-ECC RAM: silent memory faults are possible; suspect DIMMs if `cpu-epp` is on and crashes still\
|
||
\ happen.\n- BIOS update path requires a Windows-To-Go USB (no LVFS, no Linux flasher).\n\n## Authorized SSH keys (root)\n\
|
||
\n- `root@hubris` (self, RSA) — local\n- `d.toro.v@pm.me` (ed25519) — user's iMac, added 2026-04-22\n- `root@strong` (RSA)\
|
||
\ — strong's cluster-join key, added 2026-07-01 so\n `pvecm add` could authenticate without a password prompt\n\nOpenSSH\
|
||
\ on `0.0.0.0:22`. Netbird's built-in SSH server is on `100.122.38.109:22022` and bypasses `authorized_keys` (OIDC/browser).\
|
||
\ See [SSH access](../infrastructure/ssh-access.md) for the dual-server gotcha.\n\n## Related\n- [Containers index](../containers/index.md)\n\
|
||
- [DNS split-horizon](../infrastructure/dns.md)\n- [Caddy reverse proxy](../containers/121-caddy.md)\n- [Mesh / Netbird](../infrastructure/mesh.md)\n\
|
||
- [Media permissions](../infrastructure/media-permissions.md)\n- [Monitoring](../infrastructure/monitoring.md)\n- [Backups\
|
||
\ (disabled)](../infrastructure/backups.md)\n- [Operations cheatsheet](../../../.agents/operations/commands.md)\n- [Investigation:\
|
||
\ 2026-04-21 crash loop](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md)\n- [strong — Proxmox host](strong.md)\n\
|
||
\n## Changelog\n\n### 2026-07-01 — strong joined as a 2nd cluster node (\"Homelab\")\nUser reformatted `strong` (formerly\
|
||
\ a Linux dev workstation, `192.168.178.181`) to Proxmox VE 9.2.3. Cluster/OS hostname on that box is `strong` (left as-is\
|
||
\ from install). Bootstrapped root SSH on strong from a one-time console password (installed hubris's existing trusted\
|
||
\ key set: `root@hubris`, `d.toro.v@pm.me`), then generated a keypair on strong and pre-authorized it here (`root@strong`)\
|
||
\ so `pvecm add 192.168.8.77 --use_ssh 1` (run from strong) could join without an interactive password prompt. No cabling/routing\
|
||
\ changes needed — strong reaches hubris's corosync address (`192.168.8.77`) via the existing Fritz!Box static route.\
|
||
\ Cluster now 2 nodes, quorate, **no QDevice** (explicit choice — see [Cluster](#cluster) above for the quorum tradeoff\
|
||
\ this implies). strong hosts no guests yet; this is Phase 1 of the [library-SSD migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md),\
|
||
\ nothing further from that plan has been executed.\n\n### 2026-06-02 — Slate AX retired; SODOLA switch added; network\
|
||
\ restructured\nReplaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network\
|
||
\ support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at\
|
||
\ `192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged\
|
||
\ LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed.\
|
||
\ Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md)\
|
||
\ and [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).\n\n### 2026-05-14 — LXC 109 (syncthing)\
|
||
\ decommissioned\nUser destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy\
|
||
\ 109 --purge` cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing`\
|
||
\ was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up.\
|
||
\ Entry moved to the \"recently destroyed\" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology);\
|
||
\ references stripped from [README](../../../README.md), [media-permissions](../infrastructure/media-permissions.md),\
|
||
\ [vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md).\n\n### 2026-05-14\
|
||
\ — network performance baseline captured\nFirst explicit speed snapshot: WAN ↓113.5 / ↑19.9 Mbit (24.6 ms), `eno1` 1\
|
||
\ Gb full-duplex negotiated, intra-host `vmbr0` ~34.7 Gbit/s host↔LXC and ~34.8 Gbit/s LXC↔LXC (single TCP stream, zero\
|
||
\ retransmits). `iperf3` + `speedtest-cli` installed on host. Noted `eno1` `rx_errors` at 1.62 M (~1.7 % of 96 M RX packets\
|
||
\ in 14 d uptime) plus 10.9 k `align_errors` — flagged for follow-up; expect to recheck the trend in ~1 week, suspect\
|
||
\ patch cable / switch port first if still climbing. See new \"Network performance baseline\" section above.\n\n### 2026-05-14\
|
||
\ — LXC 116 (heaper) decommissioned\nUser destroyed the heaper LXC. No `116.conf.bak` left behind in `/etc/pve/lxc/`.\
|
||
\ Data subtree `/mnt/library/heaper` (224 MiB) retained for archaeology — see [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology)\
|
||
\ and [media-permissions](../infrastructure/media-permissions.md).\n\n### 2026-05-13 — TCP BBR + fq qdisc enabled\n`/etc/sysctl.d/99-bbr.conf`\
|
||
\ switches `net.ipv4.tcp_congestion_control` from `cubic` to `bbr` and `net.core.default_qdisc` from `fq_codel` to `fq`.\
|
||
\ Also bumps `rmem_max`/`wmem_max` to 64 MiB and widens `tcp_rmem`/`tcp_wmem`. `tcp_bbr` module pinned at boot via `/etc/modules-load.d/bbr.conf`.\
|
||
\ Triggered by Nextcloud client downloads from a WiFi laptop pulling ~2 MB/s despite a 152 Mbps link — server-side baseline\
|
||
\ through Caddy with BBR is ~400 MB/s single-stream loopback, so any client-perceived single-stream improvement is pure\
|
||
\ congestion-control win. Touches every LXC's outbound TCP since they all share this kernel.\n\n### 2026-04-29 — relocated\
|
||
\ to better-ventilated spot\nUser physically moved the host to a new location with improved airflow. Post-move idle baseline\
|
||
\ (45 min uptime, light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C,\
|
||
\ nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C, ACPI zone 47–49 °C. Compares well against the 2026-04-23\
|
||
\ thermal-pad steady-state (nvme0 sensor1 60–61 °C). Watch the lifetime NVMe warning-time counter over the coming days\
|
||
\ for confirmation. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation).\n\
|
||
\n### 2026-04-28 — Phase 1 WiFi failover\nHost now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failover,\
|
||
\ metric 200) on the GL-AXT1800-714-5G AP. Installed `wpasupplicant`+`iw`; added `wlp3s0` stanza to `/etc/network/interfaces`\
|
||
\ with `wpa-conf`; ARP isolation sysctls in `post-up`. Built `wan-failover.service` to remove the vmbr0 default route\
|
||
\ on `eno1` carrier loss, since the bridge's carrier doesn't follow `eno1` (the LXC veths keep it `1`). LXC/VM guests\
|
||
\ are still LAN-only — Phase 2 will migrate them.\n\n### 2026-04-28 — wiki started\nThis wiki created. Live state at this\
|
||
\ date: 14 LXCs running (109 syncthing stopped), 1 VM, kernel `6.14.11-4-pve`, uptime 3 d 0 h post drive-removal A/B test.\
|
||
\ Compared to memory snapshot from a week ago, **destroyed**: LXC 100 (yunohost arr), 106 (flaresolverr), 107 (marimo),\
|
||
\ 110 (photoprism), 111 (karakeep), 112 (immich), 115 (reticulum). 100 + 106 destroyed per the planned 2026-04-21 \\*arr\
|
||
\ migration retention; the others removed since.\n\n### 2026-04-23 — SSD cooling + thermal pads installed\nThermal pads\
|
||
\ on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 60–61 °C, nvme1 38–40 °C. Zero new warning-time minutes\
|
||
\ after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict).\n\
|
||
\n### 2026-04-22 — drive removal A/B test\nRemoved external USB backup drive (Silicon Motion `090c:2320`). Disabled the\
|
||
\ four `backup-library*.timer` units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on\
|
||
\ the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive,\
|
||
\ repeated crashes despite UAS blacklist + mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the\
|
||
\ primary contributor; `cpu-epp` remains as belt-and-suspenders thermal protection. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\n\
|
||
\n### 2026-04-22 — `cpu-epp.service` ordering bug fixed\nWas `After=multi-user.target` + `WantedBy=multi-user.target`\
|
||
\ — queued behind `pve-guests.service`, so the hottest boot window (20+ guests starting on `performance`) preceded EPP\
|
||
\ application. Now `After=sysinit.target` + `Before=pve-guests.service`.\n\n### 2026-04-21 — crash-capture + RAS telemetry\
|
||
\ enabled\n`60-crash-capture.conf`, softdog `soft_panic=1`, RuntimeWatchdog 15 s. `rasdaemon` installed and enabled. Pure\
|
||
\ silicon hangs still leave no trace; this catches everything else.\n\n### 2026-04-21 — `cpu-epp.service` deployed\nPinned\
|
||
\ governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz.\
|
||
\ First fix in the [crash-loop incident](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\n"
|
||
entity_slug: host:hubris
|
||
tags:
|
||
- host
|
||
at_glance:
|
||
role: Proxmox VE 9.1.2 hypervisor
|
||
hardware: GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS (Phoenix APU), 12 vCPU / ~28 GiB RAM, 2× Samsung 990 EVO Plus NVMe
|
||
(one SSD primary, one for library LVM). 2× Realtek RTL8125 NICs (r8169).
|
||
bios: 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See investigations.
|
||
uplink: 'vmbr1 (slave: eno1) → SODOLA switch → Fritz!Box 7590. DHCP-reserved 192.168.178.10/24, gateway 192.168.178.1.'
|
||
homelab_bridge: vmbr0 — portless internal bridge, 192.168.8.77/24 + 192.168.8.1/24 alias (LXC default gateway). All 16
|
||
LXCs and the HAOS VM are on vmbr0. Proxmox routes between vmbr0 and vmbr1; Fritz!Box has a static route 192.168.8.0/24
|
||
→ 192.168.178.10.
|
||
wifi: disabled 2026-06-02 — wlp3s0 removed from /etc/network/interfaces, wpa config deleted. Was used as a failover to
|
||
the now-retired Slate AX AP.
|
||
mesh: 'Netbird wt0 100.122.38.109/16. Resolver: 100.122.38.109 (the local netbird daemon, which forwards to LAN/upstream
|
||
and learns *.hubris.network answers via that path). See mesh.'
|
||
ui: https://proxmox.hubris.network (via caddy) or https://192.168.8.77:8006.
|
||
changelog:
|
||
- date: '2026-07-01'
|
||
title: strong joined as a 2nd cluster node ("Homelab")
|
||
body: 'User reformatted `strong` (formerly a Linux dev workstation, `192.168.178.181`) to Proxmox VE 9.2.3. Cluster/OS
|
||
hostname on that box is `strong` (left as-is from install). Bootstrapped root SSH on strong from a one-time console
|
||
password (installed hubris''s existing trusted key set: `root@hubris`, `d.toro.v@pm.me`), then generated a keypair on
|
||
strong and pre-authorized it here (`root@strong`) so `pvecm add 192.168.8.77 --use_ssh 1` (run from strong) could join
|
||
without an interactive password prompt. No cabling/routing changes needed — strong reaches hubris''s corosync address
|
||
(`192.168.8.77`) via the existing Fritz!Box static route. Cluster now 2 nodes, quorate, **no QDevice** (explicit choice
|
||
— see [Cluster](#cluster) above for the quorum tradeoff this implies). strong hosts no guests yet; this is Phase 1 of
|
||
the [library-SSD migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md), nothing
|
||
further from that plan has been executed.'
|
||
- date: '2026-06-02'
|
||
title: Slate AX retired; SODOLA switch added; network restructured
|
||
body: 'Replaced GL.iNet Slate AX sub-router with SODOLA 5-Port 2.5Gbit managed switch. Fritz!OS 8.x lacks second-IP-network
|
||
support on LAN ports, so Proxmox now acts as the subnet router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at
|
||
`192.168.178.10/24`; `vmbr0` is a portless internal bridge holding all LXCs/VMs with `192.168.8.1` as an alias (unchanged
|
||
LXC gateway). Fritz!Box static route `192.168.8.0/24 → 192.168.178.10` enables inbound routing. No LXC configs changed.
|
||
Eliminated double-NAT. WiFi (`wlp3s0`) also removed — was pointing at the Slate AX SSID, no longer useful. See [network](../infrastructure/network.md)
|
||
and [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).'
|
||
- date: '2026-05-14'
|
||
title: LXC 109 (syncthing) decommissioned
|
||
body: User destroyed the syncthing LXC (had been stopped since 2026-04-21, never re-enabled). `pct destroy 109 --purge`
|
||
cleaned `vm-109-disk-0` on `local-lvm` and the `/etc/pve/lxc/109.conf` entry. Data subtree `/mnt/library/syncthing`
|
||
was already empty and retained as an empty dir. No DNS, Caddy, NFS-export, or claudio-monitor references to clean up.
|
||
Entry moved to the "recently destroyed" table in [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology);
|
||
references stripped from [README](../../../README.md), [media-permissions](../infrastructure/media-permissions.md),
|
||
[vms/100-zimaos](../vms/100-zimaos.md), and [containers/102-nfs-export](../containers/102-nfs-export.md).
|
||
- date: '2026-05-14'
|
||
title: network performance baseline captured
|
||
body: 'First explicit speed snapshot: WAN ↓113.5 / ↑19.9 Mbit (24.6 ms), `eno1` 1 Gb full-duplex negotiated, intra-host
|
||
`vmbr0` ~34.7 Gbit/s host↔LXC and ~34.8 Gbit/s LXC↔LXC (single TCP stream, zero retransmits). `iperf3` + `speedtest-cli`
|
||
installed on host. Noted `eno1` `rx_errors` at 1.62 M (~1.7 % of 96 M RX packets in 14 d uptime) plus 10.9 k `align_errors`
|
||
— flagged for follow-up; expect to recheck the trend in ~1 week, suspect patch cable / switch port first if still climbing.
|
||
See new "Network performance baseline" section above.'
|
||
- date: '2026-05-14'
|
||
title: LXC 116 (heaper) decommissioned
|
||
body: User destroyed the heaper LXC. No `116.conf.bak` left behind in `/etc/pve/lxc/`. Data subtree `/mnt/library/heaper`
|
||
(224 MiB) retained for archaeology — see [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology)
|
||
and [media-permissions](../infrastructure/media-permissions.md).
|
||
- date: '2026-05-13'
|
||
title: TCP BBR + fq qdisc enabled
|
||
body: '`/etc/sysctl.d/99-bbr.conf` switches `net.ipv4.tcp_congestion_control` from `cubic` to `bbr` and `net.core.default_qdisc`
|
||
from `fq_codel` to `fq`. Also bumps `rmem_max`/`wmem_max` to 64 MiB and widens `tcp_rmem`/`tcp_wmem`. `tcp_bbr` module
|
||
pinned at boot via `/etc/modules-load.d/bbr.conf`. Triggered by Nextcloud client downloads from a WiFi laptop pulling
|
||
~2 MB/s despite a 152 Mbps link — server-side baseline through Caddy with BBR is ~400 MB/s single-stream loopback, so
|
||
any client-perceived single-stream improvement is pure congestion-control win. Touches every LXC''s outbound TCP since
|
||
they all share this kernel.'
|
||
- date: '2026-04-29'
|
||
title: relocated to better-ventilated spot
|
||
body: 'User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime,
|
||
light load): k10temp Tctl **47.2 °C**, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite
|
||
38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C, ACPI zone 47–49 °C. Compares well against the 2026-04-23 thermal-pad steady-state
|
||
(nvme0 sensor1 60–61 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-29-physical-relocation).'
|
||
- date: '2026-04-28'
|
||
title: Phase 1 WiFi failover
|
||
body: 'Host now dual-homed: LAN `192.168.8.77` (primary) + WiFi `192.168.8.141` (failover, metric 200) on the GL-AXT1800-714-5G
|
||
AP. Installed `wpasupplicant`+`iw`; added `wlp3s0` stanza to `/etc/network/interfaces` with `wpa-conf`; ARP isolation
|
||
sysctls in `post-up`. Built `wan-failover.service` to remove the vmbr0 default route on `eno1` carrier loss, since the
|
||
bridge''s carrier doesn''t follow `eno1` (the LXC veths keep it `1`). LXC/VM guests are still LAN-only — Phase 2 will
|
||
migrate them.'
|
||
- date: '2026-04-28'
|
||
title: wiki started
|
||
body: 'This wiki created. Live state at this date: 14 LXCs running (109 syncthing stopped), 1 VM, kernel `6.14.11-4-pve`,
|
||
uptime 3 d 0 h post drive-removal A/B test. Compared to memory snapshot from a week ago, **destroyed**: LXC 100 (yunohost
|
||
arr), 106 (flaresolverr), 107 (marimo), 110 (photoprism), 111 (karakeep), 112 (immich), 115 (reticulum). 100 + 106 destroyed
|
||
per the planned 2026-04-21 \*arr migration retention; the others removed since.'
|
||
- date: '2026-04-23'
|
||
title: SSD cooling + thermal pads installed
|
||
body: Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 60–61 °C, nvme1 38–40 °C. Zero new
|
||
warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See
|
||
[investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md#2026-04-23-thermal-pad-verdict).
|
||
- date: '2026-04-22'
|
||
title: drive removal A/B test
|
||
body: 'Removed external USB backup drive (Silicon Motion `090c:2320`). Disabled the four `backup-library*.timer` units,
|
||
commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant
|
||
root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist
|
||
+ mount-on-demand. **Result so far:** 3+ days uptime — the drive looks like the primary contributor; `cpu-epp` remains
|
||
as belt-and-suspenders thermal protection. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).'
|
||
- date: '2026-04-22'
|
||
title: '`cpu-epp.service` ordering bug fixed'
|
||
body: Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `pve-guests.service`, so the hottest
|
||
boot window (20+ guests starting on `performance`) preceded EPP application. Now `After=sysinit.target` + `Before=pve-guests.service`.
|
||
- date: '2026-04-21'
|
||
title: crash-capture + RAS telemetry enabled
|
||
body: '`60-crash-capture.conf`, softdog `soft_panic=1`, RuntimeWatchdog 15 s. `rasdaemon` installed and enabled. Pure
|
||
silicon hangs still leave no trace; this catches everything else.'
|
||
- date: '2026-04-21'
|
||
title: '`cpu-epp.service` deployed'
|
||
body: Pinned governor=`powersave`, EPP=`balance_power` at boot. Stopped the host idling at ~95 °C with everything pinned
|
||
at 4.4 GHz. First fix in the [crash-loop incident](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||
- slug: hosts/strong
|
||
title: '`strong` — Proxmox host'
|
||
content: "# `strong` — Proxmox host\n\nSecond node in the `Homelab` cluster, alongside [hubris](hubris.md). Formerly\na\
|
||
\ Linux dev workstation nicknamed \"ludo\" (or \"ludo-mini\") — reformatted to\nProxmox VE on 2026-07-01. No LXCs/VMs\
|
||
\ deployed on it yet.\n\n## At a glance\n\n- **Role:** Proxmox VE 9.2.3 cluster member (kernel `7.0.12-1-pve`) — hosts\
|
||
\ [7 LXCs](../containers/index.md) migrated from hubris (Phase 1+2, 2026-07-05)\n- **Naming:** the OS/cluster hostname\
|
||
\ is `strong` (leftover from install,\n kept as-is). This wiki page and `inventory.yaml` also use `strong` as of\n 2026-07-01\
|
||
\ — earlier the same day the inventory entry was briefly named\n `ludo-mini`, but that was renamed so the hostname bootstrap\
|
||
\ needs\n (`hosts/$(hostname).yaml`) would just resolve. \"Ludo\"/\"ludo-mini\" remains\n the machine's everyday nickname;\
|
||
\ some older docs (investigations, the\n library-SSD migration plan) still refer to it that way — that's fine,\n those\
|
||
\ are historical.\n- **Hardware:** AMD Ryzen 7 PRO 6850U, 16 threads, 28 GiB RAM.\n - `nvme0n1` — MasonSemi MC3100 1TB\
|
||
\ (boot/OS): `pve-root` 96G, `pve-data`\n (thinpool) 815G, 8G swap.\n - `nvme1n1` — WD_BLACK SN7100 2TB, added 2026-07-01.\
|
||
\ Arrived with an\n existing APFS partition (previously used in a Mac) — wiped via\n `pvesh set nodes/strong/disks/wipedisk\
|
||
\ --disk /dev/nvme1n1`, then turned\n into its own LVM-thin pool via `pvesh create nodes/strong/disks/lvmthin`\n \
|
||
\ (handles pvcreate/vgcreate/lvcreate + storage.cfg registration in one\n step). Registered as Proxmox storage `ludo-lvm`\
|
||
\ (VG `ludo-lvm`, thinpool\n `ludo-lvm`, ~1.8 TiB usable — kept the `ludo-lvm` name since that's what\n was actually\
|
||
\ created live; not worth renaming), restricted `nodes strong`\n in `/etc/pve/storage.cfg` — same pattern as hubris's\
|
||
\ `library` pool.\n Empty so far; this is separate from the [library-SSD migration\n plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)'s\n\
|
||
\ planned drive move from hubris (that hasn't happened) — this is\n general-purpose VM/CT capacity.\n- **Network:**\
|
||
\ `vmbr0` is bridged straight onto the household LAN —\n `192.168.178.181/24`, gateway `192.168.178.1`. Unlike hubris,\
|
||
\ there is no\n internal/uplink bridge split yet — `vmbr0` here plays the role hubris's\n `vmbr1` plays there. Reachable\
|
||
\ from the homelab subnet (`192.168.8.0/24`)\n via the existing Fritz!Box static route through hubris; no new cabling\n\
|
||
\ or routing was needed to join the cluster.\n- **Storage:** `local` + `local-lvm` (cluster-wide default names, shared\n\
|
||
\ with hubris by convention, actually separate per-node volumes). Hubris's\n `library` lvmthin pool is restricted to\
|
||
\ `nodes hubris` in\n `/etc/pve/storage.cfg` and does not exist here.\n- **Mesh:** Netbird not installed — fresh OS wiped\
|
||
\ whatever the old\n workstation had. Reachable today only via LAN routing (confirmed DNS for\n `*.hubris.network` already\
|
||
\ resolves correctly here via `192.168.8.2`).\n Add to Netbird if off-LAN access to this host itself (distinct from any\n\
|
||
\ future guests) is needed.\n- **Homelab-context client enrollment:** done 2026-07-01 via\n `bootstrap.sh --no-secrets`\
|
||
\ (reused the operator's existing Gitea PAT for\n the initial clone). `/opt/homelab-context`, the `homelab` CLI, and\
|
||
\ the\n 5-min sync timer are live; `homelab whoami` resolves correctly. See\n [agent-enrollment.md](../../../.agents/operations/agent-enrollment.md).\n\
|
||
- **Age key / secrets:** issued the same day over plain LAN (no Netbird\n needed — see the `--no-mesh` bootstrap.sh fix\
|
||
\ below). Key lives at\n `/etc/age/key.txt`; pubkey `age1rtwvdct6avjkr3cyxv3vue3vqx4d524fjfr3vk7xrnvyrylnry5sm54sn4`\n\
|
||
\ recorded in `inventory.yaml`. Not yet a recipient on any actual secret\n (`hello.yaml`, `gitea-pat.yaml`, etc.) —\
|
||
\ that's a separate grant, see\n [\"Granting a secret to a new client\"](../../../.agents/operations/agent-enrollment.md#granting-a-secret-to-a-new-client).\n\
|
||
\n## Cluster membership\n\nJoined hubris's single-node cluster (`Homelab`) via `pvecm add` on\n2026-07-01. See [hosts/hubris.md#cluster](hubris.md#cluster)\
|
||
\ for the full\ncluster picture, node IDs, and the quorum tradeoff (2 nodes, no QDevice —\neither node going down freezes\
|
||
\ management on the survivor).\n\n## SSH\n\nRoot login via the same key set trusted on hubris (`root@hubris`,\n`d.toro.v@pm.me`)\
|
||
\ — installed 2026-07-01 by appending to\n`/root/.ssh/authorized_keys` (now symlinked to `/etc/pve/priv/authorized_keys`\n\
|
||
post cluster-join, so it's cluster-synced same as hubris). No password auth\nneeded going forward.\n\n## Related\n\n-\
|
||
\ [hubris — Proxmox host](hubris.md)\n- [Library SSD migration plan](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)\
|
||
\ — the larger project this is Phase 1 of (filename kept as-is, historical)\n- [Network](../infrastructure/network.md)\n\
|
||
- [SSH access](../infrastructure/ssh-access.md)\n- [Agent enrollment](../../../.agents/operations/agent-enrollment.md)\n\
|
||
\n## Changelog\n\n### 2026-07-05 — Phase 2d: grimmory (130) migrated to strong\nMigrated book library from hubris (192.168.8.213)\
|
||
\ to strong (192.168.8.247).\nRsync'd /books (2.6G) from hubris library SSD to ludo-lvm/media. Container\nmounts /mnt/media_local\
|
||
\ → /mnt/library locally. Caddy backend for\nbooks.hubris.network updated.\n\n### 2026-07-05 — Phase 2: arriman (122)\
|
||
\ + jellyfin (101) migrated; library on ludo-lvm\nMigrated arriman (→ 192.168.8.245) and jellyfin (→ 192.168.8.246) to\
|
||
\ strong.\nCreated 1.5T thin volume on ludo-lvm for media library (/mnt/media_local).\nRsync'd 363G of media data (movies,\
|
||
\ tv, anime, downloads, music) from hubris\nlibrary SSD to ludo-lvm. Both containers now mount /mnt/media_local directly\n\
|
||
as local ext4 — no NFS cross-node dependency. Jellyfin gets Radeon 680M iGPU\n(for VAAPI) + RX 7600 dGPU passthrough via\
|
||
\ dev0/dev1. Caddy backends updated\nfor media, jellyseerr, qbit, sab. Hubris freed 16 GiB RAM (8+8).\nDhcp scope narrowed\
|
||
\ to 192.168.8.100-239.\n\n### 2026-07-05 — house (LXC 129) migrated to strong; DHCP scope narrowed\nMigrated Yuvomi family\
|
||
\ planner from hubris (192.168.8.212) to strong\n(192.168.8.244). Followed same restore pattern as elementsynapse (Phase\
|
||
\ 1b).\nDiscovered IP conflict: teddycloud (LXC 131) had 192.168.8.243 via DHCP\n(scope was 192.168.8.241-254). Narrowed\
|
||
\ DHCP scope to 192.168.8.100-239,\ngave teddycloud static IP 192.168.8.150. Caddy updated: house.hubris.network\n→ 192.168.8.244:3000,\
|
||
\ teddy.hubris.network → 192.168.8.150:8443.\n\n### 2026-07-05 — elementsynapse (LXC 118) migrated to strong\nMigrated\
|
||
\ Matrix/Synapse + Element client from hubris (192.168.8.239) to a new\nLXC on strong (192.168.8.242). Added vmbr1 on\
|
||
\ strong at 192.168.8.241/28\n(portless internal bridge) for strong-hosted guests. Enabled IP forwarding,\nproxy ARP on\
|
||
\ hubris vmbr0, and specific /32 routes for strong's guest subnet.\nCaddy updated: element.hubris.network → 192.168.8.242:8080.\n\
|
||
VPS traefik updated: matrix.hubris.network backend → 192.168.8.242:8008.\nThis is Phase 1a of the strong migration plan\
|
||
\ — see .hermes/plans/2026-07-05_strong-migration-assessment.md.\n\n### 2026-07-01 — age key issued over LAN; 3 bugs found/fixed\
|
||
\ in bootstrap.sh\nRe-ran bootstrap without `--no-secrets` to get a real age key. Hit three real bugs live, fixed all\
|
||
\ three in `bootstrap.sh` and re-ran clean:\n1. The `mcp`-CLI pipx-install step and the (unused, `--with-hermes`-only)\
|
||
\ Goose installer both called `sudo -u <user>` unconditionally — fails with \"sudo: command not found\" on a minimal root-only\
|
||
\ image with no `sudo` binary at all. Added a `run_as()` helper that only shells out to `sudo` when there's a real distinct\
|
||
\ invoking user.\n2. `sops` isn't an apt/dnf package (matches what `agent-enrollment.md`'s manual-install recipe already\
|
||
\ does) — the auto-installer tried `apt-get install sops` and failed outright. Added `install_sops_binary()`, fetching\
|
||
\ the GitHub release binary directly on both dnf and apt paths.\n3. Bigger one: running without `--no-secrets` unconditionally\
|
||
\ tries to install + interactively connect Netbird (device-code SSO), even though the very next check already knows how\
|
||
\ to accept plain LAN reachability instead. Over SSH with nobody watching, this hangs forever — had to manually kill a\
|
||
\ stuck `netbird up` process. Added `--no-mesh`, which skips the Netbird install/connect step but keeps the LAN-fallback\
|
||
\ path for secrets issuance. This run used `bootstrap.sh --no-mesh` and completed cleanly: `mesh: lan`, age key installed,\
|
||
\ `mcp` CLI installed via pipx (proving fix #1 too).\n\nResult: age key at `/etc/age/key.txt`, pubkey recorded in `inventory.yaml`.\
|
||
\ Not yet granted access to any actual secret file — see the note above.\n\n### 2026-07-01 — enrolled as a homelab-context\
|
||
\ client\nRan `bootstrap.sh --no-secrets` (reused the operator's existing personal Gitea PAT for the initial clone rather\
|
||
\ than minting a fresh read-only one). Installed git, cloned `/opt/homelab-context`, installed the 5-min systemd sync\
|
||
\ timer, symlinked `homelab` CLI and `AGENTS.md`. Skipped age-key/secrets issuance and Netbird per operator choice — but\
|
||
\ bootstrap's own connectivity check reported `mesh: lan`, i.e. the secrets-issuance endpoint is already reachable over\
|
||
\ plain LAN, so re-running without `--no-secrets` later wouldn't require a Netbird join. Known gap: the `mcp` pipx CLI\
|
||
\ install step silently failed (`sudo: command not found` — bootstrap.sh's pipx step assumes a `sudo` binary even when\
|
||
\ already root; harmless, only affects the `homelab mcp <tool>` shell subcommand).\n\n### 2026-07-01 — inventory identity\
|
||
\ renamed ludo-mini → strong\nDiscovered while starting client enrollment: `bootstrap.sh` looks up\n`hosts/$(hostname).yaml`,\
|
||
\ and the OS hostname here is `strong`, not\n`ludo-mini`. Renaming the OS hostname was ruled out (already a cluster\n\
|
||
member — Proxmox doesn't support in-place node rename, only leave+rejoin).\nRenamed the wiki/inventory side instead: `inventory.yaml`\
|
||
\ key, this page\n(`hosts/ludo-mini.md` → `hosts/strong.md`), README, ssh-access.md all now\nsay `strong`. \"Ludo\"/\"\
|
||
ludo-mini\" is still fine as a spoken nickname.\n\n### 2026-07-01 — 2nd NVMe added; new LVM-thin pool `ludo-lvm`\nUser\
|
||
\ added a WD_BLACK SN7100 2TB (`nvme1n1`), previously used in a Mac\n(arrived with an EFI + APFS partition table — confirmed\
|
||
\ disposable, wiped).\nUsed Proxmox's own disk-management API rather than raw LVM commands:\n`pvesh set nodes/strong/disks/wipedisk\
|
||
\ --disk /dev/nvme1n1` to clear the old\npartition table/signatures, then `pvesh create nodes/strong/disks/lvmthin\n--name\
|
||
\ ludo-lvm --device /dev/disk/by-id/nvme-WD_BLACK_SN7100_2TB_251663803202\n--add_storage 1` to create the PV/VG/thinpool\
|
||
\ and register it as Proxmox\nstorage in one step. Result: storage ID `ludo-lvm`, ~1.8 TiB, `content\nrootdir,images`,\
|
||
\ `nodes strong` (mirrors hubris's `library` node-restriction\npattern — this pool only physically exists here). Empty\
|
||
\ — no VM/CT disks\nplaced on it yet.\n\n### 2026-07-01 — Proxmox install; joined Homelab cluster\nReformatted from Linux\
|
||
\ workstation to Proxmox VE 9.2.3. SSH keys seeded from\nhubris's trusted set (root password used once, then discarded).\
|
||
\ Joined the\nexisting `Homelab` cluster via `pvecm add 192.168.8.77 --use_ssh 1` from\nthis node's side, using key-based\
|
||
\ SSH pre-authorized in both directions —\nno interactive password prompt needed for the join itself. Cluster now 2\n\
|
||
nodes (`hubris`, `strong`), quorate, no QDevice. Decided to leave hostname\nas `strong` and skip a QDevice for now — both\
|
||
\ revisitable later.\n"
|
||
entity_slug: host:strong
|
||
tags:
|
||
- host
|
||
at_glance:
|
||
role: Proxmox VE 9.2.3 cluster member (kernel 7.0.12-1-pve) — hosts 7 LXCs migrated from hubris
|
||
naming: the OS/cluster hostname is strong (leftover from install,
|
||
hardware: AMD Ryzen 7 PRO 6850U, 16 threads, 28 GiB RAM.
|
||
network: vmbr0 is bridged straight onto the household LAN —
|
||
storage: local + local-lvm (cluster-wide default names, shared
|
||
mesh: Netbird not installed — fresh OS wiped whatever the old
|
||
homelab-context_client_enrollment: done 2026-07-01 via
|
||
age_key___secrets: issued the same day over plain LAN (no Netbird
|
||
changelog:
|
||
- date: '2026-07-05'
|
||
title: 'Phase 2d: grimmory (130) migrated to strong'
|
||
body: Migrated book library from hubris (192.168.8.213) to strong (192.168.8.247). Rsync'd /books (2.6G) from hubris library
|
||
SSD to ludo-lvm/media. Container mounts /mnt/media_local → /mnt/library locally. Caddy backend for books.hubris.network
|
||
updated.
|
||
- date: '2026-07-05'
|
||
title: 'Phase 2: arriman (122) + jellyfin (101) migrated; library on ludo-lvm'
|
||
body: Migrated arriman (→ 192.168.8.245) and jellyfin (→ 192.168.8.246) to strong. Created 1.5T thin volume on ludo-lvm
|
||
for media library (/mnt/media_local). Rsync'd 363G of media data (movies, tv, anime, downloads, music) from hubris library
|
||
SSD to ludo-lvm. Both containers now mount /mnt/media_local directly as local ext4 — no NFS cross-node dependency. Jellyfin
|
||
gets Radeon 680M iGPU (for VAAPI) + RX 7600 dGPU passthrough via dev0/dev1. Caddy backends updated for media, jellyseerr,
|
||
qbit, sab. Hubris freed 16 GiB RAM (8+8). Dhcp scope narrowed to 192.168.8.100-239.
|
||
- date: '2026-07-05'
|
||
title: house (LXC 129) migrated to strong; DHCP scope narrowed
|
||
body: 'Migrated Yuvomi family planner from hubris (192.168.8.212) to strong (192.168.8.244). Followed same restore pattern
|
||
as elementsynapse (Phase 1b). Discovered IP conflict: teddycloud (LXC 131) had 192.168.8.243 via DHCP (scope was 192.168.8.241-254).
|
||
Narrowed DHCP scope to 192.168.8.100-239, gave teddycloud static IP 192.168.8.150. Caddy updated: house.hubris.network
|
||
→ 192.168.8.244:3000, teddy.hubris.network → 192.168.8.150:8443.'
|
||
- date: '2026-07-05'
|
||
title: elementsynapse (LXC 118) migrated to strong
|
||
body: 'Migrated Matrix/Synapse + Element client from hubris (192.168.8.239) to a new LXC on strong (192.168.8.242). Added
|
||
vmbr1 on strong at 192.168.8.241/28 (portless internal bridge) for strong-hosted guests. Enabled IP forwarding, proxy
|
||
ARP on hubris vmbr0, and specific /32 routes for strong''s guest subnet. Caddy updated: element.hubris.network → 192.168.8.242:8080.
|
||
VPS traefik updated: matrix.hubris.network backend → 192.168.8.242:8008. This is Phase 1a of the strong migration plan
|
||
— see .hermes/plans/2026-07-05_strong-migration-assessment.md.'
|
||
- date: '2026-07-01'
|
||
title: age key issued over LAN; 3 bugs found/fixed in bootstrap.sh
|
||
body: 'Re-ran bootstrap without `--no-secrets` to get a real age key. Hit three real bugs live, fixed all three in `bootstrap.sh`
|
||
and re-ran clean: 1. The `mcp`-CLI pipx-install step and the (unused, `--with-hermes`-only) Goose installer both called
|
||
`sudo -u <user>` unconditionally — fails with "sudo: command not found" on a minimal root-only image with no `sudo`
|
||
binary at all. Added a `run_as()` helper that only shells out to `sudo` when there''s a real distinct invoking user.
|
||
2. `sops` isn''t an apt/dnf package (matches what `agent-enrollment.md`''s manual-install recipe already does) — the
|
||
auto-installer tried `apt-get install sops` and failed outright. Added `install_sops_binary()`, fetching the GitHub
|
||
release binary directly on both dnf and apt paths. 3. Bigger one: running without `--no-secrets` unconditionally tries
|
||
to install + interactively connect Netbird (device-code SSO), even though the very next check already knows how to accept
|
||
plain LAN reachability instead. Over SSH with nobody watching, this hangs forever — had to manually kill a stuck `netbird
|
||
up` process. Added `--no-mesh`, which skips the Netbird install/connect step but keeps the LAN-fallback path for secrets
|
||
issuance. This run used `bootstrap.sh --no-mesh` and completed cleanly: `mesh: lan`, age key installed, `mcp` CLI installed
|
||
via pipx (proving fix #1 too). Result: age key at `/etc/age/key.txt`, pubkey recorded in `inventory.yaml`. Not yet granted
|
||
access to any actual secret file — see the note above.'
|
||
- date: '2026-07-01'
|
||
title: enrolled as a homelab-context client
|
||
body: 'Ran `bootstrap.sh --no-secrets` (reused the operator''s existing personal Gitea PAT for the initial clone rather
|
||
than minting a fresh read-only one). Installed git, cloned `/opt/homelab-context`, installed the 5-min systemd sync
|
||
timer, symlinked `homelab` CLI and `AGENTS.md`. Skipped age-key/secrets issuance and Netbird per operator choice — but
|
||
bootstrap''s own connectivity check reported `mesh: lan`, i.e. the secrets-issuance endpoint is already reachable over
|
||
plain LAN, so re-running without `--no-secrets` later wouldn''t require a Netbird join. Known gap: the `mcp` pipx CLI
|
||
install step silently failed (`sudo: command not found` — bootstrap.sh''s pipx step assumes a `sudo` binary even when
|
||
already root; harmless, only affects the `homelab mcp <tool>` shell subcommand).'
|
||
- date: '2026-07-01'
|
||
title: inventory identity renamed ludo-mini → strong
|
||
body: 'Discovered while starting client enrollment: `bootstrap.sh` looks up `hosts/$(hostname).yaml`, and the OS hostname
|
||
here is `strong`, not `ludo-mini`. Renaming the OS hostname was ruled out (already a cluster member — Proxmox doesn''t
|
||
support in-place node rename, only leave+rejoin). Renamed the wiki/inventory side instead: `inventory.yaml` key, this
|
||
page (`hosts/ludo-mini.md` → `hosts/strong.md`), README, ssh-access.md all now say `strong`. "Ludo"/"ludo-mini" is still
|
||
fine as a spoken nickname.'
|
||
- date: '2026-07-01'
|
||
title: 2nd NVMe added; new LVM-thin pool `ludo-lvm`
|
||
body: 'User added a WD_BLACK SN7100 2TB (`nvme1n1`), previously used in a Mac (arrived with an EFI + APFS partition table
|
||
— confirmed disposable, wiped). Used Proxmox''s own disk-management API rather than raw LVM commands: `pvesh set nodes/strong/disks/wipedisk
|
||
--disk /dev/nvme1n1` to clear the old partition table/signatures, then `pvesh create nodes/strong/disks/lvmthin --name
|
||
ludo-lvm --device /dev/disk/by-id/nvme-WD_BLACK_SN7100_2TB_251663803202 --add_storage 1` to create the PV/VG/thinpool
|
||
and register it as Proxmox storage in one step. Result: storage ID `ludo-lvm`, ~1.8 TiB, `content rootdir,images`, `nodes
|
||
strong` (mirrors hubris''s `library` node-restriction pattern — this pool only physically exists here). Empty — no VM/CT
|
||
disks placed on it yet.'
|
||
- date: '2026-07-01'
|
||
title: Proxmox install; joined Homelab cluster
|
||
body: Reformatted from Linux workstation to Proxmox VE 9.2.3. SSH keys seeded from hubris's trusted set (root password
|
||
used once, then discarded). Joined the existing `Homelab` cluster via `pvecm add 192.168.8.77 --use_ssh 1` from this
|
||
node's side, using key-based SSH pre-authorized in both directions — no interactive password prompt needed for the join
|
||
itself. Cluster now 2 nodes (`hubris`, `strong`), quorate, no QDevice. Decided to leave hostname as `strong` and skip
|
||
a QDevice for now — both revisitable later.
|
||
- slug: vms/100-zimaos
|
||
title: 100 — `zimaos`
|
||
content: '# 100 — `zimaos`
|
||
|
||
|
||
ZimaOS (IceWhale / CasaOS-family NAS distro), installed as a Proxmox VM to evaluate it as a potential primary NAS frontend
|
||
in front of `/mnt/library` — alongside the existing fleet ([nextcloud (114)](../containers/114-nextcloud.md), [jellyfin
|
||
(101)](../containers/101-jellyfin.md), [mule-images (120)](../containers/120-mule-images.md)).
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Type:** QEMU VM
|
||
|
||
- **ZimaOS version:** `v1.6.1` (build 2026-04-21, kernel 6.12.25)
|
||
|
||
- **IP:** `192.168.8.195` (DHCP lease from the LAN router — see "Open items" below)
|
||
|
||
- **Hostname (LAN DNS):** [`zimaos.hubris.network`](../infrastructure/dns.md) → [caddy (121)](../containers/121-caddy.md)
|
||
`192.168.8.175` → VM `192.168.8.195:80`. TLS terminates at Caddy (Let''s Encrypt via IONOS DNS-01).
|
||
|
||
- **Resources:** 4 vCPU, 8 GiB RAM (balloon min 2 GiB), 64 GiB OS disk on `local-lvm`
|
||
|
||
- **BIOS / machine:** OVMF (UEFI), `q35`, `virtio-scsi-pci`
|
||
|
||
- **EFI disk:** none — ZimaOS installer writes a removable-media fallback (`/EFI/BOOT/BOOTX64.EFI`), so a persistent EFI
|
||
vars disk isn''t required. PVE warns "no efidisk configured! Using temporary efivars disk" on start; harmless.
|
||
|
||
- **ISO:** `/var/lib/vz/template/iso/zimaos-x86_64-1.6.1_installer.iso` (1.48 GB, from `IceWhaleTech/ZimaOS` GitHub releases;
|
||
upstream publishes SHA256 only for the `.img`/`.raucb` variants, not the `.iso`)
|
||
|
||
|
||
## Web UI
|
||
|
||
|
||
- **LAN:** <https://zimaos.hubris.network> (via [caddy (121)](../containers/121-caddy.md))
|
||
|
||
- **Netbird mesh (off-LAN):** same URL — works because hubris advertises `192.168.8.0/24` into the mesh, and once a peer
|
||
resolves `zimaos.hubris.network → 192.168.8.175`, the request tunnels home. DNS resolution from a mesh peer needs **either**
|
||
a netbird Management nameserver group (`hubris.network` match domain → `192.168.8.180`), or an `/etc/hosts` override on
|
||
the peer (`192.168.8.175 zimaos.hubris.network`). See [mesh](../infrastructure/mesh.md).
|
||
|
||
- **Direct fallback:** <http://192.168.8.195> — backend ZimaOS bundles its own Caddy (`Server: Caddy`, `Via: ZimaOS-Gateway`),
|
||
which the front-end Caddy on LXC 121 reverse-proxies to. Two Caddys in series.
|
||
|
||
|
||
## SSH
|
||
|
||
|
||
Disabled by default. To enable:
|
||
|
||
1. Toggle SSH on in the ZimaOS web UI ("Settings → SSH"), OR
|
||
|
||
2. From the VM console, switch to TTY2 with **Alt+F2** and run `passwd` to set a root password first.
|
||
|
||
|
||
## Storage strategy
|
||
|
||
|
||
**Initial assumption (didn''t work).** PVE 9''s native virtiofs share with the host''s existing `library` dir mapping
|
||
(used by [108-haos](108-haos.md)) was the obvious path. Attaching it took one line — but ZimaOS''s kernel ships *without*
|
||
the virtiofs module (`modinfo virtiofs` → `Module virtiofs not found`). The kvm-side device was attached fine; the guest
|
||
just can''t mount it. Detached on 2026-05-14.
|
||
|
||
|
||
**What we ended up with.** A dedicated NFS re-export LXC [102-nfs-export](../containers/102-nfs-export.md) bind-mounts
|
||
`/mnt/library` from the host and serves it over NFSv4 to ZimaOS. Performance is within ~2% of host-served NFS (LXC adds
|
||
namespace isolation only, not an IO hop), with the upside that no NFS/RPC daemons run on the bare-metal host. Read-only
|
||
during evaluation; promote to rw once the UI test confirms behavior.
|
||
|
||
|
||
**Filesystem layout inside ZimaOS.** `/etc/fstab` lives on a writable overlay so persistent mounts via fstab DO survive
|
||
reboots, but the *standard* path to add a network share is the ZimaOS Files UI (Connect Network Storage → NFS). UID/GID
|
||
mapping is squashed at the server to `www-data:media` (33:10000) — the same identity Nextcloud and mule-images use — so
|
||
any rw activity from ZimaOS records as a normal `media`-group write.
|
||
|
||
|
||
**Caveats discovered during install.**
|
||
|
||
- Root filesystem is **squashfs / read-only**; only `/DATA` is writable.
|
||
|
||
- `dtoro` identity is `uid 999, gid 1000(samba), groups samba+wheel`. Primary group is **`samba`** (not `dtoro` — that
|
||
group doesn''t exist). SSH key needs `chown dtoro:samba` to satisfy sshd StrictModes.
|
||
|
||
- `dtoro`''s home directory is literally `/DATA` (not `/home/dtoro`). Authorized keys go at `/DATA/.ssh/authorized_keys`.
|
||
|
||
- `/etc` is a writable overlay rooted at `/mnt/overlay/upper_etc`. Mounts added to `/etc/fstab` persist across reboots.
|
||
|
||
- `mount.nfs4` is a stripped busybox variant — rejects `_netdev`, `nofail`, `actimeo=N`. The working minimal fstab line:
|
||
`192.168.8.200:/mnt/library /media/library nfs rw,vers=4 0 0`. **Mount path matters** — see the 2026-05-15 Changelog entry
|
||
on why the NFS export is mounted at `/media/library` and not `/DATA/library`.
|
||
|
||
- No `qemu-guest-agent` is bundled — `qm guest cmd` calls fail. IP discovery via `qm monitor` + `screendump` instead.
|
||
|
||
- No package manager at runtime — no `apt`, no `opkg`. Apps come via ZimaOS''s CasaOS-style appstore only.
|
||
|
||
|
||
The alternative (dedicated virtual data disk on the `library` lvmthin pool, e.g. `qm set 100 --scsi1 library:1024`) was
|
||
rejected because it would start empty and force content migration; the NFS path keeps `/mnt/library` as the single source
|
||
of truth shared with the existing 8 LXCs.
|
||
|
||
|
||
## Open items
|
||
|
||
|
||
- **DHCP → static IP fixed (2026-06-03).** ZimaOS IP drifted from `.195` (Slate AX) → `.103` (Technitium) after the DHCP
|
||
migration, causing Caddy 502s. Fixed by injecting a static systemd-networkd config and restarting the VM. IP now pinned
|
||
at `192.168.8.195`. See [changelog](#2026-06-03--static-ip-set-to-195-dhcp-drift-fixed).
|
||
|
||
- **No Authentik wiring.** [authentik (124)](../containers/106-auth-outpost.md) isn''t enforcing auth in front of ZimaOS
|
||
yet — ZimaOS handles its own first-run wizard. The Caddyfile block uses bare `reverse_proxy` rather than the `import authentik`
|
||
pattern used by e.g. artifacto; layer it in once the wizard is complete and a static admin user exists.
|
||
|
||
- **No PBS backup.** No Proxmox Backup Server configured on hubris today; this VM is not backed up.
|
||
|
||
- **qemu-guest-agent not installed.** ZimaOS''s installer doesn''t bundle it, so `qm guest cmd 100 ...` returns "QEMU
|
||
guest agent is not running". IP discovery during this install was done via console screendump → `qm monitor` → `screendump`.
|
||
|
||
|
||
## Related
|
||
|
||
- [108 — `haos-16.3`](108-haos.md) — the existing VM, same q35/OVMF pattern; also a virtiofs `library` consumer
|
||
|
||
- [DNS](../infrastructure/dns.md) — split-horizon entry lives here
|
||
|
||
- [Caddy (121)](../containers/121-caddy.md) — would front this if/when promoted
|
||
|
||
- [Media permissions](../infrastructure/media-permissions.md) — relevant when wiring `/mnt/library`
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-03 — Static IP set to `.195`; DHCP drift fixed
|
||
|
||
ZimaOS had drifted from `.195` (Slate AX DHCP) → `.103` (Technitium DHCP), causing Caddy 502s. Injected `/etc/systemd/network/10-static.network`
|
||
into overlay (match `en*/eth*`, address `192.168.8.195/24`, gateway `.1`, DNS `.2`). VM restarted; verified reachable
|
||
at `.195`. Caddy (`zimaos.hubris.network`) now returns 200. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).
|
||
|
||
|
||
### 2026-05-15 — NFS mount relocated to `/media/library` (UI delete fix)
|
||
|
||
|
||
Symptom: deleting any file or folder inside the NFS-served `library` tree from the ZimaOS Files UI failed with `rename
|
||
... invalid cross-device link`. Root cause: ZimaOS''s "Move to Trash" is implemented as `rename(2)` into a `.trash/` directory
|
||
derived from the *drive root*, and icewhale-files identifies the drive root with the regex `^/media/([^/]+)`. With the
|
||
NFS mounted at `/DATA/library`, the visible UI path was `/media/ZimaOS-HD/library/foo` → drive `ZimaOS-HD` → trash at
|
||
`/media/ZimaOS-HD/.trash` (which lives on the local ext4 `/dev/sda8`). NFS → ext4 rename = `EXDEV`.
|
||
|
||
|
||
Fix: remount the NFS export at `/media/library` so it sits as its own `/media/<name>` segment. Now icewhale-files extracts
|
||
drive=`library`, computes trash at `/media/library/.trash`, and the rename stays within the NFS — no cross-device. The
|
||
local-storage daemon (LSBLK-based) doesn''t surface the NFS in the Files-app *sidebar* (drives sidebar only lists block
|
||
devices), but the icewhale-files daemon picks the mount up as its own "storage" tile in the Files view, and crucially
|
||
exposes a direct **Delete** action there (no trash flow, since ZimaOS treats non-LSBLK mounts as foreign and skips trash
|
||
for them). For a 4 TB NFS share, immediate delete is arguably the right behavior anyway.
|
||
|
||
|
||
Concrete change:
|
||
|
||
- `/etc/fstab` line changed from `… /DATA/library nfs rw,vers=4 0 0` to `… /media/library nfs rw,vers=4 0 0`.
|
||
|
||
- New mountpoint dir `/var/lib/casaos_data/.media/library` (= `/media/library` via the existing `/media` bind).
|
||
|
||
- Old `/DATA/library` mountpoint dir removed.
|
||
|
||
|
||
Side effect: navigating via `ZimaOS-HD → library` in the UI now shows an empty/missing folder. Users access the library
|
||
via the separate `library` tile in the Files app instead.
|
||
|
||
|
||
### 2026-05-14 — Mesh-reachable from netbird peers
|
||
|
||
Verified `https://zimaos.hubris.network` reachable from `republic-laptop` (100.122.78.140) through the existing hubris→netbird
|
||
subnet advertisement (`192.168.8.0/24`). No new infrastructure required — the path was already wired by [mesh](../infrastructure/mesh.md).
|
||
Open: configure a netbird Management nameserver group for `hubris.network` so peers don''t need per-device `/etc/hosts`
|
||
overrides.
|
||
|
||
|
||
### 2026-05-14 — Phase 2 promoted to rw
|
||
|
||
NFS export on [nfs-export (102)](../containers/102-nfs-export.md) flipped from `ro` to `rw` after Files UI evaluation
|
||
passed. Verified the squash: a write from ZimaOS appears on hubris''s `/mnt/library` as `www-data:media` (uid 33, gid
|
||
10000), matching the existing tree convention.
|
||
|
||
|
||
### 2026-05-14 — Phase 2: NFS mount via new LXC 102 (read-only)
|
||
|
||
Virtiofs path abandoned — ZimaOS kernel 6.12.25 ships without the virtiofs module (verified by `modinfo virtiofs` returning
|
||
"Module not found"). Pivoted to NFSv4: built [nfs-export (102)](../containers/102-nfs-export.md) as a dedicated 512-MiB
|
||
privileged Debian LXC bind-mounting `/mnt/library` and serving it on `192.168.8.200:2049` with `ro,all_squash,anonuid=33,anongid=10000`
|
||
to the LAN subnet. Detached the now-useless `virtiofs0` from VM 100. ZimaOS-side mount via Files UI (NFS protocol, server
|
||
`192.168.8.200`, path `/mnt/library`). Promote export to `rw` after Files UI evaluation passes.
|
||
|
||
|
||
### 2026-05-14 — Wired through Caddy (HTTPS)
|
||
|
||
Added `zimaos.hubris.network` site block to `/etc/caddy/Caddyfile` on [caddy (121)](../containers/121-caddy.md): bare
|
||
`reverse_proxy 192.168.8.195` + IONOS DNS-01 TLS, same pattern as plato/jellyfin. dnsmasq entry repointed from `192.168.8.195`
|
||
to `192.168.8.175`. Let''s Encrypt cert issued on first request. Caddy commit `a219176` pending push to `dtoro/caddy-conf`.
|
||
|
||
|
||
### 2026-05-14 — VM created, ZimaOS 1.6.1 installed (Phase 1)
|
||
|
||
`qm create 100` with q35/OVMF, no EFI disk, 4 vCPU / 8 GiB / 64 GiB on `local-lvm`. Installed via the official ISO (manual
|
||
console install). Web UI verified at `http://192.168.8.195`. `onboot=1`, `startup order=20`. dnsmasq entry `zimaos.hubris.network
|
||
→ 192.168.8.195` initially added direct-to-VM on [authentik (124)](../containers/106-auth-outpost.md) (later repointed
|
||
— see above). `/mnt/library` is **not** yet shared into the VM; Phase 2 (virtiofs) is gated on UI evaluation.
|
||
|
||
'
|
||
entity_slug: vm:zimaos
|
||
tags:
|
||
- vm
|
||
at_glance:
|
||
type: QEMU VM
|
||
zimaos_version: v1.6.1
|
||
ip: 192.168.8.195
|
||
hostname_(lan_dns): zimaos.hubris.network → caddy (121) 192.168.8.175 → VM 192.168.8.195:80. TLS terminates at Caddy (Let's
|
||
Encrypt via IONOS DNS-01).
|
||
resources: 4 vCPU, 8 GiB RAM (balloon min 2 GiB), 64 GiB OS disk on local-lvm
|
||
bios___machine: OVMF (UEFI), q35, virtio-scsi-pci
|
||
efi_disk: none — ZimaOS installer writes a removable-media fallback (/EFI/BOOT/BOOTX64.EFI), so a persistent EFI vars
|
||
disk isn't required. PVE warns "no efidisk configured! Using temporary efivars disk" on start; harmless.
|
||
iso: /var/lib/vz/template/iso/zimaos-x86_64-1.6.1_installer.iso
|
||
changelog:
|
||
- date: '2026-06-03'
|
||
title: Static IP set to `.195`; DHCP drift fixed
|
||
body: ZimaOS had drifted from `.195` (Slate AX DHCP) → `.103` (Technitium DHCP), causing Caddy 502s. Injected `/etc/systemd/network/10-static.network`
|
||
into overlay (match `en*/eth*`, address `192.168.8.195/24`, gateway `.1`, DNS `.2`). VM restarted; verified reachable
|
||
at `.195`. Caddy (`zimaos.hubris.network`) now returns 200. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).
|
||
- date: '2026-05-15'
|
||
title: NFS mount relocated to `/media/library` (UI delete fix)
|
||
body: 'Symptom: deleting any file or folder inside the NFS-served `library` tree from the ZimaOS Files UI failed with
|
||
`rename ... invalid cross-device link`. Root cause: ZimaOS''s "Move to Trash" is implemented as `rename(2)` into a `.trash/`
|
||
directory derived from the *drive root*, and icewhale-files identifies the drive root with the regex `^/media/([^/]+)`.
|
||
With the NFS mounted at `/DATA/library`, the visible UI path was `/media/ZimaOS-HD/library/foo` → drive `ZimaOS-HD`
|
||
→ trash at `/media/ZimaOS-HD/.trash` (which lives on the local ext4 `/dev/sda8`). NFS → ext4 rename = `EXDEV`. Fix:
|
||
remount the NFS export at `/media/library` so it sits as its own `/media/<name>` segment. Now icewhale-files extracts
|
||
drive=`library`, computes trash at `/media/library/.trash`, and the rename stays within the NFS — no cross-device. The
|
||
local-storage daemon (LSBLK-based) doesn''t surface the NFS in the Files-app *sidebar* (drives sidebar only lists block
|
||
devices), but the icewhale-files daemon picks the mount up as its own "storage" tile in the Files view, and crucially
|
||
exposes a direct **Delete** action there (no trash flow, since ZimaOS treats non-LSBLK mounts as foreign and skips trash
|
||
for them). For a 4 TB NFS share, immediate delete is arguably the right behavior anyway. Concrete change: - `/etc/fstab`
|
||
line changed from `… /DATA/library nfs rw,vers=4 0 0` to `… /media/library nfs rw,vers=4 0 0`. - New mountpoint dir
|
||
`/var/lib/casaos_data/.media/library` (= `/media/library` via the existing `/media` bind). - Old `/DATA/library` mountpoint
|
||
dir removed. Side effect: navigating via `ZimaOS-HD → library` in the UI now shows an empty/missing folder. Users access
|
||
the library via the separate `library` tile in the Files app instead.'
|
||
- date: '2026-05-14'
|
||
title: Mesh-reachable from netbird peers
|
||
body: 'Verified `https://zimaos.hubris.network` reachable from `republic-laptop` (100.122.78.140) through the existing
|
||
hubris→netbird subnet advertisement (`192.168.8.0/24`). No new infrastructure required — the path was already wired
|
||
by [mesh](../infrastructure/mesh.md). Open: configure a netbird Management nameserver group for `hubris.network` so
|
||
peers don''t need per-device `/etc/hosts` overrides.'
|
||
- date: '2026-05-14'
|
||
title: Phase 2 promoted to rw
|
||
body: 'NFS export on [nfs-export (102)](../containers/102-nfs-export.md) flipped from `ro` to `rw` after Files UI evaluation
|
||
passed. Verified the squash: a write from ZimaOS appears on hubris''s `/mnt/library` as `www-data:media` (uid 33, gid
|
||
10000), matching the existing tree convention.'
|
||
- date: '2026-05-14'
|
||
title: 'Phase 2: NFS mount via new LXC 102 (read-only)'
|
||
body: 'Virtiofs path abandoned — ZimaOS kernel 6.12.25 ships without the virtiofs module (verified by `modinfo virtiofs`
|
||
returning "Module not found"). Pivoted to NFSv4: built [nfs-export (102)](../containers/102-nfs-export.md) as a dedicated
|
||
512-MiB privileged Debian LXC bind-mounting `/mnt/library` and serving it on `192.168.8.200:2049` with `ro,all_squash,anonuid=33,anongid=10000`
|
||
to the LAN subnet. Detached the now-useless `virtiofs0` from VM 100. ZimaOS-side mount via Files UI (NFS protocol, server
|
||
`192.168.8.200`, path `/mnt/library`). Promote export to `rw` after Files UI evaluation passes.'
|
||
- date: '2026-05-14'
|
||
title: Wired through Caddy (HTTPS)
|
||
body: 'Added `zimaos.hubris.network` site block to `/etc/caddy/Caddyfile` on [caddy (121)](../containers/121-caddy.md):
|
||
bare `reverse_proxy 192.168.8.195` + IONOS DNS-01 TLS, same pattern as plato/jellyfin. dnsmasq entry repointed from
|
||
`192.168.8.195` to `192.168.8.175`. Let''s Encrypt cert issued on first request. Caddy commit `a219176` pending push
|
||
to `dtoro/caddy-conf`.'
|
||
- date: '2026-05-14'
|
||
title: VM created, ZimaOS 1.6.1 installed (Phase 1)
|
||
body: '`qm create 100` with q35/OVMF, no EFI disk, 4 vCPU / 8 GiB / 64 GiB on `local-lvm`. Installed via the official
|
||
ISO (manual console install). Web UI verified at `http://192.168.8.195`. `onboot=1`, `startup order=20`. dnsmasq entry
|
||
`zimaos.hubris.network → 192.168.8.195` initially added direct-to-VM on [authentik (124)](../containers/106-auth-outpost.md)
|
||
(later repointed — see above). `/mnt/library` is **not** yet shared into the VM; Phase 2 (virtiofs) is gated on UI evaluation.'
|
||
- slug: vms/108-haos
|
||
title: 108 — `haos-16.3`
|
||
content: "# 108 — `haos-16.3`\n\nHome Assistant OS — the only VM on hubris (HAOS doesn't run cleanly in an LXC, hence the\
|
||
\ qm tenant).\n\n## At a glance\n- **Type:** QEMU VM\n- **HAOS version:** 16.3 (last verified)\n- **IP:** `192.168.8.101`\n\
|
||
- **Resources:** 4 GiB RAM, 32 GiB boot disk\n- **Public hostname:** [`home.hubris.network`](../infrastructure/dns.md)\
|
||
\ → [caddy (121)](../containers/121-caddy.md) → `192.168.8.101:8123`\n\n## Auth\n\nNative OIDC via the HACS integration\
|
||
\ `christiaangoossens/hass-oidc-auth` (repo `https://github.com/christiaangoossens/hass-oidc-auth`).\n\nKey gotchas:\n\
|
||
- HAOS containers don't honor the Network-panel DNS. Set Supervisor DNS via:\n ```\n ha dns options --servers \"dns://192.168.8.180\"\
|
||
\ --servers \"dns://1.1.1.1\"\n ```\n so OIDC discovery resolves internally to [authentik (124)](../containers/106-auth-outpost.md).\n\
|
||
- Authentik app slug in the discovery URL is whatever was set in Authentik — confirm via the DB rather than guessing.\
|
||
\ User set `home-assistant` (with hyphen).\n- YAML config:\n - `features.automatic_user_linking: true` — link to existing\
|
||
\ HA users by `preferred_username` match (otherwise a duplicate is created).\n - `features.default_redirect: true` —\
|
||
\ skip the welcome-splash so users land on the normal HA login page.\n- Run `ha core restart` after each config change\
|
||
\ — HA caches DNS and OIDC discovery across frontend reloads.\n\n## Telemetry\n\nHA pulls Proxmox metrics via the official\
|
||
\ Proxmox VE integration. As of 2026-04-21 [claudio-monitor](../infrastructure/monitoring.md) stopped publishing to MQTT/REST\
|
||
\ (commit `82f0596`) — HA gets metrics from PVE directly; claudio-monitor focuses on alerting.\n\n## Related\n- [Authentik\
|
||
\ (124)](../containers/106-auth-outpost.md)\n- [Caddy (121)](../containers/121-caddy.md)\n- [DNS](../infrastructure/dns.md)\n\
|
||
- [Monitoring](../infrastructure/monitoring.md)\n\n## Changelog\n\n### 2026-04-28 — wiki entry created\nInitial documentation.\n\
|
||
\n### 2026-04-21 — wired into Authentik via HACS hass-oidc-auth\nSupervisor DNS pointed at LXC 124 dnsmasq; YAML features\
|
||
\ for auto-linking + default redirect.\n"
|
||
entity_slug: vm:haos
|
||
tags:
|
||
- vm
|
||
at_glance:
|
||
type: QEMU VM
|
||
haos_version: '16.3'
|
||
ip: 192.168.8.101
|
||
resources: 4 GiB RAM, 32 GiB boot disk
|
||
public_host: home.hubris.network → caddy (121) → 192.168.8.101:8123
|
||
changelog:
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-21'
|
||
title: wired into Authentik via HACS hass-oidc-auth
|
||
body: Supervisor DNS pointed at LXC 124 dnsmasq; YAML features for auto-linking + default redirect.
|
||
- slug: infrastructure/auto-deploy
|
||
title: Auto-deploy — gitea-webhook pipelines
|
||
content: "# Auto-deploy — gitea-webhook pipelines\n\nSeveral configs and apps in the lab live in `dtoro/*` repos on [gitea\
|
||
\ (104)](../containers/104-gitea.md) and auto-redeploy on push. All pipelines follow one of two shapes.\n\n## Two shapes\n\
|
||
\n### Shape A — checkout IS the working tree (config repos)\n\n`/etc/<thing>` or `/var/lib/<thing>/...` is itself a `git\
|
||
\ clone`. Push triggers `git pull` + a reload command. Used for pure-config repos where re-cloning is cheap.\n\n### Shape\
|
||
\ B — receiver outside the app repo (compose stacks)\n\nThe app repo at `/opt/<thing>` is the working tree, but the deploy\
|
||
\ tooling (`webhook.py`, `deploy.sh`, systemd unit) lives in a sibling `/opt/<thing>-deploy/` so the app repo stays portable.\
|
||
\ Push triggers `git pull` + `docker compose up -d --build`. Returns 202 immediately and runs the build in a daemon thread\
|
||
\ because docker builds exceed gitea's request timeout.\n\n## Common\n\n- All receivers validate `X-Gitea-Signature` HMAC-SHA256\
|
||
\ against a per-pipeline secret in `/etc/<thing>-deploy/secret`.\n- All filter to `refs/heads/main` (or `master` for older\
|
||
\ repos). Gitea's \"test delivery\" button sends `ref=main` (without `refs/heads/`) — those will log \"ignoring ref main\"\
|
||
\ and 204. Real pushes work. **Don't \"fix\" the ref filter to accept both** — it'd also accept PR merges from side branches\
|
||
\ that got fast-forwarded.\n- Gitea's `app.ini` `[webhook] ALLOWED_HOST_LIST` must include every receiver IP. Currently:\n\
|
||
\ - `127.0.0.1` (gitea customizations on [LXC 104](../containers/104-gitea.md))\n - `192.168.8.175` ([caddy (121)](../containers/121-caddy.md))\n\
|
||
\ - `192.168.8.205` ([apps (105)](../containers/105-apps.md) — Artifacto)\n - ~~`192.168.8.230` (claudio-bot — destroyed\
|
||
\ 2026-06-04)~~\n - `192.168.8.136` ([mule-images (120)](../containers/120-mule-images.md))\n - `192.168.8.77` ([hubris\
|
||
\ host](../hosts/hubris.md) — backup-library)\n - ~~`192.168.8.190` ([plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology))~~\
|
||
\ (destroyed 2026-06-28)\n - `192.168.8.211` ([trmnl (128)](../containers/128-trmnl.md) — terminalito)\n\n **Don't strip\
|
||
\ these when editing app.ini.**\n\n- Git creds for root-run deploy services live in `/etc/<thing>-deploy/git-credentials`\
|
||
\ (mode 600) and are wired via `credential.helper = store --file=/etc/<thing>-deploy/git-credentials` in the repo's `.git/config`.\
|
||
\ Necessary because the unit typically runs with `ProtectHome=true`, which blocks `/root`.\n\n## Pipelines\n\n| Repo \
|
||
\ | Target | Shape | Receiver \
|
||
\ | Webhook id | Reload action |\n| ------------------------------- | --------------------------------------------\
|
||
\ | ----- | ------------------------------------- | ---------- | ------------- |\n| `dtoro/caddy-conf` |\
|
||
\ [caddy (121)](../containers/121-caddy.md) `/etc/caddy/` | A | `http://192.168.8.175:9797/deploy` | 2 | `caddy validate`\
|
||
\ + `systemctl reload caddy` |\n| `dtoro/gitea-customizations` | [gitea (104)](../containers/104-gitea.md) `/var/lib/gitea/custom/`\
|
||
\ | A | `http://127.0.0.1:9797/deploy` (loopback) | (orig) | `systemctl restart gitea` if templates changed |\n| `dtoro/mule-image`\
|
||
\ | [mule-images (120)](../containers/120-mule-images.md) `/opt/mule-image/` | B | `http://192.168.8.136:9797/deploy`\
|
||
\ | 6 | `docker compose up -d --build` |\n| `dtoro/Artifacto` | [apps (105)](../containers/105-apps.md)\
|
||
\ `/opt/artifacto/` | B | `http://192.168.8.205:9798/deploy` | 7 | `docker compose up -d --build` |\n| ~~`dtoro/Plato`~~\
|
||
\ | ~~[plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology) `/opt/plato/app/`~~\
|
||
\ (destroyed 2026-06-28) | ⊘ | `http://192.168.8.190:9799/deploy` (dead) | 8 (removed) | Repo archived — LXC destroyed\
|
||
\ |\n| `dtoro/claudio-bot` | ~~[claudio-bot (123)](../containers/archive/123-claudio-bot.md)~~ (destroyed\
|
||
\ 2026-06-04) | ⊘ | `http://192.168.8.230:9797/deploy` (dead) | (archived) | Repo archived — LXC destroyed |\n| `dtoro/backup-library`\
|
||
\ | [hubris host](../hosts/hubris.md) `/opt/backup-library/` | A | `http://192.168.8.77:9798/deploy` | (orig)\
|
||
\ | runs `deploy.sh` (preserves admin-edited `/etc/restic/include-*.list`) |\n| `dtoro/oikos` → homelab-mcp | [apps\
|
||
\ (105)](../containers/105-apps.md) `/opt/homelab-mcp/` | B | `http://192.168.8.205:9811/deploy` | 10 (deprecated) | ~~reinstalls\
|
||
\ `homelab-mcp.service` + restart~~ → replaced by Go Docker stack on mac-mini |\n| `dtoro/oikos` → secrets-issuance\
|
||
\ | [apps (105)](../containers/105-apps.md) `/opt/secrets-issuance/` | B | `http://192.168.8.205:9821/deploy` | 11 (deprecated)\
|
||
\ | ~~reinstalls `secrets-issuance.service` + restart~~ → replaced by `internal/secrets/` Go package |\n| `dtoro/terminalito`\
|
||
\ | [trmnl (128)](../containers/128-trmnl.md) `/opt/terminalito/` | B | `http://192.168.8.211:9797/deploy`\
|
||
\ | 12 | reinstalls units + `systemctl restart trmnl-plugins` |\n| `dtoro/oikos` → oikos-console | [apps (105)](../containers/105-apps.md)\
|
||
\ `/opt/oikos-console/` | B | `http://192.168.8.205:9831/deploy` | 14 | reinstalls `oikos-console.service` + restart —\
|
||
\ see [oikos/console/deploy/README.md](../../../oikos/console/deploy/README.md) |\n\n> Note: `dtoro/oikos` has\
|
||
\ **three webhooks** firing on the same push.\n> Each owns its own clone on LXC 105. They don't conflict because each\n\
|
||
> deploy.sh only touches its own service unit + venv.\n\n> **Not yet wired:** `dtoro/claudio-monitor` (push, then `/opt/claudio-monitor/scripts/deploy.sh`\
|
||
\ manually). The former authentik LXC (124) is destroyed — Authentik runs on the [VPS](../../../hosts/netbird-vps.yaml).\
|
||
\ DNS moved to [Technitium on dns (107)](../containers/107-dns.md).\n\n## When you change a tracked config\n\nAlways commit\
|
||
\ + push. Local-only edits drift. Common ones:\n\n- `/etc/caddy/Caddyfile` ↔ `dtoro/caddy-conf` (auto-deploys)\n- `/var/lib/gitea/custom/`\
|
||
\ ↔ `dtoro/gitea-customizations` (auto-deploys)\n- `/opt/artifacto/` ↔ `dtoro/Artifacto` (auto-deploys)\n- `/opt/mule-image/`\
|
||
\ ↔ `dtoro/mule-image` (auto-deploys)\n- ~~`/opt/plato/app/` ↔ `dtoro/Plato`~~ (destroyed 2026-06-28)\n- ~~`/opt/claudio-bot/`\
|
||
\ ↔ `dtoro/claudio-bot`~~ (destroyed 2026-06-04)\n- `/opt/backup-library/` ↔ `dtoro/backup-library` (auto-deploys)\n-\
|
||
\ `/opt/homelab-mcp/` + `/opt/secrets-issuance/` ↔ `dtoro/oikos` (auto-deploys both, see [homelab-context](homelab-context.md))\n\
|
||
\n## Per-pipeline notes / gotchas\n\n### caddy-conf\n- Repo includes `scripts/webhook/install.sh`. Editing the systemd\
|
||
\ unit *inside the repo* does **not** auto-reinstall — re-run `install.sh` manually after unit edits.\n- The unit has\
|
||
\ `ReadWritePaths=/etc/caddy` — load-bearing (`ProtectSystem=full` would otherwise block `git pull`).\n\n### gitea-customizations\n\
|
||
- Receiver is on **loopback** (`127.0.0.1:9797`), not the LXC IP.\n- Online3DViewer binary assets are NOT tracked; `deploy.sh`\
|
||
\ fetches them on first run.\n\n### mule-image / Artifacto\n- Async deploy (returns 202) — gitea would otherwise time\
|
||
\ out the request. Logs: `pct exec <id> -- journalctl -u <thing>-deploy-webhook -f`.\n- **Cloning from inside the LXC\
|
||
\ must use the internal gitea IP** (`http://192.168.8.121:3000/...`). `https://git.hubris.network` hits a connection reset\
|
||
\ from inside [apps (105)](../containers/105-apps.md) (Caddy routing / TLS hairpin not configured for this LXC). Configured\
|
||
\ `origin` on the in-LXC checkout is the internal URL.\n- Manual deploy: `pct exec <id> -- /opt/<thing>-deploy/deploy.sh`.\n\
|
||
- Health: `pct exec <id> -- curl -s http://127.0.0.1:<port>/health` → `ok`.\n- Setup tokens used to register the webhook\
|
||
\ (e.g., `artifacto-deploy-setup`, `artifacto-deploy-setup-2`, `artifacto-cleanup` on user `dtoro`) need manual revocation\
|
||
\ in the Gitea UI → Settings → Applications → Manage Access Tokens. Gitea's `/users/{u}/tokens` endpoints require basic\
|
||
\ auth (not bearer), so cleanup couldn't be automated.\n\n### backup-library\n- Currently the only deploy that targets\
|
||
\ the host directly (`192.168.8.77:9798`).\n- `deploy.sh` is careful to preserve admin edits to `/etc/restic/include-*.list`\
|
||
\ — canonical source is `config/` in the repo, but the install path is treated as authoritative once `deploy.sh` has run.\n\
|
||
\n### homelab-mcp / secrets-issuance\n- Both ride a single push to `dtoro/oikos`. Two clones on LXC 105\n (`/opt/homelab-mcp`,\
|
||
\ `/opt/secrets-issuance`) — each is an independent\n Shape-B target with its own webhook receiver.\n- The deploy script\
|
||
\ restarts the service it just updated. Because the\n webhook receiver itself is a separate systemd unit (`*-deploy.service`),\n\
|
||
\ it does NOT restart itself — but `deploy.sh` running `systemctl\n restart homelab-mcp-deploy.service` (or the secrets-issuance\
|
||
\ one)\n would create a kill-self loop. The current `deploy.sh` is careful\n to only restart the main service.\n- Both\
|
||
\ services consume `/opt/homelab-context` for their runtime data\n (inventory, secret recipient lookup). That clone is\
|
||
\ **the same clone\n every other client has** — kept fresh by `homelab-context-sync.timer`,\n not by these webhooks.\n\
|
||
\n## Custom-built binaries that overlap apt-managed paths\n\nIf a pipeline (or any out-of-band build) drops a binary into\
|
||
\ a path that an apt package also owns — most commonly `/usr/bin/<name>` — then the next `apt upgrade` of the corresponding\
|
||
\ package will silently clobber the custom build. That's exactly how [LXC 121 caddy](../containers/121-caddy.md) went\
|
||
\ down for ~10 min on 2026-05-21: an xcaddy build with `caddy-dns/ionos` lived at `/usr/bin/caddy` and Debian's caddy\
|
||
\ 2.11.2→2.11.3 apt upgrade replaced it with a vanilla 2.11.3 that couldn't parse the Caddyfile.\n\nTwo patterns are acceptable,\
|
||
\ pick one when authoring a pipeline that ships a non-apt binary:\n\n1. **Ship the build as a `.deb` with an epoch-bumped\
|
||
\ version.** Use `dpkg-deb --build` (or `nfpm`) to package the binary as `Package: <name>`, `Version: 1:<upstream>-hubris<n>`.\
|
||
\ The epoch (`1:`) means it beats any non-epoch upstream version regardless of point bumps, so `apt upgrade` is a no-op\
|
||
\ for that package. Used by caddy: `caddy 1:2.11.3-hubris1` (see commit `2026-05-21` in [121-caddy.md](../containers/121-caddy.md)).\n\
|
||
\n2. **Hold the apt package.** `apt-mark hold <pkg>` on the LXC during pipeline install; apt will refuse to upgrade it.\
|
||
\ Simpler than `.deb` packaging but: (a) the hold flag isn't preserved by `dpkg -i` of a new version, (b) it's invisible\
|
||
\ unless you check `apt-mark showhold`, (c) you have to remember to `apt-mark unhold` when you intentionally want a new\
|
||
\ version. The new `homelab apt-audit` subcommand surfaces holds across the fleet so they don't get forgotten.\n\nIf you're\
|
||
\ not sure what's already lurking, run `homelab apt-audit --fleet` and look at the `NONAPT` column — that's a count of\
|
||
\ binaries in `/usr/bin/{caddy,docker,jellyfin}` + `/usr/local/bin/*` that no apt package owns.\n\n## Related\n- [Gitea\
|
||
\ (104)](../containers/104-gitea.md) — webhook source for all of these\n- [Caddy (121)](../containers/121-caddy.md), [apps\
|
||
\ (105)](../containers/105-apps.md), [mule-images (120)](../containers/120-mule-images.md), [hubris host](../hosts/hubris.md)\
|
||
\ — webhook targets\n- [Backups (disabled)](backups.md)\n- [Operations cheatsheet](../../../.agents/operations/commands.md)\
|
||
\ — `homelab apt-audit` / `homelab apt-upgrade` reference\n\n## Changelog\n\n### 2026-06-28 — Plato pipeline decommissioned\n\
|
||
LXC 126 destroyed, webhook id 8 on `dtoro/Plato` removed. `192.168.8.190` removed from gitea `app.ini` `ALLOWED_HOST_LIST`.\n\
|
||
\n### 2026-06-24 — terminalito pipeline added\nWebhook id 12 on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy`\
|
||
\ on [trmnl (128)](../containers/128-trmnl.md). Shape B (`server/deploy/webhook.py` receiver, in-repo `server/deploy/deploy.sh`;\
|
||
\ secret `/etc/terminalito-deploy/secret`). `app.ini` `ALLOWED_HOST_LIST` extended with `192.168.8.211`. Verified end-to-end\
|
||
\ with a push. Repo-local `credential.helper` in `/opt/terminalito/.git/config` (the unit can't read root's global git\
|
||
\ config).\n\n### 2026-05-20 — homelab-mcp + secrets-issuance pipelines added\nWebhook ids 10 + 11 on `dtoro/oikos`\
|
||
\ (ports `9811` + `9821` on [apps (105)](../containers/105-apps.md)). Two webhooks on one repo — each owns its own clone\
|
||
\ (`/opt/homelab-mcp`, `/opt/secrets-issuance`) and only restarts its own service. See [homelab-context](homelab-context.md)\
|
||
\ for why both services live in one repo.\n\n### 2026-05-13 — Plato pipeline added\nWebhook id 8 on `dtoro/Plato` (port\
|
||
\ `9799` on [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology)). `app.ini` `ALLOWED_HOST_LIST`\
|
||
\ extended to include `192.168.8.190`.\n\n### 2026-04-28 — wiki entry created\nInitial documentation. Six active pipelines.\n\
|
||
\n### 2026-04-22 — Artifacto pipeline added\nWebhook id 7 on `dtoro/Artifacto` (port 9798 on apps). `app.ini` `ALLOWED_HOST_LIST`\
|
||
\ extended.\n\n### 2026-06-04 — claudio-bot pipeline decommissioned\nLXC 123 destroyed, `dtoro/claudio-bot` archived.\
|
||
\ Webhook port 9797 dead.\n\n### 2026-04-21 — mule-image + claudio-bot pipelines added\nWebhook id 6; receiver on apps'\
|
||
\ sibling `/opt/mule-deploy/`. Same shape used for claudio-bot.\n\n### 2026-04-20 — caddy-conf + gitea-customizations\
|
||
\ + backup-library pipelines shipped\nInitial three. Set the conventions everything else follows.\n"
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-06-28'
|
||
title: Plato pipeline decommissioned
|
||
body: LXC 126 destroyed, webhook id 8 on `dtoro/Plato` removed. `192.168.8.190` removed from gitea `app.ini` `ALLOWED_HOST_LIST`.
|
||
- date: '2026-06-24'
|
||
title: terminalito pipeline added
|
||
body: Webhook id 12 on `dtoro/terminalito` → `http://192.168.8.211:9797/deploy` on [trmnl (128)](../containers/128-trmnl.md).
|
||
Shape B (`server/deploy/webhook.py` receiver, in-repo `server/deploy/deploy.sh`; secret `/etc/terminalito-deploy/secret`).
|
||
`app.ini` `ALLOWED_HOST_LIST` extended with `192.168.8.211`. Verified end-to-end with a push. Repo-local `credential.helper`
|
||
in `/opt/terminalito/.git/config` (the unit can't read root's global git config).
|
||
- date: '2026-05-20'
|
||
title: homelab-mcp + secrets-issuance pipelines added
|
||
body: Webhook ids 10 + 11 on `dtoro/oikos` (ports `9811` + `9821` on [apps (105)](../containers/105-apps.md)).
|
||
Two webhooks on one repo — each owns its own clone (`/opt/homelab-mcp`, `/opt/secrets-issuance`) and only restarts its
|
||
own service. See [homelab-context](homelab-context.md) for why both services live in one repo.
|
||
- date: '2026-05-13'
|
||
title: Plato pipeline added
|
||
body: Webhook id 8 on `dtoro/Plato` (port `9799` on [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology)).
|
||
`app.ini` `ALLOWED_HOST_LIST` extended to include `192.168.8.190`.
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation. Six active pipelines.
|
||
- date: '2026-04-22'
|
||
title: Artifacto pipeline added
|
||
body: Webhook id 7 on `dtoro/Artifacto` (port 9798 on apps). `app.ini` `ALLOWED_HOST_LIST` extended.
|
||
- date: '2026-06-04'
|
||
title: claudio-bot pipeline decommissioned
|
||
body: LXC 123 destroyed, `dtoro/claudio-bot` archived. Webhook port 9797 dead.
|
||
- date: '2026-04-21'
|
||
title: mule-image + claudio-bot pipelines added
|
||
body: Webhook id 6; receiver on apps' sibling `/opt/mule-deploy/`. Same shape used for claudio-bot.
|
||
- date: '2026-04-20'
|
||
title: caddy-conf + gitea-customizations + backup-library pipelines shipped
|
||
body: Initial three. Set the conventions everything else follows.
|
||
- slug: infrastructure/backups
|
||
title: Backups — restic on external drive (DEPRECATED — superseded)
|
||
content: "# Backups — restic on external drive (DEPRECATED — superseded)\n\n> **DEPRECATED 2026-07-01.** Superseded by the\
|
||
\ **rclone → Proton Drive** off-host mirror on\n> [LXC 132 `rclone`](../containers/132-rclone.md). That job finally closes\
|
||
\ the off-host / 3-2-1 gap\n> this page flagged for months. The restic-on-USB job below is kept for archaeology; it has\
|
||
\ been\n> **DISABLED since 2026-04-22** and is not coming back in its old form.\n\n## Current backup — rclone → Proton\
|
||
\ Drive (LXC 132)\n\n- **Where:** [LXC 132 `rclone`](../containers/132-rclone.md) (`192.168.8.214`), `/mnt/library`\n\
|
||
\ mounted **read-only**.\n- **What:** plain `rclone sync` (Proton mirrors local; browsable files, no versioning) of the\n\
|
||
\ folders listed in `/etc/rclone-backup/folders.list`, to `proton:library-backup/…`.\n- **When:** monthly — `rclone-backup.timer`\
|
||
\ (`OnCalendar=*-*-01 03:00`).\n- **UI:** rclone Web GUI on `192.168.8.214:5572` (LAN-only, no auth).\n- **Encryption:**\
|
||
\ Proton's built-in E2E (no rclone `crypt` overlay).\n- **Runs / logs:** `/var/log/rclone-backup/` + `runs.jsonl`.\n-\
|
||
\ **Still a single off-host target** (Proton only). Not yet a full 3-2-1 (no second independent\n copy), but strictly\
|
||
\ better than the previous \"no off-host copy at all.\"\n\nSee [132-rclone](../containers/132-rclone.md) for the full\
|
||
\ design.\n\n---\n\n## Legacy — restic on external drive (DISABLED 2026-04-22)\n\nChunked monthly restic backup of `/mnt/library`'s\
|
||
\ irreplaceable subset. **Disabled 2026-04-22** as part of the [hubris crash-loop A/B test](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\n\
|
||
\n## Status\n\n**DISABLED 2026-04-22.** All four timers `systemctl disable --now`'d:\n- `backup-library@homecloud.timer`\n\
|
||
- `backup-library@images.timer`\n- `backup-library@small.timer`\n- `backup-library-check.timer`\n\nFstab entry commented\
|
||
\ out. USB drive de-authorized and physically removed. `backup-library-deploy.service` left enabled (harmless webhook\
|
||
\ receiver).\n\n**Reason:** the host hang recurred 2026-04-22 18:42 after 30h despite the `cpu-epp` fix, the UAS blacklist,\
|
||
\ and mount-on-demand. User wants to confirm host stability without the drive at all (was stable 33 days before the drive\
|
||
\ arrived). See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\n\n**To re-enable:**\
|
||
\ uncomment fstab line, `systemctl enable --now` the four timers, re-attach drive.\n\n## Design\n\nMonthly rolling snapshots\
|
||
\ onto a 2 TB external USB drive. Chunked across the month so no single run pushes the Samsung 990 EVO Plus 4 TB (which\
|
||
\ backs `/mnt/library`) into thermal danger.\n\nRetention per tag: `--keep-last 3 --keep-monthly 12 --keep-yearly 3` with\
|
||
\ `--group-by host,tags,paths`.\n\n## Components\n\n- **Repo:** `dtoro/backup-library`\n- **Checkout:** `/opt/backup-library`\
|
||
\ on the [hubris host](../hosts/hubris.md)\n- **Auto-deploys** via gitea webhook → `http://192.168.8.77:9798/deploy`.\
|
||
\ See [auto-deploy](auto-deploy.md). [Gitea (104)](../containers/104-gitea.md) `app.ini` `ALLOWED_HOST_LIST` includes\
|
||
\ `192.168.8.77` for this.\n- **Restic repo:** `/mnt/backup/restic-library`. Passphrase `/etc/restic/passphrase` (mode\
|
||
\ 600). **Escrow in password manager — loss = permanent data loss.**\n- **External drive:** `/dev/sda1` ext4 label `backup-library`\
|
||
\ UUID `ff46e775-1ba1-4892-82c9-e5cac5be933a`. Fstab uses `noauto` + `nofail,x-systemd.device-timeout=10s,errors=remount-ro`.\n\
|
||
\n## Mount-on-demand\n\n`/usr/local/sbin/backup-usb.sh attach|detach|status`. All three backup units (`backup-library.service`,\
|
||
\ `backup-library@.service`, `backup-library-check.service`) have `ExecStartPre=backup-usb.sh attach` and `ExecStopPost=backup-usb.sh\
|
||
\ detach`. The helper toggles `/sys/bus/usb/devices/*/authorized` by matching vendor:product `090c:2320`, then mounts/unmounts\
|
||
\ `/mnt/backup`. Drive is de-authorized when not backing up — no UAS keepalive, no kernel error-recovery paths firing\
|
||
\ against a flaky bridge.\n\n## UAS blacklist\n\n`/etc/modprobe.d/usb-storage-quirks.conf`:\n```\noptions usb-storage\
|
||
\ quirks=090c:2320:u\n```\nForces Bulk-Only Transport (BOT) instead of UAS for the SMI bridge. Confirm with `dmesg | grep\
|
||
\ \"UAS is ignored\"`.\n\n## Schedule\n\nThree timers, one per chunk, staggered ~10 days apart so each disk zone gets\
|
||
\ a long cooldown:\n\n| Timer | When | Include list | Approx\
|
||
\ size |\n| ---------------------------------- | -------------- | ------------------------------------ | ----------- |\n\
|
||
| `backup-library@homecloud.timer` | day 1 / month | `/etc/restic/include-homecloud.list` | ~315 G |\n| `backup-library@images.timer`\
|
||
\ | day 10 / month | `/etc/restic/include-images.list` | ~103 G |\n| `backup-library@small.timer` \
|
||
\ | day 20 / month | `/etc/restic/include-small.list` (docs / books / music / notes / repos / marimo / heaper) | ~11 G\
|
||
\ |\n\nSnapshots tagged `chunk-<group>` so forget/prune treats each series independently.\n\nAd-hoc full run (kept for\
|
||
\ manual use): `systemctl start backup-library.service` (no arg → uses `/etc/restic/include.list`, tag `monthly`).\n\n\
|
||
Yearly integrity: `backup-library-check.timer` (`OnCalendar=yearly`) runs full `restic check --read-data`.\n\n## Thermal\
|
||
\ caps\n\nBaked into the systemd units:\n- `IOReadBandwidthMax=/mnt/library 50M`\n- `IOWriteBandwidthMax=/mnt/backup 30M`\n\
|
||
- `--read-concurrency=1` on restic.\n\n## Wrapper\n\n`/usr/local/sbin/backup-library.sh` — preflight → unlock → backup\
|
||
\ → forget/prune (`--group-by host,tags,paths`) → `check --read-data-subset=5%` → notify. Takes optional `<group>` arg\
|
||
\ or `GROUP=` env.\n\n## Notifications\n\n~~POST to claudio-bot (123) `http://192.168.8.230:9090/notify`~~ — IPC endpoint\
|
||
\ dead since 2026-06-04. When backups are re-enabled, wire notifications to Hermes `send_message` via Matrix instead.\n\
|
||
\n`OnFailure=notify-failure@%n.service` on the backup unit fires a synchronous notify as belt-and-suspenders for cases\
|
||
\ where the wrapper itself died before reaching its own notify.\n\n## Recovery\n\nRunbook at `/usr/share/doc/backup-library/RECOVERY.md`\
|
||
\ (or in the repo at `doc/RECOVERY.md`). Covers `restic snapshots/ls/find/restore/mount`, uid/gid gotcha, cross-host recovery.\n\
|
||
\n## Known SPOF\n\nSingle drive. RECOVERY.md flags the 3-2-1 gap. Mitigations (second drive, cloud repo via `restic copy`)\
|
||
\ were not implemented before this job was retired — the **off-host copy is now provided by [rclone → Proton Drive (LXC\
|
||
\ 132)](../containers/132-rclone.md)** instead. A second independent copy is still outstanding.\n\n## Drive history\n\n\
|
||
The `Silicon Motion Portable SSD` (vid:pid `090c:2320`) drops under sustained heavy writes through a hub chain. Bypass\
|
||
\ all hubs / use a rear motherboard USB 3 port if attaching it again.\n\nAfter it was first attached on 2026-04-19, hubris\
|
||
\ crashed twice in 2.5 days (46h then 12h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic,\
|
||
\ OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1.\
|
||
\ The UAS blacklist + mount-on-demand mitigations didn't fully eliminate it (recurrence 2026-04-22), prompting drive removal\
|
||
\ as the cleaner test. See [investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\n\n\
|
||
## Thermal monitoring\n\nMoved out of this repo to `dtoro/claudio-monitor` on 2026-04-21 (commit `50dc213`). See [monitoring](monitoring.md).\n\
|
||
\n## Related\n- [Hubris host](../hosts/hubris.md)\n- ~~[claudio-bot (123)](../containers/archive/123-claudio-bot.md)~~\
|
||
\ (destroyed 2026-06-04)\n- [Monitoring](monitoring.md)\n- [Auto-deploy](auto-deploy.md)\n- [Investigation: 2026-04-21\
|
||
\ crash loop](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md)\n\n## Changelog\n\n### 2026-07-01\
|
||
\ — DEPRECATED; superseded by rclone → Proton Drive (LXC 132)\nOff-host backup moved to a plain `rclone sync` mirror on\
|
||
\ the new [LXC 132 `rclone`](../containers/132-rclone.md) (`/mnt/library` → Proton Drive, monthly, LAN Web GUI). This\
|
||
\ finally provides the off-host copy the \"Known SPOF\" note wanted. The restic-on-USB units on hubris remain `disabled`\
|
||
\ (drive already removed 2026-04-22); page restructured to lead with the current job and demote restic to \"Legacy\".\n\
|
||
\n### 2026-04-28 — wiki entry created\nInitial documentation. Status remains DISABLED.\n\n### 2026-04-22 — DISABLED\n\
|
||
Drive removed as the A/B test in the [crash investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).\
|
||
\ Timers disabled, fstab commented, drive de-authorized.\n\n### 2026-04-21 — UAS blacklist + mount-on-demand shipped;\
|
||
\ root-caused host hangs to drive\nDrive identified as the source of the hangs after hubris crashed twice in 2.5 days.\
|
||
\ UAS blacklist forces BOT; helper script toggles `/sys/bus/usb/.../authorized` so the drive is de-authorized when not\
|
||
\ backing up. Recovery drill (restore 188KB PDF + hash compare) had passed earlier. Bug fixed in `backup-library.sh`:\
|
||
\ `python3 -c '…' KEY=VAL` does NOT pass env vars — env-var prefix must precede the command. Caused false-failure even\
|
||
\ after successful backups.\n\n### 2026-04-20 — deployed; redesigned for thermal-gentleness\nInitial deploy. First backup\
|
||
\ attempt died at 18:56 (USB drive dropped off the bus during heavy writes); after re-plugging, restic resumed and completed\
|
||
\ at 21:21.\n"
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-07-01'
|
||
title: DEPRECATED; superseded by rclone → Proton Drive (LXC 132)
|
||
body: Off-host backup moved to a plain `rclone sync` mirror on the new [LXC 132 `rclone`](../containers/132-rclone.md)
|
||
(`/mnt/library` → Proton Drive, monthly, LAN Web GUI). This finally provides the off-host copy the "Known SPOF" note
|
||
wanted. The restic-on-USB units on hubris remain `disabled` (drive already removed 2026-04-22); page restructured to
|
||
lead with the current job and demote restic to "Legacy".
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation. Status remains DISABLED.
|
||
- date: '2026-04-22'
|
||
title: DISABLED
|
||
body: Drive removed as the A/B test in the [crash investigation](../../sources/investigations/archive/2026-04-21-hubris-crash-loop.md).
|
||
Timers disabled, fstab commented, drive de-authorized.
|
||
- date: '2026-04-21'
|
||
title: UAS blacklist + mount-on-demand shipped; root-caused host hangs to drive
|
||
body: 'Drive identified as the source of the hangs after hubris crashed twice in 2.5 days. UAS blacklist forces BOT; helper
|
||
script toggles `/sys/bus/usb/.../authorized` so the drive is de-authorized when not backing up. Recovery drill (restore
|
||
188KB PDF + hash compare) had passed earlier. Bug fixed in `backup-library.sh`: `python3 -c ''…'' KEY=VAL` does NOT
|
||
pass env vars — env-var prefix must precede the command. Caused false-failure even after successful backups.'
|
||
- date: '2026-04-20'
|
||
title: deployed; redesigned for thermal-gentleness
|
||
body: Initial deploy. First backup attempt died at 18:56 (USB drive dropped off the bus during heavy writes); after re-plugging,
|
||
restic resumed and completed at 21:21.
|
||
- slug: infrastructure/dns
|
||
title: DNS — split-horizon `*.hubris.network`
|
||
content: '# DNS — split-horizon `*.hubris.network`
|
||
|
||
|
||
LAN clients resolve `*.hubris.network` to the [Caddy reverse proxy](../containers/121-caddy.md) (`192.168.8.175`). Public
|
||
clients resolve to the IONOS VPS (`82.165.190.79`) via an IONOS wildcard, where they hit the [VPS traefik public ingress](ingress.md).
|
||
|
||
|
||
There is **no wildcard on the LAN side**. Every subdomain needs an explicit entry.
|
||
|
||
|
||
## Components
|
||
|
||
|
||
- **Authoritative public DNS:** IONOS. `*.hubris.network → 82.165.190.79` (was `74.118.126.4` until 2026-04-22).
|
||
|
||
- **LAN authoritative for `hubris.network` records:** [Technitium DNS](https://technitium.com) on [dns (107)](../containers/107-dns.md)
|
||
at `192.168.8.2:53`. Syncs A records to the NetBird managed DNS zone via cron (see [dns-sync.py](../../../scripts/dns-sync.py)).
|
||
Formerly dnsmasq on [authentik (124)](../containers/106-auth-outpost.md) (decommissioned 2026-06-04).
|
||
|
||
- **PVE host** (`192.168.8.77`): resolver is the local Netbird daemon at `100.122.38.109:53`, which forwards to the LAN/upstream
|
||
and learns hubris.network answers via that path. `netbird status` says "Nameservers: 0/0 Available" — confirming netbird
|
||
does NOT manage a hubris.network zone; it just caches whatever the system resolver returns.
|
||
|
||
- **Some LXCs** keep router DNS (`192.168.8.1`) or Tailscale MagicDNS (`100.100.100.100`), both of which return the public
|
||
IONOS A record. Those LXCs need either a `/etc/hosts` override or local dnsmasq — see [mesh migration](mesh.md) for which
|
||
technique applies where.
|
||
|
||
|
||
## Live entries (as of 2026-06-04)
|
||
|
||
|
||
```
|
||
|
||
address=/auth.hubris.network/82.165.190.79 # → VPS, not Caddy (Authentik migrated 2026-05-31)
|
||
|
||
address=/sso.hubris.network/192.168.8.175 # → Caddy → LAN forward-auth outpost (106); added 2026-06-01
|
||
|
||
address=/git.hubris.network/192.168.8.175
|
||
|
||
address=/media.hubris.network/192.168.8.175
|
||
|
||
address=/paperless.hubris.network/192.168.8.175
|
||
|
||
address=/books.hubris.network/192.168.8.175
|
||
|
||
address=/home.hubris.network/192.168.8.175
|
||
|
||
address=/cloud.hubris.network/192.168.8.175
|
||
|
||
address=/matrix.hubris.network/192.168.8.175
|
||
|
||
address=/proxmox.hubris.network/192.168.8.175
|
||
|
||
address=/docker.hubris.network/192.168.8.175
|
||
|
||
address=/jellyseerr.hubris.network/192.168.8.175
|
||
|
||
address=/qbit.hubris.network/192.168.8.175
|
||
|
||
address=/sab.hubris.network/192.168.8.175
|
||
|
||
address=/blog.hubris.network/192.168.8.175
|
||
|
||
address=/photos.hubris.network/192.168.8.175
|
||
|
||
address=/photos-new.hubris.network/192.168.8.175
|
||
|
||
address=/artifacto.hubris.network/192.168.8.175
|
||
|
||
address=/zimaos.hubris.network/192.168.8.175
|
||
|
||
address=/nfs-export.hubris.network/192.168.8.200
|
||
|
||
```
|
||
|
||
|
||
Note: `nfs-export.hubris.network` is the only `.hubris.network` entry that points to a non-HTTP service (NFSv4 on port
|
||
2049). It bypasses [caddy (121)](../containers/121-caddy.md) because NFS is L4, not HTTP — Caddy has nothing to do.
|
||
|
||
|
||
## Why split-horizon
|
||
|
||
|
||
The IONOS wildcard points at the VPS for public ingress (per-host routers in [VPS traefik](ingress.md)). The VPS only
|
||
routes hostnames it knows — anything else 404s. So LAN clients pointing at the public IP are a dead end for any service
|
||
that isn''t explicitly published. The [Technitium DNS](dns.md) override on `192.168.8.2` keeps LAN traffic on the home
|
||
Caddy.
|
||
|
||
|
||
## The gotcha that cost a debug session (2026-04-22)
|
||
|
||
|
||
Creating a new Caddyfile site block is necessary but **not sufficient**. Without the Technitium entry on [dns (107)](../containers/107-dns.md),
|
||
LAN queries fall through to upstream, get the public IONOS answer, and time out. Symptom: "subdomain doesn''t load" even
|
||
though Caddy config + cert are fine.
|
||
|
||
|
||
## Recipe — adding a new subdomain
|
||
|
||
|
||
1. Edit `/etc/caddy/Caddyfile` on [caddy (121)](../containers/121-caddy.md), commit + push to `dtoro/caddy-conf`. Webhook
|
||
reloads caddy. See [auto-deploy](auto-deploy.md).
|
||
|
||
2. Add the A record in the [Technitium UI](http://192.168.8.2) at `dns (107)` — the NetBird managed DNS zone sync picks
|
||
it up within ~10 minutes via cron. Or add directly to the NetBird managed zone via API if you need it faster.
|
||
|
||
3. Verify: `dig @192.168.8.2 +short <new>.hubris.network` → `192.168.8.175`.
|
||
|
||
4. On macOS clients, flush: `sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder`.
|
||
|
||
|
||
> The Technitium config on LXC 107 is the single source of truth. Never hand-edit the NetBird managed zone directly —
|
||
the [`scripts/dns-sync.py`](../../../scripts/dns-sync.py) cron on 107 reconciles them and reaps stale records. See [dns.md
|
||
changelog 2026-06-03](#2026-06-03--single-authoring-source-technitium--netbird-managed-zone-sync).
|
||
|
||
|
||
## Public path — what does and doesn''t follow the LAN map
|
||
|
||
|
||
- Hostnames published in [VPS traefik dynamic config](ingress.md) (currently `artifacto.hubris.network`, `blog.hubris.network`)
|
||
reach a real backend over the netbird mesh.
|
||
|
||
- Anything else with a `*.hubris.network` URL hits the VPS but isn''t routed anywhere — returns 404.
|
||
|
||
- `netbird.hubris.network` is its own thing — TCP passthrough at the VPS, served by netbird-proxy. Doesn''t follow the
|
||
file-provider router pattern.
|
||
|
||
|
||
## Long-term plan
|
||
|
||
|
||
Either:
|
||
|
||
- Move split-horizon DNS to the LAN router so `*.hubris.network → 192.168.8.175` is answered for every LAN client. Eliminates
|
||
per-LXC overrides.
|
||
|
||
- Or, once the [Tailscale → Netbird migration](mesh.md) completes, every LXC''s resolver becomes the netbird daemon, which
|
||
already learns hubris.network answers via the system resolver chain.
|
||
|
||
|
||
## Related
|
||
|
||
- [Caddy (121)](../containers/121-caddy.md) — every LAN entry points here
|
||
|
||
- [Ingress (VPS traefik)](ingress.md) — public-side counterpart
|
||
|
||
- [Mesh migration](mesh.md) — per-LXC DNS workarounds during the transition
|
||
|
||
- [DNS server (107)](../containers/107-dns.md) — Technitium, current DNS authority
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-28 — `plato.hubris.network` removed
|
||
|
||
Plato (LXC 126) decommissioned. Technitium entry deleted; dns-sync cron reaped the NetBird managed zone record.
|
||
|
||
|
||
### 2026-06-17 — Fritz!Box DNSv4 server set to Technitium; old limitation resolved
|
||
|
||
Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs — the limitation noted below is resolved.
|
||
Configured at Fritz!Box Internet → Filter → DNS Server → DNSv4 Server = `192.168.8.2` (User-defined).
|
||
|
||
Authentik LXC 124 (192.168.8.180) destroyed — Authentik runs on VPS, DNS on Technitium (107).
|
||
|
||
- Caddy: `auth.hubris.network`, `authentik` snippet, and `sso.hubris.network` all proxied to VPS
|
||
|
||
- `header_up Host auth.hubris.network` added to strip `:443` from upstream Host header
|
||
|
||
- All 9 LXCs'' /etc/hosts updated: `auth.hubris.network → 192.168.8.175` (Caddy proxy)
|
||
|
||
- Inventory: removed `hosts.authentik`, renamed `dnsmasq` service → `dns`
|
||
|
||
- Docs: `124-authentik.md` deleted; dns.md references updated to Technitium (107)
|
||
|
||
The "delete NetBird managed zone → forward everything to Technitium" plan was **abandoned** — NetBird''s DNS defeats it:
|
||
it **won''t apply a nameserver group that contains the peer''s own mesh IP** (the Mac''s `100.122.234.17` → `Nameservers:
|
||
0/0 Available`), and nameserver-group forwarding to Technitium never actually took effect for mesh peers (the **managed
|
||
zone was doing all the real work**; disabling it broke all mesh resolution). So the model is now:
|
||
|
||
|
||
- **Technitium (`192.168.8.2`) is the single place you author DNS** (UI/API, MX/SPF/CAA, full zone).
|
||
|
||
- A **sync job on [dns (107)](../containers/107-dns.md)** (`/opt/dns-sync/sync.py`, cron */10) reconciles Technitium''s
|
||
named A-records → the **NetBird managed DNS zone** via the NetBird API (`/api/dns/zones/{id}/records`, PAT in sops `secrets/netbird-pat.yaml`).
|
||
Mesh peers keep using the managed zone (which works); non-mesh LAN clients query Technitium directly; undefined names
|
||
fall to the public IONOS wildcard (`.79`) — correct.
|
||
|
||
- This **killed the manual drift** that caused the whole `auth`/`sso`/`nfs-export` saga. Never hand-edit the NetBird managed
|
||
zone again — edit Technitium; the sync propagates.
|
||
|
||
|
||
**Cleanup done same day:** removed the inert Mac-Mini Technitium secondary (mesh-only, served nobody); reverted the primary''s
|
||
`zoneTransfer=Allow`; fixed `home-lab-dns` group → `[192.168.8.2]` (dropped the self-referencing Mac IP → now `1/1 Available`);
|
||
deleted the vestigial `Proxmox Names` group.
|
||
|
||
|
||
> Reference: [scripts/dns-sync.py](../../../scripts/dns-sync.py). The sync''s source of truth is Technitium; it **deletes**
|
||
NetBird records absent from Technitium (so obsolete names like `files`, `photos-new` get reaped).
|
||
|
||
|
||
### 2026-06-06 — dns-sync cron finally installed (had been dormant since 2026-06-04 deployment)
|
||
|
||
The `dns-sync.py` script on LXC 107 had been placed at `/opt/dns-sync/sync.py` on 2026-06-04 but **no crontab was configured**
|
||
— the sync had never run automatically. The NetBird managed DNS zone was only in sync because manual runs happened during
|
||
incident debugging.
|
||
|
||
|
||
**Fixed:** added `/etc/cron.d/dns-sync` (`*/10 * * * * root python3 /opt/dns-sync/sync.py >> /var/log/dns-sync.log 2>&1`).
|
||
|
||
|
||
Also added a Caddy backend health check cron on hubris (`/etc/cron.d/caddy-backend-health`) that runs `scripts/check-caddy-backends.sh`
|
||
every 10 minutes.
|
||
|
||
|
||
### 2026-06-02 — 8 LXCs moved from DHCP to static IP
|
||
|
||
All LXCs that Caddy reverse-proxies to by IP were on `ip=dhcp` and could float on reboot (arriman got a different lease
|
||
mid-session and broke). Fixed via `pct set` + in-LXC `/etc/network/interfaces`. Affected: 101 jellyfin, 103 paperless,
|
||
104 gitea, 105 apps, 114 nextcloud, 118 elementsynapse, 120 mule-images, 121 caddy, 122 arriman. See [arriman changelog](../containers/122-arriman.md#changelog).
|
||
|
||
|
||
### 2026-06-01 — dnsmasq replaced by Technitium on [dns (107)](../containers/107-dns.md); LXC 124 retired
|
||
|
||
Split-horizon DNS moved off [124](../containers/106-auth-outpost.md) to a dedicated **Technitium** LXC at **`192.168.8.2`**
|
||
(zone: specific A overrides + wildcard→VPS + replicated MX/SPF/CAA). NetBird `home-lab-dns` nameserver group cut over
|
||
to `192.168.8.2` (with `.180` as a now-dead fallback). dnsmasq stopped, all names verified via Technitium, **LXC 124 shut
|
||
down**. **Caveat:** the [NetBird managed DNS zone](../containers/106-auth-outpost.md) still answers most app names *directly*
|
||
(bypassing the nameserver group) — three overlapping DNS sources remain; see the single-source-of-truth decision (Phase
|
||
4). **Action needed:** update router DHCP DNS from the dead `.180` → `192.168.8.2` for any plain-LAN (non-mesh) clients.
|
||
|
||
|
||
### 2026-05-31 — `auth.hubris.network` re-pointed to the VPS (`82.165.190.79`)
|
||
|
||
Authentik migrated off LXC 124 onto the VPS (see [investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)).
|
||
The dnsmasq entry changed from `192.168.8.175` (home Caddy) to `82.165.190.79` (VPS traefik). This is the first LAN entry
|
||
that intentionally points at the VPS rather than Caddy — `auth` is now a genuinely public service served directly from
|
||
the VPS. **Gotcha logged:** the NetBird per-client resolver (`100.122.255.254`) caches dnsmasq answers and does **not**
|
||
clear on `netbird down/up`; clients needed `/etc/hosts` overrides or `resolvectl flush-caches` to pick up the change.
|
||
Since the service is now fully public, the long-term cleaner option is to drop the override entirely and let it fall through
|
||
to the IONOS wildcard (which also points at the VPS).
|
||
|
||
|
||
### 2026-05-14 — `nfs-export.hubris.network` added (direct, non-HTTP)
|
||
|
||
NFSv4 export server [nfs-export (102)](../containers/102-nfs-export.md) at `192.168.8.200`. Direct entry, not Caddy-fronted
|
||
— NFS is L4, no HTTP reverse-proxy meaningful.
|
||
|
||
|
||
### 2026-05-14 — `zimaos.hubris.network` added (Caddy-fronted, standard pattern)
|
||
|
||
New LAN entry for [100-zimaos](../vms/100-zimaos.md) → [caddy (121)](../containers/121-caddy.md) → `192.168.8.195`. Briefly
|
||
pointed direct-to-VM during install for the initial smoke-test, then re-pointed once a Caddyfile block was added (`reverse_proxy
|
||
192.168.8.195` + IONOS DNS-01 TLS).
|
||
|
||
|
||
### 2026-05-13 — `plato.hubris.network` added; `files.hubris.network` removed
|
||
|
||
New LAN-only entry for [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology). Same day, the `files.hubris.network`
|
||
entry for the just-decommissioned seafile experiment was dropped; queries now fall through to the public IONOS answer
|
||
(no LAN backend).
|
||
|
||
|
||
### 2026-05-12 — `files.hubris.network` added (since removed 2026-05-13)
|
||
|
||
Originally added for the seafile (LXC 125) Nextcloud-replacement evaluation. Pointed at 192.168.8.175 (Caddy reverse-proxied
|
||
to 192.168.8.185:80). Entry removed when the experiment was torn down a day later.
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Initial documentation. 16 active entries.
|
||
|
||
|
||
### 2026-04-22 — IONOS wildcard moved 74.118.126.4 → 82.165.190.79
|
||
|
||
Public path now lands on the VPS traefik, not the old yunohost. Necessary for the [public ingress](ingress.md) pattern.
|
||
The LAN dead-end semantics didn''t change — public DNS still doesn''t help LAN clients reach LAN-only services.
|
||
|
||
|
||
### 2026-04-22 — three caddy sites without DNS entries (jellyseerr, qbit, sab)
|
||
|
||
Caddy + certs were working but LAN resolution failed because the dnsmasq lines weren''t added. Lesson recorded; entries
|
||
added later that day.
|
||
|
||
|
||
### 2026-04-21 — dnsmasq stood up on LXC 124
|
||
|
||
Co-located with Authentik. Initial entries cover everything routed through Caddy.
|
||
|
||
'
|
||
entity_slug: service:dns
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-06-28'
|
||
title: '`plato.hubris.network` removed'
|
||
body: Plato (LXC 126) decommissioned. Technitium entry deleted; dns-sync cron reaped the NetBird managed zone record.
|
||
- date: '2026-06-17'
|
||
title: Fritz!Box DNSv4 server set to Technitium; old limitation resolved
|
||
body: 'Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs — the limitation noted below is
|
||
resolved. Configured at Fritz!Box Internet → Filter → DNS Server → DNSv4 Server = `192.168.8.2` (User-defined). Authentik
|
||
LXC 124 (192.168.8.180) destroyed — Authentik runs on VPS, DNS on Technitium (107). - Caddy: `auth.hubris.network`,
|
||
`authentik` snippet, and `sso.hubris.network` all proxied to VPS - `header_up Host auth.hubris.network` added to strip
|
||
`:443` from upstream Host header - All 9 LXCs'' /etc/hosts updated: `auth.hubris.network → 192.168.8.175` (Caddy proxy)
|
||
- Inventory: removed `hosts.authentik`, renamed `dnsmasq` service → `dns` - Docs: `124-authentik.md` deleted; dns.md
|
||
references updated to Technitium (107) The "delete NetBird managed zone → forward everything to Technitium" plan was
|
||
**abandoned** — NetBird''s DNS defeats it: it **won''t apply a nameserver group that contains the peer''s own mesh IP**
|
||
(the Mac''s `100.122.234.17` → `Nameservers: 0/0 Available`), and nameserver-group forwarding to Technitium never actually
|
||
took effect for mesh peers (the **managed zone was doing all the real work**; disabling it broke all mesh resolution).
|
||
So the model is now: - **Technitium (`192.168.8.2`) is the single place you author DNS** (UI/API, MX/SPF/CAA, full zone).
|
||
- A **sync job on [dns (107)](../containers/107-dns.md)** (`/opt/dns-sync/sync.py`, cron */10) reconciles Technitium''s
|
||
named A-records → the **NetBird managed DNS zone** via the NetBird API (`/api/dns/zones/{id}/records`, PAT in sops `secrets/netbird-pat.yaml`).
|
||
Mesh peers keep using the managed zone (which works); non-mesh LAN clients query Technitium directly; undefined names
|
||
fall to the public IONOS wildcard (`.79`) — correct. - This **killed the manual drift** that caused the whole `auth`/`sso`/`nfs-export`
|
||
saga. Never hand-edit the NetBird managed zone again — edit Technitium; the sync propagates. **Cleanup done same day:**
|
||
removed the inert Mac-Mini Technitium secondary (mesh-only, served nobody); reverted the primary''s `zoneTransfer=Allow`;
|
||
fixed `home-lab-dns` group → `[192.168.8.2]` (dropped the self-referencing Mac IP → now `1/1 Available`); deleted the
|
||
vestigial `Proxmox Names` group. > Reference: [scripts/dns-sync.py](../../../scripts/dns-sync.py). The sync''s source
|
||
of truth is Technitium; it **deletes** NetBird records absent from Technitium (so obsolete names like `files`, `photos-new`
|
||
get reaped).'
|
||
- date: '2026-06-06'
|
||
title: dns-sync cron finally installed (had been dormant since 2026-06-04 deployment)
|
||
body: The `dns-sync.py` script on LXC 107 had been placed at `/opt/dns-sync/sync.py` on 2026-06-04 but **no crontab was
|
||
configured** — the sync had never run automatically. The NetBird managed DNS zone was only in sync because manual runs
|
||
happened during incident debugging. **Fixed:** added `/etc/cron.d/dns-sync` (`*/10 * * * * root python3 /opt/dns-sync/sync.py
|
||
>> /var/log/dns-sync.log 2>&1`). Also added a Caddy backend health check cron on hubris (`/etc/cron.d/caddy-backend-health`)
|
||
that runs `scripts/check-caddy-backends.sh` every 10 minutes.
|
||
- date: '2026-06-02'
|
||
title: 8 LXCs moved from DHCP to static IP
|
||
body: 'All LXCs that Caddy reverse-proxies to by IP were on `ip=dhcp` and could float on reboot (arriman got a different
|
||
lease mid-session and broke). Fixed via `pct set` + in-LXC `/etc/network/interfaces`. Affected: 101 jellyfin, 103 paperless,
|
||
104 gitea, 105 apps, 114 nextcloud, 118 elementsynapse, 120 mule-images, 121 caddy, 122 arriman. See [arriman changelog](../containers/122-arriman.md#changelog).'
|
||
- date: '2026-06-01'
|
||
title: dnsmasq replaced by Technitium on [dns (107)](../containers/107-dns.md); LXC 124 retired
|
||
body: 'Split-horizon DNS moved off [124](../containers/106-auth-outpost.md) to a dedicated **Technitium** LXC at **`192.168.8.2`**
|
||
(zone: specific A overrides + wildcard→VPS + replicated MX/SPF/CAA). NetBird `home-lab-dns` nameserver group cut over
|
||
to `192.168.8.2` (with `.180` as a now-dead fallback). dnsmasq stopped, all names verified via Technitium, **LXC 124
|
||
shut down**. **Caveat:** the [NetBird managed DNS zone](../containers/106-auth-outpost.md) still answers most app names
|
||
*directly* (bypassing the nameserver group) — three overlapping DNS sources remain; see the single-source-of-truth decision
|
||
(Phase 4). **Action needed:** update router DHCP DNS from the dead `.180` → `192.168.8.2` for any plain-LAN (non-mesh)
|
||
clients.'
|
||
- date: '2026-05-31'
|
||
title: '`auth.hubris.network` re-pointed to the VPS (`82.165.190.79`)'
|
||
body: Authentik migrated off LXC 124 onto the VPS (see [investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)).
|
||
The dnsmasq entry changed from `192.168.8.175` (home Caddy) to `82.165.190.79` (VPS traefik). This is the first LAN
|
||
entry that intentionally points at the VPS rather than Caddy — `auth` is now a genuinely public service served directly
|
||
from the VPS. **Gotcha logged:** the NetBird per-client resolver (`100.122.255.254`) caches dnsmasq answers and does
|
||
**not** clear on `netbird down/up`; clients needed `/etc/hosts` overrides or `resolvectl flush-caches` to pick up the
|
||
change. Since the service is now fully public, the long-term cleaner option is to drop the override entirely and let
|
||
it fall through to the IONOS wildcard (which also points at the VPS).
|
||
- date: '2026-05-14'
|
||
title: '`nfs-export.hubris.network` added (direct, non-HTTP)'
|
||
body: NFSv4 export server [nfs-export (102)](../containers/102-nfs-export.md) at `192.168.8.200`. Direct entry, not Caddy-fronted
|
||
— NFS is L4, no HTTP reverse-proxy meaningful.
|
||
- date: '2026-05-14'
|
||
title: '`zimaos.hubris.network` added (Caddy-fronted, standard pattern)'
|
||
body: New LAN entry for [100-zimaos](../vms/100-zimaos.md) → [caddy (121)](../containers/121-caddy.md) → `192.168.8.195`.
|
||
Briefly pointed direct-to-VM during install for the initial smoke-test, then re-pointed once a Caddyfile block was added
|
||
(`reverse_proxy 192.168.8.195` + IONOS DNS-01 TLS).
|
||
- date: '2026-05-13'
|
||
title: '`plato.hubris.network` added; `files.hubris.network` removed'
|
||
body: New LAN-only entry for [plato (126)](../containers/index.md#recently-destroyed-kept-for-archaeology). Same day,
|
||
the `files.hubris.network` entry for the just-decommissioned seafile experiment was dropped; queries now fall through
|
||
to the public IONOS answer (no LAN backend).
|
||
- date: '2026-05-12'
|
||
title: '`files.hubris.network` added (since removed 2026-05-13)'
|
||
body: Originally added for the seafile (LXC 125) Nextcloud-replacement evaluation. Pointed at 192.168.8.175 (Caddy reverse-proxied
|
||
to 192.168.8.185:80). Entry removed when the experiment was torn down a day later.
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation. 16 active entries.
|
||
- date: '2026-04-22'
|
||
title: IONOS wildcard moved 74.118.126.4 → 82.165.190.79
|
||
body: Public path now lands on the VPS traefik, not the old yunohost. Necessary for the [public ingress](ingress.md) pattern.
|
||
The LAN dead-end semantics didn't change — public DNS still doesn't help LAN clients reach LAN-only services.
|
||
- date: '2026-04-22'
|
||
title: three caddy sites without DNS entries (jellyseerr, qbit, sab)
|
||
body: Caddy + certs were working but LAN resolution failed because the dnsmasq lines weren't added. Lesson recorded; entries
|
||
added later that day.
|
||
- date: '2026-04-21'
|
||
title: dnsmasq stood up on LXC 124
|
||
body: Co-located with Authentik. Initial entries cover everything routed through Caddy.
|
||
- slug: infrastructure/homelab-context
|
||
title: Homelab context distribution
|
||
content: "# Homelab context distribution\n\nThe cross-client context-and-secrets system that makes every agent (Claude\n\
|
||
Code, Hermes Agent, future MCP-capable clients) on every machine in the lab\nself-locating and able to read the same source\
|
||
\ of truth.\n\nOperational walkthrough for enrolling a new client lives in\n[operations/agent-enrollment.md](../../../.agents/operations/agent-enrollment.md);\
|
||
\ this\npage is the architecture reference.\n\n## What's where\n\n| Piece | Host | Path | Role |\n| --- | --- | --- |\
|
||
\ --- |\n| Source of truth | [gitea (104)](../containers/104-gitea.md) | `dtoro/oikos.git` | Inventory + wiki +\
|
||
\ service code |\n| Per-client clone | every enrolled client | `/opt/homelab-context/` | Read by `homelab` CLI, MCP server,\
|
||
\ Hermes Agent |\n| `homelab` CLI | every enrolled client | `/usr/local/bin/homelab` → `/opt/homelab-context/bin/homelab`\
|
||
\ (symlink) | Operator surface for enroll/secret/ssh/pct |\n| Per-client age key | every enrolled client | `/etc/age/key.txt`\
|
||
\ (0600 root) | Decrypts SOPS-encrypted secrets the client is a recipient on |\n| MCP server | [apps (105)](../containers/105-apps.md)\
|
||
\ | `homelab-mcp.service` on port 9810 (https://mcp.hubris.network/mcp) | 14 tools: 8 context (get_host, search_docs,\
|
||
\ …) + 5 read-only management (get_service_status, tail_log, …) + list_my_secrets |\n| Secrets-issuance | [apps (105)](../containers/105-apps.md)\
|
||
\ | `secrets-issuance.service` on port 9820 (https://secrets.hubris.network/issue) | Generates per-client age keypair\
|
||
\ on first bootstrap; idempotent; admin-token-gated `/revoke` |\n| Sync timer | every enrolled client | `homelab-context-sync.timer`\
|
||
\ (Linux) / `network.hubris.homelab-context-sync.plist` (macOS) | `git pull --ff-only` every 5 min |\n| Encrypted secrets\
|
||
\ | `dtoro/oikos` | `secrets/*.yaml` (SOPS+age) | Recipients declared in `.sops.yaml` |\n| Read-only context PAT\
|
||
\ | `dtoro` Gitea user, scope `read:repository` | given to operators out-of-band | Bootstrap-only — for the initial clone\
|
||
\ before SOPS works |\n| Write-scoped PAT | `secrets/gitea-pat.yaml` (SOPS) | `homelab refresh-creds` swaps it into `/etc/homelab-context/git-credentials`\
|
||
\ | All post-bootstrap pushes (client lifecycle, wiki edits) |\n\n## Data flow\n\n```\n dtoro/oikos\
|
||
\ (gitea)\n │\n ┌────────── push ────────┤ ◀── git push (write PAT or SSH)\n \
|
||
\ │ │\n │ ┌────── push ──────┘\n │ │ │\n │ │ \
|
||
\ ▼ webhook (push event)\n │ │ ┌─── homelab-mcp-deploy ──── (LXC 105:9811)\n │ \
|
||
\ │ └─── secrets-issuance-deploy (LXC 105:9821)\n │ │ │\n │ │ \
|
||
\ ▼\n │ │ git pull → deploy.sh → restart service\n │ │\n │ └── on every client:\n\
|
||
\ │ timer (5 min) → git pull --ff-only into /opt/homelab-context\n │\n ▼\n homelab CLI /\
|
||
\ MCP server reads /opt/homelab-context for everything\n```\n\n## Why two clones on LXC 105\n\nThe MCP server and secrets-issuance\
|
||
\ each have their own clone\n(`/opt/homelab-mcp`, `/opt/secrets-issuance`) **in addition to**\n`/opt/homelab-context`.\
|
||
\ Reasons:\n\n- The deploy webhook for each service updates its own clone, runs\n `deploy.sh` from there, and re-installs\
|
||
\ the systemd unit. Mixing this\n with the client-context clone would create a circular dependency\n (deploy reinstalls\
|
||
\ the unit that pulled it).\n- The MCP server reads its data from `/opt/homelab-context` (the same path\n every client\
|
||
\ uses) so changes to inventory propagate identically. Code\n changes live in `/opt/homelab-mcp` and trigger a service\
|
||
\ restart.\n\n## Mesh / network gates\n\n- Both services bind `0.0.0.0:<port>`. The trust boundary is\n `MESH_SUBNETS`\
|
||
\ in the service's environment + nftables (planned). Today\n `MESH_SUBNETS=100.122.0.0/16,100.64.0.0/10,192.168.8.0/24`\
|
||
\ — Netbird +\n Tailscale + the homelab LAN. Adjust if the LAN ever has untrusted\n devices.\n- Caddy fronts both with\
|
||
\ Let's Encrypt certs via the IONOS DNS challenge:\n `mcp.hubris.network` → `192.168.8.205:9810`,\n `secrets.hubris.network`\
|
||
\ → `192.168.8.205:9820`. Off-LAN clients on\n Netbird reach them via the `192.168.8.0/24` network resource routed\n\
|
||
\ through the PVE peer ([mesh.md](mesh.md)).\n- Clients with default-public DNS (workstations not on Netbird, LXCs\n\
|
||
\ using router DNS) need a `/etc/hosts` override pointing\n `mcp.hubris.network` and `secrets.hubris.network` at the\
|
||
\ caddy LXC\n (`192.168.8.175`) — same caveat as every other `*.hubris.network`\n service, see [dns.md](dns.md).\n\n\
|
||
## Secrets model\n\n- Each enrolled client gets one **age private key** issued by\n secrets-issuance on first bootstrap.\
|
||
\ The key file stays root-only on\n the client; the public key is committed to `inventory.yaml` (and\n becomes a recipient\
|
||
\ on SOPS-encrypted files via `.sops.yaml`).\n- `secrets/*.yaml` are SOPS+age. Recipients per file are pinned in\n `.sops.yaml`\
|
||
\ `creation_rules` by `path_regex`. Re-encrypting a file is\n `sops updatekeys -y secrets/<name>.yaml`.\n- The MCP server's\
|
||
\ `list_my_secrets(caller_pubkey)` tool returns only\n secret *names* a given pubkey can decrypt — the server never sees\n\
|
||
\ plaintext. Decryption is local-on-client (`homelab secret <name>`\n shells out to `sops -d` with the client's key).\n\
|
||
- The \"all-clients\" secrets (`hello.yaml` for the bootstrap decrypt\n test, `gitea-pat.yaml` for the write-scoped PAT)\
|
||
\ are auto-granted to\n every newly enrolled client by `homelab client add --finalize-pubkey`\n (which appends the pubkey\
|
||
\ to the matching `.sops.yaml` rule and runs\n `sops updatekeys`).\n- **Removal does not erase past disclosure.** Revoking\
|
||
\ a client via\n `homelab client remove` shreds the issuance-side key, denylists the\n hostname, removes them from the\
|
||
\ recipient list, and re-keys all\n shared secrets — but anything they already decrypted to disk is out of\n your control.\
|
||
\ Rotate the underlying credential if compromise is\n suspected.\n\n## Why this design\n\n- **One source of truth** keeps\
|
||
\ inventory, code, secrets, and docs\n versioned together. A `git log` of `inventory.yaml` is the history of\n the homelab.\n\
|
||
- **Per-client age keys** scale better than a shared admin secret —\n removing a client is a real revocation (for new\
|
||
\ ciphertext), not just\n removing them from a wiki page.\n- **MCP layer over the same clone** gives MCP-capable agents\
|
||
\ structured\n query (`find_service`, `search_docs`) without forcing non-MCP tools to\n go without — anything can still\
|
||
\ `cat` the markdown.\n- **Sync timer rather than push fan-out** keeps the failure mode\n contained: one client's webhook\
|
||
\ outage doesn't block a push from\n landing on the others. Sub-5-min staleness is fine for docs and rare\n enough for\
|
||
\ secrets that we don't need lower latency.\n\n## Related\n\n- [Operations: agent enrollment](../../../.agents/operations/agent-enrollment.md)\
|
||
\ — the\n step-by-step for adding a new client\n- [Auto-deploy](auto-deploy.md) — the `homelab-mcp` + `secrets-issuance`\n\
|
||
\ pipelines (and the rest of the lab's webhook pipelines)\n- [Mesh](mesh.md) — Netbird / Tailscale paths and the `192.168.8.0/24`\n\
|
||
\ network resource\n- [Apps (105)](../containers/105-apps.md) — where both services run\n- [Gitea (104)](../containers/104-gitea.md)\
|
||
\ — the source of truth\n\n## Changelog\n\n### 2026-05-20 — system live across hubris, apps, republic-laptop\nPhase 1\
|
||
\ of the [cross-client context plan](../../../README.md) merged. Three\nclients enrolled end-to-end: PAT-based bootstrap,\
|
||
\ age-key issuance, SOPS\ndecrypt verified on each. Webhook auto-deploy for both LXC 105 services\nwired (hook ids 10\
|
||
\ + 11). `homelab refresh-creds` + atomic\n`client add --finalize-pubkey` grant flow live so new clients are one\nceremony\
|
||
\ instead of four manual steps. Outstanding: bootstrap mac-mini\n(macOS, exercises launchd) + ludo-mini + the remaining\
|
||
\ LXCs;\nHermes Agent integration so the agent uses inventory at chat-time.\n"
|
||
entity_slug: service:homelab-mcp
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-05-20'
|
||
title: system live across hubris, apps, republic-laptop
|
||
body: 'Phase 1 of the [cross-client context plan](../../../README.md) merged. Three clients enrolled end-to-end: PAT-based
|
||
bootstrap, age-key issuance, SOPS decrypt verified on each. Webhook auto-deploy for both LXC 105 services wired (hook
|
||
ids 10 + 11). `homelab refresh-creds` + atomic `client add --finalize-pubkey` grant flow live so new clients are one
|
||
ceremony instead of four manual steps. Outstanding: bootstrap mac-mini (macOS, exercises launchd) + ludo-mini + the
|
||
remaining LXCs; Hermes Agent integration so the agent uses inventory at chat-time.'
|
||
- slug: infrastructure/ingress
|
||
title: Public ingress — VPS traefik + cert mirror
|
||
content: "# Public ingress — VPS traefik + cert mirror\n\nHow home services reach the open internet without exposing the\
|
||
\ home network. Two-stage pattern: traefik on the IONOS VPS terminates TLS at the public edge, then reverse-proxies over\
|
||
\ the netbird mesh to home Caddy / direct backends.\n\n## The shape\n\n```\nPublic client\n │ *.hubris.network → 82.165.190.79\
|
||
\ (IONOS wildcard)\n ▼\n[VPS traefik] priority 1: HostSNI(*) → netbird-proxy:8443 ← netbird control plane\n \
|
||
\ priority 10: per-host HTTP routers ← home services\n │ HTTP over netbird mesh\n ▼\n[Home\
|
||
\ backend on 192.168.8.x]\n```\n\nLAN clients resolve via the [Technitium DNS on dns (107)](dns.md) → `192.168.8.175`\
|
||
\ → home [Caddy (121)](../containers/121-caddy.md), unchanged. The two paths are independent.\n\n## Why this shape\n\n\
|
||
- VPS traefik already has a HostSNI(*) TCP passthrough at priority 1 so netbird's own ingress (`netbird.hubris.network`,\
|
||
\ future `*.proxy.hubris.network`) is unaffected.\n- Per-hostname HTTP file-provider routers at priority 10 win over the\
|
||
\ passthrough for the listed hosts and let traefik terminate TLS itself for those.\n- Traefik's own ACME (`letsencrypt`\
|
||
\ resolver) fails on this box: HostSNI(*) grabs TLS-ALPN-01 challenges before traefik's `allowACMEByPass=true` can respond.\
|
||
\ Solution: home Caddy obtains certs via IONOS DNS-01 (no conflict) and the VPS *mirrors* the result over.\n\n## Components\n\
|
||
\n### On the VPS (`82.165.190.79`)\n\n- `/opt/traefik-dynamic.yaml` — file-watched dynamic config. One `http.routers.<name>-public`\
|
||
\ + one `http.services.<name>-public` per exposed service, plus one entry in the top-level `tls.certificates` list per\
|
||
\ hostname.\n- Let's-Encrypt-effective directory: `/letsencrypt/` inside the traefik container, backed by the host-side\
|
||
\ docker volume `opt_netbird_traefik_letsencrypt`.\n- Backups before edits: `cp /opt/traefik-dynamic.yaml /opt/traefik-dynamic.yaml.bak.$(date\
|
||
\ +%s)` — several bak files live alongside.\n\n### On the PVE host (`192.168.8.77`)\n\n- `/usr/local/bin/hubris-public-cert-sync.sh`\
|
||
\ — runs daily via `hubris-public-cert-sync.timer`. Maps source hostname → VPS cert filenames in a bash assoc array. For\
|
||
\ each mapping: `pct pull` cert+key from [Caddy (121)](../containers/121-caddy.md)'s store, diff against the VPS copy,\
|
||
\ scp only on change.\n- Filenames are stable per host so the dynamic.yaml never needs editing on renewal — traefik file-watches\
|
||
\ and hot-reloads the cert.\n\n## Services currently exposed\n\n| Hostname | Path scope \
|
||
\ | Backend | Middlewares | Cert files on VPS\
|
||
\ |\n| ------------------------------ | -------------------------------- | --------------------------------\
|
||
\ | -------------------------------------------- | ------------------------------------------ |\n| `artifacto.hubris.network`\
|
||
\ | `/p/*`, `/static/*`, `/healthz` | `192.168.8.205:3100` | `artifacto-strip-sso` + `artifacto-ratelimit`\
|
||
\ (50 rps / 100 burst) | `fullchain.crt` / `privkey.key` |\n| `blog.hubris.network` | whole host \
|
||
\ | `192.168.8.205:8080` | `blog-ratelimit` (100 rps / 200 burst) | `blog.fullchain.crt`\
|
||
\ / `blog.privkey.key` |\n| `trmnl.hubris.network` | whole host | `192.168.8.211:9851` ([trmnl\
|
||
\ 128](../containers/128-trmnl.md)) | `trmnl-ratelimit` (20 rps / 40 burst) | `trmnl.fullchain.crt` / `trmnl.privkey.key`\
|
||
\ |\n| `house.hubris.network` | whole host | `192.168.8.212:3000` ([house 129](../containers/129-house.md))\
|
||
\ | `house-ratelimit` (30 rps / 60 burst) | `house.fullchain.crt` / `house.privkey.key` |\n\n`artifacto-strip-sso`\
|
||
\ blanks inbound `X-Authentik-*` and `X-Artifacto-Gateway` so external clients can't spoof the SSO auto-login header contract.\
|
||
\ Path split is enforced at the VPS router rule, not by home Caddy. See [Artifacto on apps (105)](../containers/105-apps.md).\n\
|
||
\n### `auth.hubris.network` — different pattern (local container, not cert-mirror)\n\nSince 2026-05-31 [Authentik runs\
|
||
\ on the VPS itself](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md), so `auth.hubris.network`\
|
||
\ is served by a **local Docker container**, not proxied to a home backend. It therefore does **not** use the file-provider\
|
||
\ + cert-mirror pattern above:\n\n- Routed via traefik **Docker provider labels** on the `authentik-server` service (`/opt/docker-compose.yml`),\
|
||
\ not `traefik-dynamic.yaml`.\n- TLS via traefik's own `letsencrypt` resolver (works here because it's a normal HTTP router,\
|
||
\ not the HostSNI passthrough).\n- Traefik reaches it over the `auth` Docker network (`172.30.1.0/24`); Postgres/Redis\
|
||
\ on that net are isolated from the netbird containers.\n- Admin UI is IP-gated: an `admin-allowlist` ipAllowList middleware\
|
||
\ on `PathPrefix(/if/admin/)` (currently `5.61.168.0/24`). Login/flow endpoints stay public.\n\nNo cert-mirror entry and\
|
||
\ no `hubris-public-cert-sync.sh` mapping is needed for `auth`.\n\n## Recipe — exposing another service\n\n1. Ensure home\
|
||
\ Caddy on [LXC 121](../containers/121-caddy.md) already serves the hostname (cert exists at `/var/lib/caddy/.local/share/caddy/certificates/acme-v02.api.letsencrypt.org-directory/<host>/`).\n\
|
||
2. Add an entry to `HOSTS` in `/usr/local/bin/hubris-public-cert-sync.sh` mapping the hostname → VPS filenames. Run once:\
|
||
\ `systemctl start hubris-public-cert-sync.service`. Confirm the cert landed.\n3. Edit `/opt/traefik-dynamic.yaml` on\
|
||
\ the VPS:\n - Add to `tls.certificates`: paths `/letsencrypt/<crt>` and `/letsencrypt/<key>`.\n - Add `http.routers.<host>-public`:\
|
||
\ `rule: 'Host(\\`<host>\\`)'` (or with path matchers if scope-gating), `entryPoints: [websecure]`, `priority: 10`, `tls:\
|
||
\ {}`, `service: <host>-public`, `middlewares: [...]`.\n - Add a ratelimit middleware under `http.middlewares` if wanted.\n\
|
||
\ - Add `http.services.<host>-public.loadBalancer.servers[0].url: 'http://<lan-ip>:<port>'`.\n4. Verify:\n ```\n \
|
||
\ ssh root@100.122.165.149 'curl -skI --resolve <host>:443:127.0.0.1 https://<host>/' # 2xx/3xx\n curl -skI --resolve\
|
||
\ <host>:443:<vps-public-ip> https://<host>/ # same\n ```\n5. **No DNS edit needed** — the\
|
||
\ IONOS wildcard already points at the VPS.\n\n## What does NOT follow this pattern\n\n- `netbird.hubris.network` (and\
|
||
\ any future `*.proxy.hubris.network`) uses the netbird-proxy / HostSNI passthrough path. Netbird handles its own cert\
|
||
\ via ACME cleanly because it *is* the passthrough target.\n\n## Related\n- [DNS split-horizon](dns.md)\n- [Caddy (121)](../containers/121-caddy.md)\
|
||
\ — cert source, internal counterpart\n- [Mesh migration](mesh.md) — netbird is the transport between VPS and home\n-\
|
||
\ [VPS hardening](vps-hardening.md) — fail2ban / nftables that the access logs feed\n- [Artifacto on apps (105)](../containers/105-apps.md)\
|
||
\ — first publicly-exposed service\n\n## Changelog\n\n### 2026-06-24 — `trmnl.hubris.network` exposed\nTRMNL plugins middleware\
|
||
\ on [trmnl (128)](../containers/128-trmnl.md). File-provider router `trmnl-public` → `192.168.8.211:9851`, `trmnl-ratelimit`\
|
||
\ (20 rps / 40 burst), cert mirrored as `trmnl.fullchain.crt`/`trmnl.privkey.key`. Verified live from the internet (200\
|
||
\ with token / 401 without). It was provisioned during a mesh outage — the `home-lab-network` (192.168.8.0/24) route had\
|
||
\ no active routing peer because the **mac-mini routing peer's netbird was down** (all home-backed public services 504'd).\
|
||
\ Bringing netbird up on mac-mini restored the route; no traefik change was needed.\n\n### 2026-05-31 — `auth.hubris.network`\
|
||
\ now served locally on the VPS\nAuthentik migrated onto the VPS ([investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)).\
|
||
\ Unlike the home-backed services above, `auth` is a local container routed via traefik Docker-provider labels with traefik-managed\
|
||
\ Let's Encrypt — no cert-mirror, no `traefik-dynamic.yaml` router. Admin UI gated by an ipAllowList middleware. Traefik\
|
||
\ gained a second Docker network (`auth`, `172.30.1.0/24`) to reach it while keeping its DB/Redis isolated from the netbird\
|
||
\ stack.\n\n### 2026-04-28 — wiki entry created\nInitial documentation.\n\n### 2026-04-23 — `blog.hubris.network` exposed\n\
|
||
WriteFreely on [apps (105)](../containers/105-apps.md). Whole host is public.\n\n### 2026-04-22 — pattern established\
|
||
\ with Artifacto\nFirst service through the file-provider router. IONOS wildcard moved to the VPS this day. Cert mirror\
|
||
\ script + timer deployed on the PVE host.\n"
|
||
entity_slug: service:caddy
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-06-24'
|
||
title: '`trmnl.hubris.network` exposed'
|
||
body: TRMNL plugins middleware on [trmnl (128)](../containers/128-trmnl.md). File-provider router `trmnl-public` → `192.168.8.211:9851`,
|
||
`trmnl-ratelimit` (20 rps / 40 burst), cert mirrored as `trmnl.fullchain.crt`/`trmnl.privkey.key`. Verified live from
|
||
the internet (200 with token / 401 without). It was provisioned during a mesh outage — the `home-lab-network` (192.168.8.0/24)
|
||
route had no active routing peer because the **mac-mini routing peer's netbird was down** (all home-backed public services
|
||
504'd). Bringing netbird up on mac-mini restored the route; no traefik change was needed.
|
||
- date: '2026-05-31'
|
||
title: '`auth.hubris.network` now served locally on the VPS'
|
||
body: Authentik migrated onto the VPS ([investigation](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md)).
|
||
Unlike the home-backed services above, `auth` is a local container routed via traefik Docker-provider labels with traefik-managed
|
||
Let's Encrypt — no cert-mirror, no `traefik-dynamic.yaml` router. Admin UI gated by an ipAllowList middleware. Traefik
|
||
gained a second Docker network (`auth`, `172.30.1.0/24`) to reach it while keeping its DB/Redis isolated from the netbird
|
||
stack.
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-23'
|
||
title: '`blog.hubris.network` exposed'
|
||
body: WriteFreely on [apps (105)](../containers/105-apps.md). Whole host is public.
|
||
- date: '2026-04-22'
|
||
title: pattern established with Artifacto
|
||
body: First service through the file-provider router. IONOS wildcard moved to the VPS this day. Cert mirror script + timer
|
||
deployed on the PVE host.
|
||
- slug: infrastructure/media-permissions
|
||
title: Media permissions — `media` GID 10000
|
||
content: "# Media permissions — `media` GID 10000\n\nStandard for any LXC reading/writing `/mnt/library` on [hubris](../hosts/hubris.md).\
|
||
\ Applied 2026-04-20.\n\n## Standard\n\nEvery LXC that mounts `/mnt/library` participates in a shared `media` group with\
|
||
\ **GID 10000**. Shared subtrees are owned by that group with the setgid bit (`drwxrwsr-x`, mode `2775`), so new files\
|
||
\ auto-inherit the right group regardless of which container wrote them.\n\n## Why\n\n`/mnt/library` is a cross-container\
|
||
\ storage pool. \\*arr writes, jellyfin reads, mulita scans, paperless ingests. Without a shared group, each container\
|
||
\ sees files as `nobody:nogroup` (unprivileged) or `www-data` (privileged 1:1) and the permission web collapses into one-off\
|
||
\ chmods. GID 10000 bridges privileged and unprivileged containers.\n\n## Onboarding a new LXC\n\n1. `pct set <id> -mp0\
|
||
\ /mnt/library,mp=/mnt/library` (if not already mounted).\n2. Inside the container:\n ```\n groupadd -g 10000 media\n\
|
||
\ usermod -aG media <service-user> # for every user that needs library access\n ```\n3. If the container is **unprivileged**\
|
||
\ (check `pct config <id> | grep unprivileged`), append this idmap block to `/etc/pve/lxc/<id>.conf` (back up first):\n\
|
||
\ ```\n lxc.idmap: u 0 100000 65536\n lxc.idmap: g 0 100000 10000\n lxc.idmap: g 10000 10000 1\n lxc.idmap:\
|
||
\ g 10001 110001 55535\n ```\n Then `pct stop <id> && pct start <id>`.\n4. For systemd services running with `User=root`\
|
||
\ (not typical), add a drop-in with `SupplementaryGroups=media`. Systemd skips `initgroups()` for `User=root`.\n5. `pct\
|
||
\ exec` sessions don't get supplementary groups (no initgroups). Use `sudo -i` or `su - <user>` inside the container to\
|
||
\ verify membership interactively. Real services use `initgroups` and work correctly.\n\n## State snapshot\n\n### Host\n\
|
||
\n- Group `media` GID 10000 exists.\n- `/etc/subgid` has `root:100000:65536` AND `root:10000:1` (second line required\
|
||
\ for unprivileged LXCs to receive GID 10000).\n- Shared subtrees owned `:media` mode `2775` (drwxrwsr-x, setgid):\n \
|
||
\ - `movies`, `tv`, `music`, `anime`, `podcasts` — jellyfin libraries\n - `audiobooks`, `audiobookshelf-metadata`, `books`,\
|
||
\ `comics` — audiobookshelf / grimmory\n - `downloads` — \\*arr stack output\n - `images` — photoprism / immich / mulita\n\
|
||
\ - `roms` — emu frontends\n - `syncthing` — empty subtree, retained for archaeology (LXC 109 destroyed 2026-05-14)\n\
|
||
- Container-specific subtrees intentionally **not** migrated (keep their own owner:group):\n - `documents` (paperless,\
|
||
\ `www-data:www-data 750`)\n - `homecloud` (nextcloud — its own permission model, easy to break)\n - `marimo` (marimo\
|
||
\ venv) — *LXC since destroyed; review whether subtree still serves a purpose*\n - `notes`, `sophia` (single-container\
|
||
\ use); `heaper` — orphaned data subtree (LXC since destroyed 2026-05-14, 224 MiB retained)\n - `repos` (owner UID 102\
|
||
\ GID 105 from inside [gitea](../containers/104-gitea.md) — don't touch)\n\n### LXCs with media-group membership\n\n|\
|
||
\ ID | Name | Priv | Media-group members |\n| ---\
|
||
\ | --------------------------------------------- | ---- | --------------------------------------------- |\n| 101 | [jellyfin](../containers/101-jellyfin.md)\
|
||
\ | **unpriv + idmap** | jellyfin |\n| 103 | [paperless](../containers/103-paperless.md)\
|
||
\ | priv | www-data |\n| 104 | [gitea](../containers/104-gitea.md) |\
|
||
\ priv | www-data, gitea |\n| 105 | [apps](../containers/105-apps.md) | priv\
|
||
\ | www-data |\n| 114 | [nextcloud](../containers/114-nextcloud.md) | priv | www-data\
|
||
\ |\n| 119 | [sophia](../containers/119-sophia.md) | priv | www-data \
|
||
\ |\n| 120 | [mule-images](../containers/120-mule-images.md) | priv | www-data \
|
||
\ |\n| 122 | [arriman](../containers/122-arriman.md) | priv | www-data, audiobookshelf,\
|
||
\ radarr, sonarr, lidarr, prowlarr, qbittorrent, bazarr, jellyseerr, mylar, jackett, overseerr, plex, arr |\n| 130 | [grimmory](../containers/130-grimmory.md)\
|
||
\ | priv | Docker container uses `GROUP_ID=10000` env var (linuxserver pattern) — no in-LXC group needed |\n| 132\
|
||
\ | [rclone](../containers/132-rclone.md) | priv | **read-only** mount; runs as root → reads all subtrees. No\
|
||
\ media group needed |\n\n> Some entries from earlier snapshots — 100 (arr-yunohost), 107 (marimo), 109 (syncthing), 110\
|
||
\ (photoprism), 112 (immich), 116 (heaper) — referenced LXCs that have since been destroyed. See [containers/index](../containers/index.md#recently-destroyed-kept-for-archaeology).\n\
|
||
\nConfig backups: `/root/101.conf.bak.*`, `/root/109.conf.bak.*` (109 destroyed 2026-05-14).\n\n## Gotchas\n\n- **[apps\
|
||
\ (105)](../containers/105-apps.md) and [grimmory (130)](../containers/130-grimmory.md) are Docker hosts.** Adding `media`\
|
||
\ to the LXC alone is *not* enough for Docker containers inside. Each Docker container needs its GID passed in explicitly:\
|
||
\ `--group-add 10000`, `user: \"<uid>:10000\"`, or `GROUP_ID=10000` (linuxserver images) in compose. Grimmory, audiobookshelf-in-docker,\
|
||
\ etc. need this per-container.\n- **`pct exec` does NOT run initgroups.** So `pct exec <id> -- id` shows only the primary\
|
||
\ group. For interactive verification, use `pct exec <id> -- sudo -i -u root id` or `su - <user> -c id`. Real systemd\
|
||
\ services work fine.\n- **systemd `User=root`** skips initgroups — explicit `SupplementaryGroups=media` drop-in needed.\n\
|
||
- **`pct restore`** or template rebuilds wipe in-container group membership and unprivileged-LXC idmap blocks. Re-apply\
|
||
\ from this page.\n- **`/etc/subgid`** must retain both `root:100000:65536` AND `root:10000:1`. Dropping the second breaks\
|
||
\ startup of any unprivileged LXC with the idmap block.\n- **\\*arr \"Set Permissions\" options** can override the setgid\
|
||
\ inheritance by explicitly chown'ing files. Leave those off, or set the group to `media`. Relevant to Sonarr/Radarr/qBittorrent\
|
||
\ on [arriman (122)](../containers/122-arriman.md).\n- **Nextcloud** files under `/mnt/library/homecloud` are deliberately\
|
||
\ NOT in the media group. NC manages its own permission model. See [nextcloud (114)](../containers/114-nextcloud.md).\n\
|
||
- **\\*arr stack on arriman** required `MEDIACENTER_GID=10000` (not 13000) in `.env` because s6-setuidgid only honors\
|
||
\ the primary PGID; `group_add:` doesn't propagate. See [arriman (122)](../containers/122-arriman.md#changelog).\n\n##\
|
||
\ Related\n- [Hubris host](../hosts/hubris.md)\n- All container pages list whether they're in the standard\n\n## Changelog\n\
|
||
\n### 2026-05-14 — LXC 109 (syncthing) destroyed\nRemoved the syncthing row from the membership table and the syncthing-as-`User=root`\
|
||
\ example from the onboarding section. `/mnt/library/syncthing` subtree was already empty and retained as an empty dir.\n\
|
||
\n### 2026-05-14 — LXC 116 (heaper) destroyed\nRemoved the heaper row from the LXC membership table and noted the orphaned\
|
||
\ `/mnt/library/heaper` subtree (224 MiB retained). See [host changelog](../hosts/hubris.md#changelog).\n\n### 2026-04-28\
|
||
\ — wiki entry created\nInitial documentation.\n\n### 2026-04-26 — `MEDIACENTER_GID` fix on [arriman (122)](../containers/122-arriman.md)\n\
|
||
qBit was erroring every torrent with \"Permission denied\" because `MEDIACENTER_GID=13000` was set as a supplementary\
|
||
\ GID via `group_add:`. Changed to 10000 (primary GID); fix described above is now standard.\n\n### 2026-04-20 — standard\
|
||
\ rolled out\nGID 10000 hostgroup, idmap blocks for unprivileged LXCs, setgid 2775 on shared subtrees, `media` membership\
|
||
\ for service users in every participating LXC.\n"
|
||
entity_slug: service:jellyfin
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-05-14'
|
||
title: LXC 109 (syncthing) destroyed
|
||
body: Removed the syncthing row from the membership table and the syncthing-as-`User=root` example from the onboarding
|
||
section. `/mnt/library/syncthing` subtree was already empty and retained as an empty dir.
|
||
- date: '2026-05-14'
|
||
title: LXC 116 (heaper) destroyed
|
||
body: Removed the heaper row from the LXC membership table and noted the orphaned `/mnt/library/heaper` subtree (224 MiB
|
||
retained). See [host changelog](../hosts/hubris.md#changelog).
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-26'
|
||
title: '`MEDIACENTER_GID` fix on [arriman (122)](../containers/122-arriman.md)'
|
||
body: qBit was erroring every torrent with "Permission denied" because `MEDIACENTER_GID=13000` was set as a supplementary
|
||
GID via `group_add:`. Changed to 10000 (primary GID); fix described above is now standard.
|
||
- date: '2026-04-20'
|
||
title: standard rolled out
|
||
body: GID 10000 hostgroup, idmap blocks for unprivileged LXCs, setgid 2775 on shared subtrees, `media` membership for
|
||
service users in every participating LXC.
|
||
- slug: infrastructure/mesh
|
||
title: Mesh — Tailscale → Netbird migration
|
||
content: "# Mesh — Tailscale → Netbird migration\n\nThe hubris fleet is migrating from Tailscale to Netbird. Netbird is\
|
||
\ the target end-state. In-progress as of 2026-04-21.\n\n## Current state\n\n- **PVE host** uses Netbird (`wt0`, `100.122.38.109/16`).\
|
||
\ Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets `*.hubris.network →\
|
||
\ 192.168.8.175` via the system resolver chain.\n- **Netbird mgmt host** (`82.165.190.79`, FQDN `inspiring-ramanujan.netbird.selfhosted`,\
|
||
\ NB IP `100.122.165.149`) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes\
|
||
\ the homelab network (`192.168.8.0/24`) via the PVE peer. This gives the mgmt host LAN access *and* split-horizon DNS\
|
||
\ for `*.hubris.network`. Useful independently of any Authentik integration.\n- **Most LXCs** still run Tailscale or use\
|
||
\ router DNS (`192.168.8.1`) / Tailscale MagicDNS (`100.100.100.100`), both of which return the *public* IONOS A record\
|
||
\ `*.hubris.network → 82.165.190.79`. The VPS only routes hostnames it actually publishes (today, `artifacto` + `blog`),\
|
||
\ so this path is a dead end for any LAN-only service.\n\n## ICE / STUN / TURN\n\n**Today** (post-2026-05-21 migration):\n\
|
||
\n- External STUN servers (Google + Cloudflare) declared under `server.Stuns` in `/opt/management.json`. Embedded STUN\
|
||
\ is no longer in use.\n- coturn runs on the VPS host (apt-installed, systemd-managed). TURN URL `turn:netbird.hubris.network:3478?transport=tcp`\
|
||
\ is advertised to peers via mgmt's `TURNConfig.Turns` block. Long-term credentials at `user=netbird:<pwd from /root/turn-pass.txt>`.\n\
|
||
- For non-symmetric peers, ICE picks direct `srflx/srflx` (P2P). For symmetric-NAT peers, ICE falls through to TURN-relay\
|
||
\ before falling back to the WSS relay (`rels://netbird.hubris.network:443`).\n\n**IONOS port-3478 caveat** (load-bearing\
|
||
\ — undocumented before 2026-05-21):\n\nIONOS upstream filters STUN-class traffic on port 3478 **for both UDP and TCP**\
|
||
\ by default. The UDP block was already known (verified 2026-05-10 with `tcpdump -ni any udp port 3478`). The TCP block\
|
||
\ was discovered 2026-05-21 when TURN-allocate retries timed out from a symmetric-NAT laptop: TCP handshake completed\
|
||
\ (carried by kernel-only SYN exchange) but data packets never reached the VPS's `ens6` interface.\n\n**Resolution**:\
|
||
\ operator must add an inbound firewall exception in the IONOS dashboard (Cloud Panel → Networks → Firewall policy on\
|
||
\ the VPS) for inbound TCP 3478 (and UDP 3478 if STUN-via-VPS-listener is ever needed; we use external STUN so this is\
|
||
\ optional). Without that exception, coturn is unreachable from outside even though host nftables + listener look fine.\n\
|
||
\n**Verifying TURN works** end-to-end from an outside peer:\n\n```python\n# python3\nimport socket, struct, secrets\n\
|
||
s = socket.create_connection((\"netbird.hubris.network\", 3478), timeout=10)\ntid = secrets.token_bytes(12)\nattrs = struct.pack(\"\
|
||
!HHI\", 0x0019, 4, 17 << 24) # REQUESTED-TRANSPORT = UDP\nmsg = struct.pack(\"!HHI\", 0x0003, len(attrs), 0x2112A442)\
|
||
\ + tid + attrs # Allocate\ns.sendall(msg)\nprint(s.recv(4096).hex()) # expect 120 bytes starting 0x0113 (Allocate Error\
|
||
\ = auth challenge)\n```\n\nA 120-byte `0x0113`-typed response means coturn is reachable and responding. An indefinite\
|
||
\ TCP timeout (after `connected from ...` prints) means the IONOS rule was reverted or narrowed.\n\n**If a peer is still\
|
||
\ on the WSS relay after this** — verify in `netbird status -d`: `Connection type: P2P` with srflx/srflx is the cone-NAT\
|
||
\ happy path; `Connection type: P2P` with relay/srflx (or relay/prflx) means the peer's network is symmetric-NAT and it's\
|
||
\ using TURN. Falling back to `Relayed (rels://...)` only happens if TURN allocate also fails — investigate the IONOS\
|
||
\ exception first.\n\n**Old combined-server note (history, kept for context):**\n\nPre-migration, the bundled `netbirdio/netbird-server`\
|
||
\ combined image silently ignored both `server.turns:` and top-level `TURNConfig:` YAML, so TURN was non-functional. That's\
|
||
\ why the architectural migration to the canonical multi-container stack (`mgmt + signal + relay + coturn`) happened.\
|
||
\ See the 2026-05-21 changelog entry below and the [Netbird-combined-no-TURN finding](https://github.com/netbirdio/netbird/issues)\
|
||
\ in homelab memory for the full discovery.\n\n## Consequence — every LXC wired to Authentik needs an internal override\n\
|
||
\nUntil each LXC is migrated to Netbird, anything that needs to reach `auth.hubris.network` (Authentik), `cloud.hubris.network`\
|
||
\ (Nextcloud), etc., must override the public answer with `192.168.8.175`.\n\nTwo techniques. Pick by HTTP-client behavior.\n\
|
||
\n### A) `/etc/hosts` override\n\nWorks for libc `getaddrinfo` clients: curl, wget, most Go/Python/Ruby apps, gitea.\n\
|
||
\n- Place the line **outside** the `# --- BEGIN PVE ---` / `# --- END PVE ---` markers. Proxmox rewrites everything inside\
|
||
\ that block on every container start.\n- Belt-and-suspenders: `/etc/systemd/system/hubris-hosts-override.service` (oneshot,\
|
||
\ enabled, idempotent).\n\n### B) Local dnsmasq\n\nRequired for clients that bypass `/etc/hosts`. **Nextcloud (PHP Guzzle\
|
||
\ + `OC\\Http\\Client\\DnsPinMiddleware`) is one** — uses `dns_get_record()`, not `getaddrinfo`. Likely candidates for\
|
||
\ the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting.\n\nRecipe:\n```\napt install\
|
||
\ dnsmasq\n\ncat > /etc/dnsmasq.d/hubris-internal.conf <<EOF\naddress=/auth.hubris.network/192.168.8.175\nserver=192.168.8.1\n\
|
||
server=1.1.1.1\ninterface=lo\nbind-interfaces\nno-hosts\nno-resolv\nEOF\n\n# Set LXC default nameservers and live resolv.conf\n\
|
||
pct set <id> --nameserver \"127.0.0.1 192.168.8.1 1.1.1.1\"\n# Then update /etc/resolv.conf inside the LXC too.\n```\n\
|
||
\n### Known overrides applied\n\n| LXC | Technique \
|
||
\ | Notes |\n| ------------------------------------------ | ---------------------------------------- | ----- |\n| [104\
|
||
\ (gitea)](../containers/104-gitea.md) | `/etc/hosts` + `hubris-hosts-override.service` | Standard |\n| [114 (nextcloud)](../containers/114-nextcloud.md)\
|
||
\ | local dnsmasq @ 127.0.0.1 + resolv.conf reorder | Guzzle bypasses /etc/hosts |\n| [105 (apps)](../containers/105-apps.md),\
|
||
\ inner containers | `extra_hosts:` in compose | Booklore, mulita, WriteFreely each ship with this |\n\n## Adding new\
|
||
\ LXCs\n\n- Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris.\n- When\
|
||
\ wiring a new app into Authentik: `cat /etc/resolv.conf` on the target LXC. If nameserver is `192.168.8.1` or `100.100.100.100`,\
|
||
\ add the hosts override. If it's the netbird daemon IP, skip.\n\n## Long-term fix\n\nEither:\n- Split-horizon DNS at\
|
||
\ LAN/router level so `*.hubris.network → 192.168.8.175` for every LAN client. Eliminates all per-LXC overrides.\n- Once\
|
||
\ Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers\
|
||
\ via the system resolver chain on the PVE host.\n\n## CRITICAL — never `docker compose up` Portainer-managed stacks\n\
|
||
\n[apps (105)](../containers/105-apps.md) runs multiple stacks deployed via the Portainer UI (`/var/lib/docker/volumes/portainer_data/_data/compose/<N>/`).\
|
||
\ Running `docker compose up -d <svc>` from the host shell on these triggers a recreate of OTHER services in the stack\
|
||
\ — labels drift, compose treats the project as reconciling, containers get destroyed and remade. **This wiped Booklore's\
|
||
\ mariadb data on 2026-04-22** (bind mount `./mariadb/config` re-initialized fresh).\n\nRecipe for container-config changes\
|
||
\ (e.g. adding `extra_hosts`) on Portainer-managed stacks:\n1. Edit the compose in Portainer UI → **Stacks → <stack> →\
|
||
\ Editor → Update stack**. Portainer handles the recreate cleanly with its own state tracking.\n2. Do NOT edit `/var/lib/docker/volumes/portainer_data/_data/compose/<N>/docker-compose.yml`\
|
||
\ directly.\n3. For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the\
|
||
\ app's built-in OIDC over container-level `extra_hosts` when possible.\n\n## Related\n- [DNS split-horizon](dns.md)\n\
|
||
- [Authentik (124)](../containers/106-auth-outpost.md) — the IdP that triggers most of these overrides\n- [Nextcloud (114)](../containers/114-nextcloud.md)\
|
||
\ — example of Technique B\n- [Gitea (104)](../containers/104-gitea.md) — example of Technique A\n- [Public ingress (VPS\
|
||
\ traefik)](ingress.md) — uses the same mesh as transport\n\n## Changelog\n\n### 2026-05-31 (later) — Authentik moved\
|
||
\ to the VPS; mesh-dependency for auth eliminated (supersedes the band-aid below)\nThe earlier same-day fix routed `auth.hubris.network`\
|
||
\ through VPS Traefik → Caddy → LXC 124 **over the mesh**. That restored service but re-created the original fragility:\
|
||
\ if the mesh is dark when management restarts, the `192.168.8.175` backend is unreachable and management crash-loops\
|
||
\ again (the \"Bootstrap note\" in the entry below). That note is now **obsolete** — Authentik was migrated onto the VPS\
|
||
\ itself, so OIDC no longer touches the mesh. The `auth-authentik` → `192.168.8.175` route and its `skip-verify` transport\
|
||
\ were removed from `/opt/traefik-dynamic.yaml`; `auth.hubris.network` is now served by a local `authentik-server` container\
|
||
\ via Traefik Docker-provider labels, and netbird-mgmt has `depends_on: authentik-server: condition: service_healthy`.\
|
||
\ The socat / reverse-SSH bootstrap dance is no longer needed. Full detail: [2026-05-31 Authentik VPS migration](../../sources/investigations/archive/2026-05-31-authentik-vps-migration.md).\n\
|
||
\n### 2026-05-31 — Netbird mesh recovered; auth.hubris.network exposed via VPS Traefik\n\n**Symptom:** `netbird-mgmt`\
|
||
\ crash-looped for ~9 days (since 2026-05-21 migration). All peers showed `Connecting`, management returned `404 (Not\
|
||
\ Found)` for gRPC → `EOF` on startup.\n\n**Root cause:** The 2026-05-21 migration configured `management.json` with `OIDCConfigEndpoint:\
|
||
\ https://auth.hubris.network/...`. The public DNS for `auth.hubris.network` already pointed to `82.165.190.79` (VPS),\
|
||
\ but the VPS Traefik had no route for that host → management got EOF on every boot, crash-looped.\n\n**Fix:**\n1. Added\
|
||
\ `auth-hubris` router to `/opt/traefik-dynamic.yaml` — `Host(auth.hubris.network)` with `certResolver: letsencrypt` →\
|
||
\ service `auth-authentik`.\n2. Service backend: `https://192.168.8.175` (Caddy on hubris LAN) via the Netbird mesh.\n\
|
||
3. ServersTransport `skip-verify` with `serverName: auth.hubris.network` + `insecureSkipVerify: true` so Traefik sends\
|
||
\ the correct SNI to Caddy.\n4. Upgraded VPS Netbird client `0.69.0 → 0.71.3` via apt (`apt install netbird=0.71.3`).\n\
|
||
5. Mesh fully recovered; management connected to peers within ~1 min.\n\n**Bootstrap note (if mesh is dark and management\
|
||
\ must restart):** If the WireGuard tunnels are fully dead AND management needs to restart (after e.g. a VPS reboot with\
|
||
\ no peer handshakes), the Traefik backend `192.168.8.175` will be unreachable and management will crash again. Recovery:\
|
||
\ temporarily open VPS port 22 via IONOS console (`nft insert rule inet hubris-fw input iifname \"ens6\" tcp dport 22\
|
||
\ accept`), SSH in, run `socat TCP-LISTEN:8443,bind=172.30.0.1,reuseaddr,fork TCP:127.0.0.1:8443 &`, then from hubris\
|
||
\ `ssh -f -N -R 127.0.0.1:8443:192.168.8.175:443 root@82.165.190.79` — this bootstraps one management start, after which\
|
||
\ the mesh self-heals.\n\n**Architecture after this change:** `auth.hubris.network` is publicly accessible (HTTPS via\
|
||
\ VPS Traefik → Caddy on hubris → Authentik LXC 124). External devices authenticating to Netbird now hit this public path.\
|
||
\ Phase 6 (Authentik as Netbird IdP) is complete and live.\n\n### 2026-05-21 — VPS migrated combined → vanilla netbird\
|
||
\ stack with external TURN\nThe combined `netbirdio/netbird-server` image was replaced with the canonical multi-container\
|
||
\ deploy (`netbirdio/management:0.71.3` + `signal:0.71.3` + `relay:0.71.3` + `dashboard:latest` + host coturn) on `/opt/docker-compose.yml`.\
|
||
\ Driver: combined image silently ignored external `TURNConfig` so symmetric-NAT peers couldn't use TURN.\n\nSame migration\
|
||
\ also swapped OIDC from the combined image's embedded Dex IdP to Authentik on [LXC 124](../containers/106-auth-outpost.md),\
|
||
\ upgrading mgmt to 0.71.3. The `store.db` schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old `opt_netbird_data`\
|
||
\ volume into the new `mgmt_data` volume). Pre-cutover backups at `/root/netbird-*.tgz` on the VPS, ~857 MB, retained\
|
||
\ for ~7d.\n\nAlso during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478\
|
||
\ inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe).\n\nThe new Authentik\
|
||
\ provider for NetBird is `Client type: Public` (PKCE-only). Confidential would break the dashboard SPA's token exchange.\
|
||
\ The Device Code grant flow is wired (see [containers/124-authentik.md](../containers/106-auth-outpost.md#device-code-grant--configured-2026-05-21))\
|
||
\ so interactive `netbird up` works — `--setup-key` is no longer required for new peers.\n\n**Post-migration JWT-issuer\
|
||
\ gotcha on existing peers** (cost ~30 min to diagnose 2026-05-21):\n\nExisting peers — registered against the old combined\
|
||
\ image's embedded Dex IdP at `https://netbird.hubris.network/oauth2` — cache the OLD expected SSH-JWT issuer in the netbird\
|
||
\ daemon's in-memory state. After the migration, incoming `netbird ssh` connections were rejected with:\n\n```\nJWT authentication\
|
||
\ failed: validate token (\n expected issuer=https://netbird.hubris.network/oauth2,\n audiences=[netbird-dashboard\
|
||
\ netbird-cli],\n actual issuer=https://auth.hubris.network/application/o/netbird/,\n audience=netbird-dashboard\n\
|
||
)\n```\n\nNeither `systemctl restart netbird` nor `netbird down && netbird up` clears the cache. Root cause: in `client/internal/engine_ssh.go`,\
|
||
\ `updateSSH()` bails out with `if e.sshServer != nil { return nil }` whenever the SSH server is already running, so mgmt-pushed\
|
||
\ JWT config updates are silently ignored. Only a full daemon-process tear-down lets the SSH server re-initialize with\
|
||
\ the new validator config:\n\n```\nsudo systemctl stop netbird\nsleep 3\nsudo systemctl start netbird\n```\n\nAfter that,\
|
||
\ `grep -iE \"issuer|audience\" /var/log/netbird/client.log | tail` shows the new Authentik issuer. Run this on every\
|
||
\ existing peer (PVE host + every LXC + every workstation) once after a future IdP swap.\n\n**Username gotcha (related):**\
|
||
\ `netbird ssh` defaults the remote username to the LOCAL one (e.g. `dtoro` from the operator's laptop). Hubris + the\
|
||
\ LXCs only have `root`, so the JWT is accepted but the session immediately fails with `user dtoro not found`. Always\
|
||
\ use the explicit `root@` prefix when invoking netbird-ssh manually:\n\n```\nnetbird ssh -p 22022 root@proxmox-server.netbird.selfhosted\n\
|
||
```\n\nThe `homelab` CLI handles this automatically via the per-host `ssh.user` field in `inventory.yaml` (defaults to\
|
||
\ `root`; set explicitly only for workstations whose login user isn't `root`).\n\nOpen follow-up: TURN-over-TLS on TCP\
|
||
\ 5349 (cert via certbot or extract Traefik's acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient\
|
||
\ for current usage.\n\n### 2026-05-10 — ICE direct p2p restored (external STUN swap)\nAll peers were `Connection type:\
|
||
\ Relayed` because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to\
|
||
\ Google + Cloudflare STUN in `/opt/config.yaml`. After `netbird down/up`, peers now report `P2P` with srflx/host candidates.\
|
||
\ Backup of pre-change config at `/opt/config.yaml.bak-20260510-185050`. Fixes a real-world Nextcloud upload bottleneck\
|
||
\ (~366 kB/s through the relay → LAN-direct on same-network peers).\n\n### 2026-04-28 — wiki entry created\nInitial documentation.\n\
|
||
\n### 2026-04-22 — Booklore mariadb data wiped (lesson recorded)\nThe \"never docker compose up Portainer-managed stacks\"\
|
||
\ rule comes from this. See [apps (105)](../containers/105-apps.md#changelog).\n\n### 2026-04-22 — netbird mgmt host joined\
|
||
\ its own mesh\n`82.165.190.79` is now a peer (`100.122.165.149`). LAN access + split-horizon DNS via PVE peer. See [VPS\
|
||
\ hardening](vps-hardening.md).\n\n### 2026-04-21 — overrides applied to gitea (104) and nextcloud (114)\nTwo techniques\
|
||
\ documented; Nextcloud forced the dnsmasq route because of Guzzle.\n"
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-05-31'
|
||
title: Netbird mesh recovered; auth.hubris.network exposed via VPS Traefik
|
||
body: '**Symptom:** `netbird-mgmt` crash-looped for ~9 days (since 2026-05-21 migration). All peers showed `Connecting`,
|
||
management returned `404 (Not Found)` for gRPC → `EOF` on startup. **Root cause:** The 2026-05-21 migration configured
|
||
`management.json` with `OIDCConfigEndpoint: https://auth.hubris.network/...`. The public DNS for `auth.hubris.network`
|
||
already pointed to `82.165.190.79` (VPS), but the VPS Traefik had no route for that host → management got EOF on every
|
||
boot, crash-looped. **Fix:** 1. Added `auth-hubris` router to `/opt/traefik-dynamic.yaml` — `Host(auth.hubris.network)`
|
||
with `certResolver: letsencrypt` → service `auth-authentik`. 2. Service backend: `https://192.168.8.175` (Caddy on hubris
|
||
LAN) via the Netbird mesh. 3. ServersTransport `skip-verify` with `serverName: auth.hubris.network` + `insecureSkipVerify:
|
||
true` so Traefik sends the correct SNI to Caddy. 4. Upgraded VPS Netbird client `0.69.0 → 0.71.3` via apt (`apt install
|
||
netbird=0.71.3`). 5. Mesh fully recovered; management connected to peers within ~1 min. **Bootstrap note (if mesh is
|
||
dark and management must restart):** If the WireGuard tunnels are fully dead AND management needs to restart (after
|
||
e.g. a VPS reboot with no peer handshakes), the Traefik backend `192.168.8.175` will be unreachable and management will
|
||
crash again. Recovery: temporarily open VPS port 22 via IONOS console (`nft insert rule inet hubris-fw input iifname
|
||
"ens6" tcp dport 22 accept`), SSH in, run `socat TCP-LISTEN:8443,bind=172.30.0.1,reuseaddr,fork TCP:127.0.0.1:8443 &`,
|
||
then from hubris `ssh -f -N -R 127.0.0.1:8443:192.168.8.175:443 root@82.165.190.79` — this bootstraps one management
|
||
start, after which the mesh self-heals. **Architecture after this change:** `auth.hubris.network` is publicly accessible
|
||
(HTTPS via VPS Traefik → Caddy on hubris → Authentik LXC 124). External devices authenticating to Netbird now hit this
|
||
public path. Phase 6 (Authentik as Netbird IdP) is complete and live.'
|
||
- date: '2026-05-21'
|
||
title: VPS migrated combined → vanilla netbird stack with external TURN
|
||
body: 'The combined `netbirdio/netbird-server` image was replaced with the canonical multi-container deploy (`netbirdio/management:0.71.3`
|
||
+ `signal:0.71.3` + `relay:0.71.3` + `dashboard:latest` + host coturn) on `/opt/docker-compose.yml`. Driver: combined
|
||
image silently ignored external `TURNConfig` so symmetric-NAT peers couldn''t use TURN. Same migration also swapped
|
||
OIDC from the combined image''s embedded Dex IdP to Authentik on [LXC 124](../containers/106-auth-outpost.md), upgrading
|
||
mgmt to 0.71.3. The `store.db` schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old `opt_netbird_data`
|
||
volume into the new `mgmt_data` volume). Pre-cutover backups at `/root/netbird-*.tgz` on the VPS, ~857 MB, retained
|
||
for ~7d. Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478
|
||
inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe). The new Authentik provider
|
||
for NetBird is `Client type: Public` (PKCE-only). Confidential would break the dashboard SPA''s token exchange. The
|
||
Device Code grant flow is wired (see [containers/124-authentik.md](../containers/106-auth-outpost.md#device-code-grant--configured-2026-05-21))
|
||
so interactive `netbird up` works — `--setup-key` is no longer required for new peers. **Post-migration JWT-issuer gotcha
|
||
on existing peers** (cost ~30 min to diagnose 2026-05-21): Existing peers — registered against the old combined image''s
|
||
embedded Dex IdP at `https://netbird.hubris.network/oauth2` — cache the OLD expected SSH-JWT issuer in the netbird daemon''s
|
||
in-memory state. After the migration, incoming `netbird ssh` connections were rejected with: ``` JWT authentication
|
||
failed: validate token ( expected issuer=https://netbird.hubris.network/oauth2, audiences=[netbird-dashboard netbird-cli],
|
||
actual issuer=https://auth.hubris.network/application/o/netbird/, audience=netbird-dashboard ) ``` Neither `systemctl
|
||
restart netbird` nor `netbird down && netbird up` clears the cache. Root cause: in `client/internal/engine_ssh.go`,
|
||
`updateSSH()` bails out with `if e.sshServer != nil { return nil }` whenever the SSH server is already running, so mgmt-pushed
|
||
JWT config updates are silently ignored. Only a full daemon-process tear-down lets the SSH server re-initialize with
|
||
the new validator config: ``` sudo systemctl stop netbird sleep 3 sudo systemctl start netbird ``` After that, `grep
|
||
-iE "issuer|audience" /var/log/netbird/client.log | tail` shows the new Authentik issuer. Run this on every existing
|
||
peer (PVE host + every LXC + every workstation) once after a future IdP swap. **Username gotcha (related):** `netbird
|
||
ssh` defaults the remote username to the LOCAL one (e.g. `dtoro` from the operator''s laptop). Hubris + the LXCs only
|
||
have `root`, so the JWT is accepted but the session immediately fails with `user dtoro not found`. Always use the explicit
|
||
`root@` prefix when invoking netbird-ssh manually: ``` netbird ssh -p 22022 root@proxmox-server.netbird.selfhosted ```
|
||
The `homelab` CLI handles this automatically via the per-host `ssh.user` field in `inventory.yaml` (defaults to `root`;
|
||
set explicitly only for workstations whose login user isn''t `root`). Open follow-up: TURN-over-TLS on TCP 5349 (cert
|
||
via certbot or extract Traefik''s acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient for current
|
||
usage.'
|
||
- date: '2026-05-10'
|
||
title: ICE direct p2p restored (external STUN swap)
|
||
body: 'All peers were `Connection type: Relayed` because the embedded STUN on the VPS was unreachable from outside (IONOS
|
||
drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in `/opt/config.yaml`. After `netbird down/up`, peers now
|
||
report `P2P` with srflx/host candidates. Backup of pre-change config at `/opt/config.yaml.bak-20260510-185050`. Fixes
|
||
a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers).'
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-22'
|
||
title: Booklore mariadb data wiped (lesson recorded)
|
||
body: The "never docker compose up Portainer-managed stacks" rule comes from this. See [apps (105)](../containers/105-apps.md#changelog).
|
||
- date: '2026-04-22'
|
||
title: netbird mgmt host joined its own mesh
|
||
body: '`82.165.190.79` is now a peer (`100.122.165.149`). LAN access + split-horizon DNS via PVE peer. See [VPS hardening](vps-hardening.md).'
|
||
- date: '2026-04-21'
|
||
title: overrides applied to gitea (104) and nextcloud (114)
|
||
body: Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle.
|
||
- slug: infrastructure/monitoring
|
||
title: Monitoring — Hermes health watchdog
|
||
content: "# Monitoring — Hermes health watchdog\n\nHomelab health monitoring via Hermes Agent on mac-mini. Replaced the\
|
||
\ legacy\n`claudio-monitor` + `claudio-bot` IPC pipeline on 2026-06-04.\n\n## Current approach\n\nTwo layers:\n\n1. **On-demand:**\
|
||
\ ask Hermes \"how's the homelab?\" or run `homelab health` — loads\n the `homelab-hardware-health` skill, checks hardware\
|
||
\ temps, LXC resources,\n service reachability, and apt/docker drift across all hosts.\n\n2. **Cron watchdog:** `homelab-health-watchdog`\
|
||
\ runs every 15 minutes via Hermes\n cron. Silent when healthy. When thresholds breach, sends an actionable alert\n\
|
||
\ to Matrix (`@dtoro:avispero`) with options the user can reply to directly\n (e.g. \"resize rootfs\", \"investigate\"\
|
||
, \"snooze 24h\"). Hermes takes action on\n the selected option via SSH.\n\nThresholds: LXC disk >80% warn/>90% critical,\
|
||
\ NVMe >60°C/>70°C, CPU >70°C/>80°C,\napt >10/>50 upgradable, services down.\n\nHome Assistant pulls PVE metrics independently\
|
||
\ via its Proxmox VE integration\n(unaffected by this change).\n\n## Legacy: claudio-monitor (deprecated 2026-06-04)\n\
|
||
\nThe old system was a bash watchdog on hubris (`claudio-monitor.timer`, every 5\nmin) that POSTed alerts to a Matrix\
|
||
\ bot (`@claudio:avispero`) via an IPC server\non LXC 123:9090. All components decommissioned:\n\n| Component | Fate |\n\
|
||
|-----------|------|\n| LXC 123 (claudio-bot) | Destroyed 2026-06-04 |\n| `dtoro/claudio-bot` | Archived (read-only) on\
|
||
\ Gitea |\n| `dtoro/claudio-monitor` | Archived (read-only) on Gitea |\n| `claudio-monitor.timer` | Disabled on hubris\
|
||
\ |\n| `/opt/claudio-monitor/` | Still on hubris (cleanup pending) |\n| `/etc/claudio-monitor/` | Still on hubris (cleanup\
|
||
\ pending) |\n\nFor the full deprecation plan, see `plans/2026-06-04_130000-deprecate-claudio-bot.md`.\n\n## Related pages\n\
|
||
- [Hubris host](../hosts/hubris.md)\n- [HAOS VM (108)](../vms/108-haos.md)\n- [Backups (disabled)](backups.md)\n- [Homelab\
|
||
\ context distribution](homelab-context.md)\n\n## Changelog\n\n### 2026-06-04 — migrated to Hermes health watchdog\nclaudio-monitor\
|
||
\ + claudio-bot IPC pipeline replaced by Hermes-native monitoring.\nOn-demand `homelab health` via extended skill; 15-min\
|
||
\ cron watchdog with actionable\nMatrix alerts. LXC 123 destroyed, repos archived.\n\n### 2026-04-28 — wiki entry created\n\
|
||
Initial documentation.\n\n### 2026-04-21 — claudio-monitor stood up; thermal-watch removed\nGeneral health monitor with\
|
||
\ per-LXC checks. MQTT/REST push paths ripped out.\n"
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-06-04'
|
||
title: migrated to Hermes health watchdog
|
||
body: claudio-monitor + claudio-bot IPC pipeline replaced by Hermes-native monitoring. On-demand `homelab health` via
|
||
extended skill; 15-min cron watchdog with actionable Matrix alerts. LXC 123 destroyed, repos archived.
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-21'
|
||
title: claudio-monitor stood up; thermal-watch removed
|
||
body: General health monitor with per-LXC checks. MQTT/REST push paths ripped out.
|
||
- slug: infrastructure/network
|
||
title: Network
|
||
content: "# Network\n\nPhysical and logical network topology for the homelab.\n\n## Why\n\nThe homelab runs on a dedicated\
|
||
\ internal subnet (`192.168.8.0/24`) isolated from the main household LAN (`192.168.178.0/24`). Isolation is enforced\
|
||
\ at Proxmox: LXC/VM traffic is bridged only on the internal `vmbr0` bridge; Proxmox routes packets out to Fritz!Box via\
|
||
\ `vmbr1`. The main LAN cannot reach homelab services directly without a Fritz!Box static route (which is configured to\
|
||
\ allow inbound).\n\nFritz!OS 8.x does not support second IP networks on LAN ports, so Proxmox (`hubris`) acts as the\
|
||
\ subnet router rather than the Fritz!Box.\n\n## Hardware\n\n| Device | Role |\n|---|---|\n| Fritz!Box 7590 | Main router\
|
||
\ / ISP gateway (`192.168.178.1`) |\n| SODOLA 5-Port 2.5Gbit Managed | Homelab switch — flat L2, all ports native |\n\
|
||
| hubris (Proxmox) | Subnet router — routes between `192.168.8.0/24` and `192.168.178.0/24` |\n\n## Topology\n\n```\n\
|
||
ISP\n └── Fritz!Box 7590 (192.168.178.1)\n │ static route: 192.168.8.0/24 → 192.168.178.10\n │\n └──\
|
||
\ SODOLA 5-Port 2.5Gbit\n ├── Port 1 uplink → Fritz!Box LAN\n ├── Port 2 hubris eno1 → vmbr1 (192.168.178.10)\n\
|
||
\ ├── Port 3 [device]\n ├── Port 4 [device]\n └── Port 5 spare\n\nhubris internal bridges:\n\
|
||
\ vmbr1 192.168.178.10/24 eno1 (uplink, DHCP-reserved) gateway 192.168.178.1\n vmbr0 192.168.8.77/24 no physical\
|
||
\ port (internal)\n 192.168.8.1/24 alias — LXC default gateway\n ├── all 16 LXCs\n └──\
|
||
\ HAOS VM\n```\n\n## Subnets\n\n| Subnet | Gateway | Purpose |\n|---|---|---|\n| `192.168.178.0/24` | `192.168.178.1`\
|
||
\ | Household LAN — laptops, phones, Fritz!Box DHCP |\n| `192.168.8.0/24` | `192.168.8.1` (Proxmox `vmbr0` alias) | Homelab\
|
||
\ — all LXCs and VMs |\n\n## DHCP\n\n- **Household (`192.168.178.x`)**: Fritz!Box built-in DHCP. Proxmox `vmbr1` has a\
|
||
\ reservation: MAC `84:47:09:6b:e7:58` → `192.168.178.10`.\n- **Homelab (`192.168.8.x`)**: Technitium on [CT 107](../containers/107-dns.md)\
|
||
\ at `192.168.8.2`. Range `192.168.8.241–192.168.8.254`, gateway `192.168.8.1`, DNS `192.168.8.2`.\n\nStatic IPs span\
|
||
\ `.101–.239` (all LXCs, VMs, and workstations). DHCP pool narrowed to `.241–.254` (2026-06-03) to avoid overlap and IP\
|
||
\ conflicts.\n\n## DNS\n\nSplit-horizon DNS for `*.hubris.network` served by Technitium on [CT 107](../containers/107-dns.md)\
|
||
\ at `192.168.8.2:53`. See [dns.md](dns.md) for full detail.\n\n## Routing\n\nProxmox has `net.ipv4.ip_forward=1` (already\
|
||
\ enabled by PVE). Packets from LXCs on `vmbr0` destined for the internet exit via `vmbr1` → Fritz!Box. Fritz!Box masquerades\
|
||
\ all outbound WAN traffic. Fritz!Box has a static route (`192.168.8.0/24 → 192.168.178.10`) so return traffic reaches\
|
||
\ the LXCs.\n\nNo NAT on Proxmox — traffic flows without double-NAT.\n\n## Remote access\n\n- **NetBird mesh** — primary\
|
||
\ path for remote administration. Authenticated via [Authentik on the VPS](../../../vps/).\n- **Tailscale** — legacy,\
|
||
\ being phased out. See [mesh.md](mesh.md).\n\n## Related\n\n- [DNS](dns.md) — split-horizon config and entry list\n-\
|
||
\ [Ingress](ingress.md) — public entry points via VPS traefik\n- [Mesh](mesh.md) — NetBird / Tailscale VPN overlay\n-\
|
||
\ [hosts/hubris.md](../hosts/hubris.md) — Proxmox host (vmbr0/vmbr1 config)\n- [CT 107 — dns](../containers/107-dns.md)\
|
||
\ — Technitium DNS + DHCP server\n\n## Changelog\n\n### 2026-06-17 — Fritz!Box DNSv4 server set to Technitium (192.168.8.2)\n\
|
||
Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs. Configured in Fritz!Box at Internet →\
|
||
\ Filter → DNS Server → DNSv4 Server → \"Use other DNSv4 servers\" → Preferred = `192.168.8.2`. No per-device or Netbird\
|
||
\ setup needed.\nPrevious pool `.100–.240` overlapped with all static LXCs/VMs (` .101–.239`), creating IP conflict risk\
|
||
\ (DHCP could hand out an IP that a static service expects). Shrunk pool to `.241–.254` via Technitium API. No services\
|
||
\ re-IP'd. 11 stale DHCP leases in `.101–.110` will expire naturally. **Open:** ZimaOS (VM 100) holds DHCP lease `.103`\
|
||
\ but inventory expects `.195` — needs static IP set inside VM. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).\n\
|
||
\n### 2026-06-02 — Executed migration; Proxmox as subnet router\nFritz!OS 8.x does not support second IP networks on LAN\
|
||
\ ports, so the final design uses Proxmox as the router: `vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10`;\
|
||
\ `vmbr0` is a portless internal bridge with `192.168.8.1` alias as the LXC gateway. Technitium DHCP enabled for `192.168.8.100–240`.\
|
||
\ Caddy service unit was missing and recreated. See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).\n\
|
||
\n### 2026-06-01 — Initial network doc; Slate AX retired; SODOLA switch added\nReplaced the GL.iNet Slate AX sub-router\
|
||
\ with the SODOLA 5-Port 2.5Gbit managed switch. Eliminated double-NAT. See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).\n"
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-06-17'
|
||
title: Fritz!Box DNSv4 server set to Technitium (192.168.8.2)
|
||
body: Household LAN clients (192.168.178.x) now resolve `*.hubris.network` to LAN IPs. Configured in Fritz!Box at Internet
|
||
→ Filter → DNS Server → DNSv4 Server → "Use other DNSv4 servers" → Preferred = `192.168.8.2`. No per-device or Netbird
|
||
setup needed. Previous pool `.100–.240` overlapped with all static LXCs/VMs (` .101–.239`), creating IP conflict risk
|
||
(DHCP could hand out an IP that a static service expects). Shrunk pool to `.241–.254` via Technitium API. No services
|
||
re-IP'd. 11 stale DHCP leases in `.101–.110` will expire naturally. **Open:** ZimaOS (VM 100) holds DHCP lease `.103`
|
||
but inventory expects `.195` — needs static IP set inside VM. See [plan](../../../.hermes/plans/2026-06-03_223218-dhcp-pool-exclude-static-ips.md).
|
||
- date: '2026-06-02'
|
||
title: Executed migration; Proxmox as subnet router
|
||
body: 'Fritz!OS 8.x does not support second IP networks on LAN ports, so the final design uses Proxmox as the router:
|
||
`vmbr1` (eno1 → SODOLA → Fritz!Box) is the uplink at `192.168.178.10`; `vmbr0` is a portless internal bridge with `192.168.8.1`
|
||
alias as the LXC gateway. Technitium DHCP enabled for `192.168.8.100–240`. Caddy service unit was missing and recreated.
|
||
See [migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).'
|
||
- date: '2026-06-01'
|
||
title: Initial network doc; Slate AX retired; SODOLA switch added
|
||
body: Replaced the GL.iNet Slate AX sub-router with the SODOLA 5-Port 2.5Gbit managed switch. Eliminated double-NAT. See
|
||
[migration plan](../../../plans/done/2026-06-01-slate-ax-to-sodola-migration.md).
|
||
- slug: infrastructure/ssh-access
|
||
title: SSH access
|
||
content: "# SSH access\n\nHow to reach every host in the fleet from any workstation, with LAN as\nthe primary path and Netbird\
|
||
\ as the automatic backup.\n\n## Architecture\n\nSSH access relies on three layers:\n\n1. **Homelab inventory (`inventory.yaml`)**\
|
||
\ — the single source of truth\n for every host's LAN IP, Netbird addresses, SSH user, and port.\n2. **Key distribution\
|
||
\ (`ssh/deploy-keys.sh`)** — deploys workstation SSH\n public keys to hubris and every running LXC, so any key-authorized\n\
|
||
\ workstation can log in anywhere.\n3. **Config generation (`homelab ssh-config --install`)** — generates\n `~/.ssh/config.d/homelab`\
|
||
\ with short hostname aliases for every host,\n using LAN IPs (routed via Netbird's `192.168.8.0/24` subnet route when\n\
|
||
\ off-LAN) with Netbird FQDN fallbacks (`<name>-mesh`) for roaming\n workstations.\n\n### How it works\n\n- **From\
|
||
\ on-LAN:** `ssh gitea` resolves to `192.168.8.121` directly.\n- **From off-LAN (Netbird):** The same `192.168.8.121`\
|
||
\ works because\n hubris routes the `192.168.8.0/24` subnet through Netbird.\n- **Roaming workstations:** `ssh mac-mini-mesh`\
|
||
\ or `ssh republic-laptop-mesh`\n uses the Netbird FQDN as a fallback when the workstation is off its\n home subnet.\n\
|
||
\nThe `homelab ssh <host>` CLI command also has built-in LAN probing:\nit tries a 1.5s TCP connect to the LAN IP, and\
|
||
\ if that fails, falls\nback to the Netbird FQDN.\n\n## Key distribution\n\nEach workstation's SSH public key lives in\
|
||
\ the repo at:\n`ssh/authorized_keys/<hostname>.pub`\n\nTo deploy or re-deploy all workstation keys to hubris + every\
|
||
\ running LXC:\n\n```bash\n# From hubris (or via homelab pct):\nsudo bash /opt/homelab-context/ssh/deploy-keys.sh\n\n\
|
||
# Or from any workstation:\nssh root@192.168.8.77 \"bash /opt/homelab-context/ssh/deploy-keys.sh\"\n```\n\nThis script:\n\
|
||
- Reads all `.pub` files from `ssh/authorized_keys/`\n- Adds any missing keys to `/etc/pve/priv/authorized_keys` on hubris\n\
|
||
- For each running LXC, appends keys to `/root/.ssh/authorized_keys`\n- Is idempotent — skips keys already present\n\n\
|
||
## Config generation\n\nTo generate the SSH config on any workstation:\n\n```bash\nhomelab ssh-config --install\n```\n\
|
||
\nThis writes to `~/.ssh/config.d/homelab` and ensures\n`Include ~/.ssh/config.d/homelab` is present in `~/.ssh/config`.\n\
|
||
\nThe config is regenerated automatically on every `homelab sync` (which\nkicks the 5-minute context sync timer).\n\n\
|
||
## Adding a new workstation\n\nWhen onboarding a new machine:\n\n1. Hostname must match an entry in `inventory.yaml`.\n\
|
||
2. If the workstation will be on the LAN, add its `lan_ip` to\n `inventory.yaml` and push. This gives it a primary LAN\
|
||
\ entry in the\n generated SSH config.\n3. Enable SSH Remote Login:\n - **macOS:** `sudo launchctl load -w /System/Library/LaunchDaemons/ssh.plist`\n\
|
||
\ - **Linux:** `sudo systemctl enable --now sshd`\n4. Generate an SSH keypair if one doesn't exist:\n ```bash\n \
|
||
\ ssh-keygen -t ed25519 -a 100\n ```\n5. Publish the public key to the repo:\n ```bash\n cp ~/.ssh/id_ed25519.pub\
|
||
\ /opt/homelab-context/ssh/authorized_keys/<hostname>.pub\n cd /opt/homelab-context && git add ssh/authorized_keys/\
|
||
\ && git commit -m 'ssh: add <hostname> pubkey' && git push\n ```\n6. Deploy the key to all hosts:\n ```bash\n ssh\
|
||
\ root@192.168.8.77 \"cd /opt/homelab-context && git pull --ff-only && bash ssh/deploy-keys.sh\"\n ```\n7. Generate\
|
||
\ the local SSH config:\n ```bash\n homelab ssh-config --install\n ```\n\n## Hosts\n\n### Hubris + strong (PVE cluster:\
|
||
\ `Homelab`)\n\nBoth nodes share `/etc/pve/priv/authorized_keys` — it's Proxmox\ncluster-synced, so a key added on either\
|
||
\ node is authorized on both.\n\n| Detail | hubris | strong |\n|--------|--------|-----------|\n| LAN IP | `192.168.8.77`\
|
||
\ | `192.168.178.181` |\n| Cluster node name | `hubris` | `strong` (OS hostname kept as-is from install) |\n| Netbird\
|
||
\ | `100.122.38.109` (`proxmox-server.netbird.selfhosted`) | not enrolled yet |\n| Netbird SSH port | `22022` (mesh-only,\
|
||
\ OIDC auth) | n/a |\n| SSH user | `root` | `root` |\n\nAuthorized root keys currently deployed (cluster-wide):\n- `root@hubris`\
|
||
\ (self, RSA)\n- `d.toro.v@pm.me` (ed25519) — mac-mini\n- `root@strong` (RSA) — strong's own key, added 2026-07-01 for\
|
||
\ the cluster join\n\n### LXCs\n\nEvery LXC at `192.168.8.x` accepts root SSH via authorized_keys. Keys\nare managed by\
|
||
\ `ssh/deploy-keys.sh`. SSH user is `root`.\n\n| LXC | Name | LAN IP | Role |\n|-----|------|--------|------|\n| 101 |\
|
||
\ jellyfin | `192.168.8.206` | media-server |\n| 102 | nfs-export | `192.168.8.200` | storage-export |\n| 103 | paperless\
|
||
\ | `192.168.8.130` | document-archive |\n| 104 | gitea | `192.168.8.121` | git-server |\n| 105 | apps | `192.168.8.205`\
|
||
\ | docker-apps |\n| 106 | auth-outpost | `192.168.8.184` | authentik-outpost |\n| 107 | dns | `192.168.8.185` | dns-helper\
|
||
\ |\n| 114 | nextcloud | `192.168.8.224` | file-sync |\n| 118 | elementsynapse | `192.168.8.239` | matrix-server |\n|\
|
||
\ 119 | sophia | `192.168.8.157` | workshop |\n| 120 | mule-images | `192.168.8.136` | photo-management |\n| 121 | caddy\
|
||
\ | `192.168.8.175` | reverse-proxy |\n| 122 | arriman | `192.168.8.132` | arr-stack |\n\n### Workstations\n\n| Name |\
|
||
\ OS | LAN IP | Netbird FQDN | SSH user |\n|------|----|--------|--------------|----------|\n| mac-mini | macOS | `192.168.8.174`\
|
||
\ | `mac-mini-234-17.netbird.selfhosted` | `dtoro` |\n| republic-laptop | Linux | TBD | `republic-laptop.netbird.selfhosted`\
|
||
\ | `dtoro` |\n\nstrong moved out of this table 2026-07-01 — it's a Proxmox host now, see the cluster table above.\n\n\
|
||
### VPS (external)\n\n| Detail | Value |\n|--------|-------|\n| Public IP | `82.165.190.79` |\n| Netbird | `100.122.165.149`\
|
||
\ (FQDN: `netbird-ionos.netbird.selfhosted`) |\n| SSH user | `root` |\n| Access | Mesh-only — public port 22 is blocked\
|
||
\ by nftables. Key-only auth. |\n\n## VPS\n\nAccess is mesh-only. From a mesh-connected peer:\n\n```bash\nssh root@100.122.165.149\n\
|
||
ssh root@netbird-ionos.netbird.selfhosted\n# or via homelab:\nhomelab ssh netbird-vps\n```\n\n## Verification\n\n```bash\n\
|
||
# From any workstation after running homelab ssh-config --install:\nfor name in hubris gitea apps sophia paperless caddy\
|
||
\ jellyfin nextcloud; do\n ssh -o BatchMode=yes \"$name\" \"hostname\" && echo \"$name OK\"\ndone\n```\n\n## Related\n\
|
||
\n- [Mesh migration](mesh.md)\n- [VPS hardening](vps-hardening.md)\n- [Agent enrollment](../../../.agents/operations/agent-enrollment.md)\n\
|
||
- [Homelab CLI](../../../bin/homelab)\n\n## Changelog\n\n### 2026-07-01 — strong reformatted to Proxmox, joined cluster;\
|
||
\ table corrected\nstrong moved from the Workstations table to the PVE-cluster table (was showing a stale `192.168.8.133`,\
|
||
\ never actually reachable — the real LAN IP has always been `192.168.178.181`, matching hosts/strong.yaml). Root key\
|
||
\ access bootstrapped via one-time console password, then key-only going forward. See [hosts/hubris.md#cluster](../hosts/hubris.md#cluster)\
|
||
\ and [hosts/strong.md](../hosts/strong.md).\n\n### 2026-06-02 — universal SSH reachability\n\nReplaced ad-hoc per-workstation\
|
||
\ SSH configs with inventory-generated\nconfigs (`ssh/gen-config.py`, `homelab ssh-config`). Added centralized\nkey distribution\
|
||
\ (`ssh/deploy-keys.sh`, `ssh/authorized_keys/`). All\nLXCs now accept root SSH from any workstation whose pubkey is in\
|
||
\ the\nrepo. mac-mini Remote Login enabled. Netbird subnet route\n(192.168.8.0/24 via hubris) provides off-LAN reachability\
|
||
\ for all LAN\nIPs.\n\n### 2026-04-28 — wiki entry created\nInitial documentation.\n\n### 2026-04-23 — VPS SSH hardened\
|
||
\ to mesh-only\nPublic `:22` blocked at nftables. Key-only sshd.\n\n### 2026-04-22 — iMac key authorized on hubris\n`d.toro.v@pm.me`\
|
||
\ added to `/etc/pve/priv/authorized_keys`."
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog:
|
||
- date: '2026-07-01'
|
||
title: strong reformatted to Proxmox, joined cluster; table corrected
|
||
body: strong moved from the Workstations table to the PVE-cluster table (was showing a stale `192.168.8.133`, never actually
|
||
reachable — the real LAN IP has always been `192.168.178.181`, matching hosts/strong.yaml). Root key access bootstrapped
|
||
via one-time console password, then key-only going forward. See [hosts/hubris.md#cluster](../hosts/hubris.md#cluster)
|
||
and [hosts/strong.md](../hosts/strong.md).
|
||
- date: '2026-06-02'
|
||
title: universal SSH reachability
|
||
body: Replaced ad-hoc per-workstation SSH configs with inventory-generated configs (`ssh/gen-config.py`, `homelab ssh-config`).
|
||
Added centralized key distribution (`ssh/deploy-keys.sh`, `ssh/authorized_keys/`). All LXCs now accept root SSH from
|
||
any workstation whose pubkey is in the repo. mac-mini Remote Login enabled. Netbird subnet route (192.168.8.0/24 via
|
||
hubris) provides off-LAN reachability for all LAN IPs.
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-23'
|
||
title: VPS SSH hardened to mesh-only
|
||
body: Public `:22` blocked at nftables. Key-only sshd.
|
||
- date: '2026-04-22'
|
||
title: iMac key authorized on hubris
|
||
body: '`d.toro.v@pm.me` added to `/etc/pve/priv/authorized_keys`.'
|
||
- slug: infrastructure/topology
|
||
title: Topology (generated)
|
||
content: "<!-- Generated by oikos/gen-topology.py from inventory.yaml. -->\n<!-- Do NOT edit by hand - your changes will\
|
||
\ be overwritten. -->\n\n# Topology (generated)\n\nSource: [inventory.yaml](../../../inventory.yaml) — 2 hypervisors,\
|
||
\ 20 LXCs, 2 VMs, 2 workstations, 18 services.\nEdge semantics: [oikos/ontology.yaml](../../../oikos/ontology.yaml). Operating\
|
||
\ model: [OIKOS.md](../../../.agents/OIKOS.md).\n\n## Compute & ingress\n\n```mermaid\nflowchart LR\n subgraph hubris_sub[\"\
|
||
hubris (Proxmox)\"]\n trmnl[\"trmnl<br/>LXC 128<br/>trmnl-middleware<br/>192.168.8.211\"]\n nfs_export[\"\
|
||
nfs-export<br/>LXC 102<br/>storage-export<br/>192.168.8.200\"]\n paperless[\"paperless<br/>LXC 103<br/>document-archive<br/>192.168.8.130\"\
|
||
]\n gitea[\"gitea<br/>LXC 104<br/>git-server<br/>192.168.8.121\"]\n apps[\"apps<br/>LXC 105<br/>docker-apps<br/>192.168.8.205\"\
|
||
]\n auth_outpost[\"auth-outpost<br/>LXC 106<br/>authentik-gateway<br/>192.168.8.6\"]\n dns[\"dns<br/>LXC\
|
||
\ 107<br/>dns-server<br/>192.168.8.2\"]\n nextcloud[\"nextcloud<br/>LXC 114<br/>file-sync<br/>192.168.8.224\"]\n\
|
||
\ sophia[\"sophia<br/>LXC 119<br/>workshop<br/>192.168.8.109\"]\n mule_images[\"mule-images<br/>LXC 120<br/>photo-management<br/>192.168.8.136\"\
|
||
]\n caddy[\"caddy<br/>LXC 121<br/>reverse-proxy<br/>192.168.8.175\"]\n teddycloud[\"teddycloud<br/>LXC 131<br/>teddycloud<br/>192.168.8.150\"\
|
||
]\n zimaos[\"zimaos<br/>VM 100<br/>nas-frontend-eval<br/>192.168.8.195\"]\n haos[\"haos<br/>VM 108<br/>home-automation<br/>192.168.8.101\"\
|
||
]\n end\n subgraph strong_sub[\"strong (Proxmox)\"]\n house[\"house<br/>LXC 129<br/>family-planner<br/>192.168.8.244\"\
|
||
]\n jellyfin[\"jellyfin<br/>LXC 101<br/>media-server<br/>192.168.8.246\"]\n elementsynapse[\"elementsynapse<br/>LXC\
|
||
\ 118<br/>matrix-server<br/>192.168.8.242\"]\n arriman[\"arriman<br/>LXC 122<br/>arr-stack<br/>192.168.8.245\"\
|
||
]\n grimmory[\"grimmory<br/>LXC 130<br/>book-library<br/>192.168.8.247\"]\n seanime[\"seanime<br/>LXC 133<br/>anime-media-server<br/>192.168.8.248\"\
|
||
]\n romm[\"romm<br/>LXC 134<br/>rom-manager<br/>192.168.8.249\"]\n end\n rclone[\"rclone<br/>lxc<br/>backup\"\
|
||
]\n republic_laptop([republic-laptop<br/>workstation<br/>primary-dev])\n mac_mini([mac-mini<br/>workstation<br/>dev<br/>192.168.178.182])\n\
|
||
\ netbird_vps[[netbird-vps<br/>external<br/>netbird-mgmt]]\n url_artifacto([\"artifacto.hubris.network\"]) -->|routes-to|\
|
||
\ apps\n url_authentik([\"auth.hubris.network\"]) -->|routes-to| netbird_vps\n url_gitea([\"git.hubris.network\"\
|
||
]) -->|routes-to| gitea\n url_homelab_mcp([\"mcp.hubris.network\"]) -->|routes-to| apps\n url_jellyfin([\"media.hubris.network\"\
|
||
]) -->|routes-to| jellyfin\n url_matrix([\"matrix.hubris.network\"]) -->|routes-to| elementsynapse\n url_nextcloud([\"\
|
||
cloud.hubris.network\"]) -->|routes-to| nextcloud\n url_paperless([\"paperless.hubris.network\"]) -->|routes-to| paperless\n\
|
||
\ url_photos([\"photos.hubris.network\"]) -->|routes-to| mule_images\n url_proxmox_ui([\"proxmox.hubris.network\"\
|
||
]) -->|routes-to| hubris_sub\n url_secrets_issuance([\"secrets.hubris.network\"]) -->|routes-to| apps\n url_teddycloud([\"\
|
||
teddy.hubris.network\"]) -->|routes-to| teddycloud\n url_trmnl([\"trmnl.hubris.network\"]) -->|routes-to| trmnl\n \
|
||
\ url_zimaos([\"zimaos.hubris.network\"]) -->|routes-to| zimaos\n```\n\n## Storage (mounts)\n\n```mermaid\nflowchart\
|
||
\ LR\n mnt_library[(\"/mnt/library\")]\n mnt_media_local[(\"/mnt/media_local\")]\n mnt_media_local_anime[(\"\
|
||
/mnt/media_local/anime\")]\n apps[\"apps\"] -->|mounts| mnt_library\n arriman[\"arriman\"] -->|mounts| mnt_media_local\n\
|
||
\ gitea[\"gitea\"] -->|mounts| mnt_library\n grimmory[\"grimmory\"] -->|mounts| mnt_media_local\n hubris[\"hubris\"\
|
||
] -->|mounts| mnt_library\n jellyfin[\"jellyfin\"] -->|mounts| mnt_media_local\n mule_images[\"mule-images\"] -->|mounts|\
|
||
\ mnt_library\n nextcloud[\"nextcloud\"] -->|mounts| mnt_library\n paperless[\"paperless\"] -->|mounts| mnt_library\n\
|
||
\ romm[\"romm\"] -->|mounts| mnt_media_local\n seanime[\"seanime\"] -->|mounts| mnt_media_local_anime\n sophia[\"\
|
||
sophia\"] -->|mounts| mnt_library\n teddycloud[\"teddycloud\"] -->|mounts| mnt_library\n```\n\n## Archaeology (destroyed\
|
||
\ nodes)\n\n| Node | ID | Destroyed | Reason |\n|---|---|---|---|\n| plato | 126 | 2026-06-28 | notes workspace decommissioned;\
|
||
\ data retained at /mnt/library/documents/plato |\n| claudio-bot | 123 | 2026-06-04 | replaced by Hermes Agent on mac-mini;\
|
||
\ monitoring moved to homelab-health-watchdog cron |\n| mule-photos-new | 127 | 2026-05-22 | PhotoPrism test stack promoted\
|
||
\ to LXC 120 (Mulimage 2.0 merge) |\n| heaper | 116 | 2026-05-14 | decommissioned; data retained at /mnt/library/heaper\
|
||
\ |\n| syncthing | 109 | 2026-05-14 | decommissioned; library subtree was empty |\n| seafile | 125 | 2026-05-13 | Seafile\
|
||
\ Pro evaluation rejected; files.hubris.network removed from caddy + dns |\n| arr-yunohost | 100 | 2026-04-28 | migrated\
|
||
\ to docker stack on arriman (LXC 122) |\n| flaresolverr | 106 | 2026-04-28 | folded into the arriman docker compose |\n\
|
||
| marimo | 107 | 2026-04-28 | decommissioned |\n| photoprism | 110 | 2026-04-28 | replaced by mule-images (LXC 120) |\n\
|
||
| karakeep | 111 | 2026-04-28 | decommissioned |\n| immich | 112 | 2026-04-28 | replaced by mule-images (LXC 120) |\n\
|
||
| reticulum | 115 | 2026-04-28 | decommissioned |\n"
|
||
entity_slug: null
|
||
tags:
|
||
- infrastructure
|
||
at_glance: {}
|
||
changelog: []
|
||
- slug: infrastructure/vps-hardening
|
||
title: VPS hardening — `82.165.190.79` / `100.122.165.149`
|
||
content: '# VPS hardening — `82.165.190.79` / `100.122.165.149`
|
||
|
||
|
||
IONOS VPS that runs the Netbird control plane and the [public ingress traefik](ingress.md). Hardened 2026-04-23 from its
|
||
stock-Plesk state.
|
||
|
||
|
||
## At a glance
|
||
|
||
- **Hostname:** `inspiring-ramanujan.82-165-190-79.plesk.page`
|
||
|
||
- **OS:** Debian 13
|
||
|
||
- **Mesh:** netbird `100.122.165.149` (peer of the lab mesh; routes `192.168.8.0/24` via [hubris](../hosts/hubris.md)).
|
||
|
||
- **Public:** `82.165.190.79` (`ens6`).
|
||
|
||
- **Public DNS:** IONOS wildcard `*.hubris.network → 82.165.190.79`.
|
||
|
||
- **Docker stack** at `/opt/docker-compose.yml`: `traefik` (TLS/ACME) + `dashboard` + `mgmt` + `signal` + `relay` + `proxy`
|
||
— netbird-mgmt 0.71.3 vanilla deploy since 2026-05-21 (see [mesh.md changelog](mesh.md#changelog)).
|
||
|
||
- **Host services (outside docker):** `coturn` (TURN-TCP on :3478, long-term creds rendered into `/etc/turnserver.conf`
|
||
by `homelab render-vps-configs` from sops-encrypted `secrets/turn-shared-secret.yaml`).
|
||
|
||
- **Config rendering:** `/etc/turnserver.conf` + `/opt/management.json` are generated from templates in `vps/*.tmpl` on
|
||
this repo by `homelab render-vps-configs`. Secret placeholders (`{{TURN_PASSWORD}}`, `{{AUTHENTIK_CLIENT_SECRET}}`) are
|
||
substituted from sops-encrypted secrets decrypted on hubris and pushed over ssh. **Do not hand-edit those two files on
|
||
the VPS** — the next render will overwrite them.
|
||
|
||
|
||
## SSH
|
||
|
||
|
||
- Key-only (`PasswordAuthentication no`, `PermitRootLogin prohibit-password`) via drop-in at `/etc/ssh/sshd_config.d/10-hubris-hardening.conf`.
|
||
Original config backed up at `/etc/ssh/sshd_config.bak.<ts>`.
|
||
|
||
- **Mesh-only**: public `:22` is dropped by the nftables firewall. SSH reaches the VPS only over `wt0`. `ListenAddress`
|
||
itself is still `0.0.0.0` — gating is firewall-layer.
|
||
|
||
- Authorized root keys: PVE (`root@hubris`), Mac Mini (`d.toro.v@pm.me`). Add a new device with `ssh-copy-id root@100.122.165.149`
|
||
from a mesh peer **before** disabling its access paths.
|
||
|
||
|
||
## Firewall — nftables (`inet hubris-fw`)
|
||
|
||
|
||
Config at `/etc/nftables.conf`, service enabled.
|
||
|
||
|
||
- Public iface `ens6`. Wireguard iface `wt0`.
|
||
|
||
- **INPUT on `ens6`** allow-list: DHCP (67→68), rate-limited ICMP/ICMPv6, **TCP 3478** (coturn TURN-TCP, added 2026-05-21).
|
||
Everything else drops.
|
||
|
||
- `wt0` fully accepted in INPUT. `lo` accepted.
|
||
|
||
- **IONOS upstream firewall** also gates inbound traffic before it reaches `ens6`. Open ports today: TCP 80/443 (traefik),
|
||
UDP 51820 (netbird-proxy), TCP 3478 (coturn, added 2026-05-21). UDP 3478 is dropped by IONOS upstream regardless of local
|
||
nftables. See [mesh.md ICE/STUN](mesh.md) for the STUN/TURN port matrix.
|
||
|
||
- **FORWARD chain at priority `filter-10`** (runs before Docker''s FORWARD) hosts the fail2ban ban enforcement — see below.
|
||
|
||
- Set `banned4` (typed `ipv4_addr`, flag `timeout`) holds fail2ban''s drops.
|
||
|
||
- Coexists with Docker''s `ip nat` / `ip filter` tables (iptables-nft compat). **Do NOT `flush ruleset`** in this config
|
||
— it''ll wipe Docker''s state too.
|
||
|
||
|
||
## fail2ban
|
||
|
||
|
||
- **Jail `traefik-4xx`** tails `/var/log/traefik/access.log` (bind-mounted from container). Filter at `/etc/fail2ban/filter.d/traefik-4xx.conf`
|
||
matches 401/403/404/429 from `blog-public@file` or `artifacto-public@file` routers only — netbird-grpc traffic isn''t
|
||
considered.
|
||
|
||
- Tunables: `findtime=600, maxretry=30, bantime=3600`.
|
||
|
||
- **Action** at `/etc/fail2ban/action.d/nft-hubris.conf` adds/removes elements from `inet hubris-fw banned4` with per-element
|
||
timeout.
|
||
|
||
|
||
### CRITICAL invariant — wireguard / fail2ban
|
||
|
||
|
||
**Bans must never affect `wt0` or wireguard UDP.** The FORWARD chain explicitly `accept`s the following *before* the ban
|
||
check:
|
||
|
||
- `udp 51820` (wireguard)
|
||
|
||
- `udp 3478` (STUN)
|
||
|
||
- `ct state established,related`
|
||
|
||
|
||
The INPUT ban rule is scoped to `iifname "ens6"`.
|
||
|
||
|
||
Violating this takes the mesh down for every home device (they share one public IP) and the only recovery is IONOS console
|
||
→ `nft flush set inet hubris-fw banned4`.
|
||
|
||
|
||
## Traefik access log
|
||
|
||
|
||
- Written to `/var/log/traefik/access.log` on the host via a bind mount added to `/opt/docker-compose.yml` (traefik volumes
|
||
include `/var/log/traefik:/logs`) plus `--accesslog.filepath=/logs/access.log`.
|
||
|
||
- CLF format. Real client IP arrives correctly because docker userland-proxy is off — see [public ingress](ingress.md).
|
||
|
||
|
||
## Plesk / mail / FTP / Dr.Web
|
||
|
||
|
||
Stopped and disabled (not uninstalled). All of:
|
||
|
||
`dovecot`, `dovecot.socket`, `postfix`, `postfix@-`, `pc-remote`, `xinetd`, `plesk-task-manager`, `plesk-web-socket`,
|
||
`sw-cp-server`, `sw-engine`, `plesk-repaird`, `plesk-repaird.socket`, `drwebd`.
|
||
|
||
|
||
`psa.service` is masked (was a one-shot boot bootstrap). `/etc/cron.d/plesk-backup-manager-task` renamed to `.disabled`.
|
||
|
||
|
||
Reverse: `systemctl unmask psa; systemctl enable --now <svc>`.
|
||
|
||
|
||
## Auto-patching
|
||
|
||
|
||
- `unattended-upgrades` enabled (stock).
|
||
|
||
- Drop-in at `/etc/apt/apt.conf.d/52hubris-reboot.conf` sets auto-reboot at **04:00 UTC** when `/var/run/reboot-required`
|
||
is set.
|
||
|
||
- Runs inside the stock `apt-daily-upgrade.timer`.
|
||
|
||
|
||
## Recovery paths
|
||
|
||
|
||
Ordered by preference:
|
||
|
||
1. **SSH via mesh** — primary. Any mesh peer with an authorized key.
|
||
|
||
2. **IONOS web console** (my.ionos.com → VPS → Console) — uses the system password, not SSH keys. Bypasses any firewall
|
||
misconfig.
|
||
|
||
3. **IONOS rescue mode** — boot rescue, mount rootfs, edit `/etc/nftables.conf` or `/etc/ssh/sshd_config.d/10-hubris-hardening.conf`
|
||
to a known-good state, reboot.
|
||
|
||
|
||
## Related
|
||
|
||
- [Public ingress (VPS traefik)](ingress.md)
|
||
|
||
- [Mesh migration](mesh.md) — VPS as a mesh peer
|
||
|
||
- [SSH access](ssh-access.md)
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-05-21 — netbird stack migrated combined → vanilla; coturn added
|
||
|
||
Replaced the `netbirdio/netbird-server` combined image with the canonical `mgmt + signal + relay + dashboard` containers
|
||
(0.71.3). Added host-side `coturn` for external TURN, with nftables rule `iifname "ens6" tcp dport 3478 accept` and an
|
||
IONOS upstream firewall exception. Authentik on LXC 124 now provides OIDC for the netbird dashboard. Full context in [mesh.md
|
||
changelog](mesh.md#changelog).
|
||
|
||
|
||
### 2026-04-28 — wiki entry created
|
||
|
||
Initial documentation.
|
||
|
||
|
||
### 2026-04-23 — hardened
|
||
|
||
nftables firewall, mesh-only SSH, fail2ban traefik jail, Plesk disabled, auto-reboot 04:00 UTC, wireguard/fail2ban invariant
|
||
established.
|
||
|
||
'
|
||
entity_slug: host:netbird-vps
|
||
tags:
|
||
- infrastructure
|
||
at_glance:
|
||
hostname: inspiring-ramanujan.82-165-190-79.plesk.page
|
||
os: Debian 13
|
||
mesh: netbird 100.122.165.149 (peer of the lab mesh; routes 192.168.8.0/24 via hubris).
|
||
public: 82.165.190.79 (ens6).
|
||
public_dns: IONOS wildcard *.hubris.network → 82.165.190.79.
|
||
docker_stack: 'at /opt/docker-compose.yml: traefik (TLS/ACME) + dashboard + mgmt + signal + relay + proxy — netbird-mgmt
|
||
0.71.3 vanilla deploy since 2026-05-21 (see mesh.md changelog).'
|
||
host_services_(outside_docker): coturn (TURN-TCP on :3478, long-term creds rendered into /etc/turnserver.conf by homelab
|
||
render-vps-configs from sops-encrypted secrets/turn-shared-secret.yaml).
|
||
config_rendering: /etc/turnserver.conf + /opt/management.json are generated from templates in vps/*.tmpl on this repo
|
||
by homelab render-vps-configs. Secret placeholders ({{TURN_PASSWORD}}, {{AUTHENTIK_CLIENT_SECRET}}) are substituted
|
||
from sops-encrypted secrets decrypted on hubris and pushed over ssh. Do not hand-edit those two files on the VPS — the
|
||
next render will overwrite them.
|
||
changelog:
|
||
- date: '2026-05-21'
|
||
title: netbird stack migrated combined → vanilla; coturn added
|
||
body: Replaced the `netbirdio/netbird-server` combined image with the canonical `mgmt + signal + relay + dashboard` containers
|
||
(0.71.3). Added host-side `coturn` for external TURN, with nftables rule `iifname "ens6" tcp dport 3478 accept` and
|
||
an IONOS upstream firewall exception. Authentik on LXC 124 now provides OIDC for the netbird dashboard. Full context
|
||
in [mesh.md changelog](mesh.md#changelog).
|
||
- date: '2026-04-28'
|
||
title: wiki entry created
|
||
body: Initial documentation.
|
||
- date: '2026-04-23'
|
||
title: hardened
|
||
body: nftables firewall, mesh-only SSH, fail2ban traefik jail, Plesk disabled, auto-reboot 04:00 UTC, wireguard/fail2ban
|
||
invariant established.
|
||
investigations:
|
||
- slug: investigations/2026-04-21-hubris-crash-loop
|
||
title: 2026-04-21 — Hubris crash loop (thermal + USB drive)
|
||
date: '2026-04-21'
|
||
status: resolved
|
||
duration: ''
|
||
content: '# 2026-04-21 — Hubris crash loop (thermal + USB drive)
|
||
|
||
|
||
## Summary
|
||
|
||
|
||
[`hubris`](../../../wiki/hosts/hubris.md) hard-locked repeatedly on 2026-04-21 (silent CPU hangs, no panic, no OOM, no
|
||
MCE). Two contributors identified: idle CPU sitting at ~95 °C on the `performance` governor, and a USB-attached external
|
||
SSD whose UAS interaction with the AMD USB4/Thunderbolt PCIe tunnel triggered hard locks. CPU thermal addressed via `cpu-epp.service`;
|
||
drive removed 2026-04-22 as an A/B test. As of 2026-04-28 the host has 3+ days uptime — the drive looks like the primary
|
||
contributor; `cpu-epp` remains as belt-and-suspenders.
|
||
|
||
|
||
## Timeline
|
||
|
||
|
||
### 2026-04-19 — drive attached
|
||
|
||
External `Silicon Motion Portable SSD` (vid:pid `090c:2320`) attached for the new restic [backup pipeline](../../../wiki/infrastructure/backups.md).
|
||
Pre-attach uptime had been 33 days stable.
|
||
|
||
|
||
### 2026-04-19 → 2026-04-21 — first crashes
|
||
|
||
Two hard crashes in 2.5 days (46 h then 12 h uptime). Kernel logs ended abruptly with routine apparmor entries — no panic,
|
||
OOM, or MCE — the classic hard-lock signature. Preceded by `uas_eh_abort_handler` storms and xHCI resets on port 6-1.
|
||
|
||
|
||
### 2026-04-21 — crash cadence escalates
|
||
|
||
23:02 (previous night) → 12:01 the next day. Cadence: 12 h → 43 m → 8 m → 5 m. Stable since the 12:03 boot once `cpu-epp.service`
|
||
started applying `EPP=balance_power`.
|
||
|
||
|
||
### 2026-04-21 — first set of mitigations applied
|
||
|
||
|
||
- **`cpu-epp.service`** deployed. Sets `scaling_governor=powersave` + EPP=`balance_power` at boot. Drops idle Tctl from
|
||
~95 °C → ~60 °C, clocks from 4.4 GHz pinned → 1.1–2 GHz idle.
|
||
|
||
- **Crash capture**: `/etc/sysctl.d/60-crash-capture.conf` (panic on oops/hardlockup/softlockup/rcu, auto-reboot 10 s),
|
||
`/etc/modprobe.d/softdog.conf` (`soft_panic=1 soft_margin=60`), `/etc/systemd/system.conf.d/watchdog.conf` (`RuntimeWatchdogSec=15s`).
|
||
|
||
- **`rasdaemon`** installed (Debian pkg; mcelog is retired) — logs MCE / memory / PCIe AER / thermal events to `/var/lib/rasdaemon/ras-mc_event.db`.
|
||
Query `ras-mc-ctl --summary`.
|
||
|
||
- **UAS blacklist** for the drive: `/etc/modprobe.d/usb-storage-quirks.conf` → `options usb-storage quirks=090c:2320:u`.
|
||
Forces BOT instead of UAS for the SMI bridge.
|
||
|
||
- **Mount-on-demand** for the drive: `/usr/local/sbin/backup-usb.sh attach|detach|status` toggles `/sys/bus/usb/devices/*/authorized`
|
||
so the drive is de-authorized when no backup is running.
|
||
|
||
|
||
### 2026-04-22 — recurrence after 30 h 37 m
|
||
|
||
Same silent-cutoff signature at 18:42:08. Much longer than any pre-`cpu-epp` crash (12 h max), so `cpu-epp` helps but
|
||
is not sufficient on its own. [claudio-monitor](../../../wiki/infrastructure/monitoring.md) showed healthy runtimes up
|
||
to 43 s before the hang (no pre-crash degradation). No MCE / no RAS / pstore empty.
|
||
|
||
|
||
### 2026-04-22 — `cpu-epp.service` design bug fixed
|
||
|
||
Was `After=multi-user.target` + `WantedBy=multi-user.target` — queued behind `pve-guests.service`. The hottest window
|
||
of every boot (20 LXCs + 1 VM coming up) ran on the `performance` governor. Fixed: now `After=sysinit.target` + `Before=pve-guests.service`.
|
||
|
||
|
||
### 2026-04-22 — drive removed (A/B test)
|
||
|
||
User physically removed the external USB drive. [Backup timers disabled](../../../wiki/infrastructure/backups.md#status),
|
||
fstab entry commented, drive de-authorized. Goal: confirm whether the drive + UAS + AMD USB4 PCIe-tunnel interaction is
|
||
the dominant root cause.
|
||
|
||
|
||
### 2026-04-23 — SSD cooling + thermal pads installed
|
||
|
||
Cold-boot baseline (3 min uptime): nvme0n1 35 °C composite / sensor1 (controller) **53 °C**; nvme1n1 36 °C composite /
|
||
both sensors ≤36 °C. Lifetime warning-time counters at install: nvme0n1 709 min warn + 5 min crit; nvme1n1 778 min warn
|
||
+ 45 min crit — both drives had spent real time in thermal warning historically.
|
||
|
||
|
||
#### 2026-04-23 thermal-pad verdict
|
||
|
||
Resolved (3 h self-paced watch). Steady-state across 17:03 → 19:56 stayed pinned at nvme0 composite 47 °C / sensor1 60–61
|
||
°C, nvme1 38–40 °C. Zero new warning-time minutes on either drive. Controller sensor1 at 60-ish is normal idle for M.2
|
||
NVMe. **The load-bearing signal is the lifetime warning-time counter, not absolute sensor1** — watch that going forward,
|
||
not sensor1 threshold.
|
||
|
||
|
||
### 2026-04-28 — A/B test passing so far
|
||
|
||
3+ days continuous uptime with the drive removed. If uptime now exceeds ~3 days without crash, the drive/PCIe-tunnel interaction
|
||
is confirmed as the dominant root cause. Will revisit options (different drive, different USB port, different bridge)
|
||
once stability is firmly established.
|
||
|
||
|
||
### 2026-04-29 — physical relocation
|
||
|
||
User moved the host to a spot with better airflow. Post-move idle baseline (45 min uptime, light load): Tctl **47.2 °C**,
|
||
GPU 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 34–35.5 °C. Notably
|
||
nvme0 sensor1 is **27 °C cooler** than the 2026-04-23 thermal-pad steady-state of 60–61 °C. Watch lifetime warning-time
|
||
counters (`nvme smart-log` → "Warning Comp. Temperature Time") over the next week — flat trend confirms the move helped.
|
||
Doesn''t change the drive-removal A/B test, but reduces residual thermal headroom risk independently.
|
||
|
||
|
||
## Root cause
|
||
|
||
|
||
Two-contributor:
|
||
|
||
1. **Thermal**: Ryzen 5 7640HS (Phoenix APU) in a passively-cooled GMKtec NucBox sitting at ~95 °C idle under `performance`
|
||
governor. Mitigated by `cpu-epp.service`. In `amd-pstate=active` mode the governor MUST be `powersave` for EPP values
|
||
to take effect — `performance` silently ignores them.
|
||
|
||
2. **Storage**: external USB SSD''s UAS keepalive interacts badly with the AMD USB4/Thunderbolt PCIe tunnel. UAS blacklist
|
||
+ mount-on-demand reduced but did not eliminate it. Removing the drive eliminates it (so far).
|
||
|
||
|
||
Background: BIOS 1.02 (latest publicly visible from GMKtec). Vendor not on LVFS so `fwupdmgr` can''t update. Non-ECC RAM
|
||
means memory faults are silent and undetectable except by memtest86+.
|
||
|
||
|
||
## Diagnostic notes
|
||
|
||
|
||
- **Pure hardware hangs leave no trace.** `pstore` stays empty across reboots despite `panic_on_oops/hardlockup/softlockup/rcu`
|
||
+ softdog `soft_panic=1` being configured. The CPU is gone before the kernel can react. Don''t waste time on panic traces
|
||
— use crash cadence + uptime pattern as the diagnostic.
|
||
|
||
- **First check** if hubris crash-loops again with the same silent-cutoff signature: `cpu-epp.service` is active and `/sys/devices/system/cpu/cpu0/cpufreq/energy_performance_preference`
|
||
is `balance_power`.
|
||
|
||
- **Non-ECC RAM means memory faults are silent.** If `cpu-epp` is on and crashes still happen, suspect DIMM next (run
|
||
memtest86+) and/or try the commented-out `pcie_aspm=off pci=nomsi` workarounds in `/etc/default/grub`.
|
||
|
||
- **Grub escalation** (`pcie_aspm=off pcie_port_pm=off pci=nomsi`) **deliberately NOT applied** yet to keep the drive-removal
|
||
test clean. Enable only if crashes continue without the drive.
|
||
|
||
- `mcelog` is not in Debian 13 repos; `rasdaemon` is the replacement.
|
||
|
||
|
||
## BIOS update path
|
||
|
||
|
||
Checked 2026-04-21. GMKtec is **not on LVFS**, so `fwupdmgr` can''t update the NucBox M6 Ultra. BIOS 1.02 (2025-08-06)
|
||
is the latest publicly visible — GMKtec''s download center only publishes Windows drivers, no BIOS listing or changelog.
|
||
Update path is a `.exe` flasher that needs Windows (would require Windows-To-Go USB on hubris). Before committing to that,
|
||
email `service@gmktec.com` with the S/N to confirm a newer BIOS / AGESA level actually exists. `fwupd` is now installed,
|
||
so if GMKtec ever onboards to LVFS, updates appear automatically.
|
||
|
||
|
||
## Mitigations applied
|
||
|
||
|
||
| Mitigation | Status | Notes |
|
||
|
||
| ----------------------------------------- | -------- | ---------------------------------------------------- |
|
||
|
||
| `cpu-epp.service` (powersave + balance_power) | Active | Unit ordering fixed 2026-04-22 to run before guests |
|
||
|
||
| Crash capture (sysctl + softdog + watchdog) | Active | Catches everything but pure silicon hangs |
|
||
|
||
| `rasdaemon` | Active | Persistent telemetry for MCE / memory / PCIe AER |
|
||
|
||
| UAS blacklist for `090c:2320` | Active (drive currently absent) | `/etc/modprobe.d/usb-storage-quirks.conf`
|
||
|
|
||
|
||
| Mount-on-demand (`backup-usb.sh`) | Active (drive currently absent) | Backup unit `ExecStartPre`/`ExecStopPost`
|
||
hooks |
|
||
|
||
| Drive physically removed | Active | A/B test in progress |
|
||
|
||
| Thermal pads on NVMes | Active 2026-04-23 | Steady-state composite 47 °C / sensor1 60 °C |
|
||
|
||
| `pcie_aspm=off pci=nomsi` | NOT applied | Reserved for if crashes recur without the drive |
|
||
|
||
|
||
## Affected nodes
|
||
|
||
- [Hubris host](../../../wiki/hosts/hubris.md)
|
||
|
||
- [Backups (disabled)](../../../wiki/infrastructure/backups.md)
|
||
|
||
- [Monitoring](../../../wiki/infrastructure/monitoring.md)
|
||
|
||
|
||
## Open questions
|
||
|
||
- Will the host stay up indefinitely without the drive? (Test ongoing — 3+ days as of 2026-04-28.)
|
||
|
||
- If yes: which mitigation lets the drive come back? (different drive? different USB port bypassing the hub chain? rear
|
||
motherboard USB 3 port? alternative bus path?)
|
||
|
||
- Newer BIOS / AGESA from GMKtec — answered email pending. Without LVFS we need explicit vendor confirmation.
|
||
|
||
- Memtest86+ pass on the DIMMs — not yet run; deferred until the drive case is closed.
|
||
|
||
'
|
||
about_slugs:
|
||
- host:hubris
|
||
tags:
|
||
- investigation
|
||
- slug: investigations/2026-05-31-authentik-vps-migration
|
||
title: 2026-05-31 — Authentik migrated from LXC 124 to the VPS
|
||
date: '2026-05-31'
|
||
status: resolved
|
||
duration: ''
|
||
content: '# 2026-05-31 — Authentik migrated from LXC 124 to the VPS
|
||
|
||
|
||
## Summary
|
||
|
||
|
||
The NetBird management server (on the [VPS](../../../wiki/infrastructure/ingress.md)) crash-looped 1200+ times because
|
||
it fetches the Authentik OIDC discovery document on startup, and Authentik was only reachable via the NetBird mesh — which
|
||
was down *because* mgmt couldn''t start. A classic bootstrap deadlock: **mgmt needs OIDC → OIDC needs the mesh → the mesh
|
||
needs mgmt.**
|
||
|
||
|
||
Resolved by moving Authentik off [LXC 124](../../../wiki/containers/106-auth-outpost.md) onto the VPS itself, so `auth.hubris.network`
|
||
resolves to a container co-located with netbird-mgmt — no mesh dependency. A `depends_on: condition: service_healthy`
|
||
on the mgmt service makes the deadlock structurally impossible to recur.
|
||
|
||
|
||
The full Authentik Postgres DB (all users, apps, passwords, groups) was migrated, so every gated app keeps working with
|
||
no per-app reconfiguration.
|
||
|
||
|
||
## Timeline
|
||
|
||
|
||
### Trigger
|
||
|
||
`netbird status` on clients showed `Management: Disconnected`; `netbird.hubris.network` unreachable. VPS `docker logs
|
||
netbird-mgmt`:
|
||
|
||
```
|
||
|
||
Error: failed reading provided config file: /etc/netbird/management.json:
|
||
|
||
OIDC configuration request returned status 504 with response: Gateway Timeout
|
||
|
||
```
|
||
|
||
The 504 came from VPS traefik trying to reach the old `auth-authentik` backend (`https://192.168.8.175`, home Caddy) over
|
||
the down mesh.
|
||
|
||
|
||
### Fix — Phase A (restore the mesh)
|
||
|
||
1. Stood up Authentik on the VPS in `/opt/docker-compose.yml`: `authentik-postgres`, `authentik-redis`, `authentik-server`,
|
||
`authentik-worker` (image `ghcr.io/goauthentik/server:2026.5.2`), on a dedicated `auth` Docker network (`172.30.1.0/24`).
|
||
Traefik joined both `netbird` and `auth` networks.
|
||
|
||
2. Removed the `auth-authentik` → `192.168.8.175` backend from `/opt/traefik-dynamic.yaml`; Authentik now routed via Docker
|
||
provider labels.
|
||
|
||
3. Added `depends_on: authentik-server: condition: service_healthy` to the mgmt service.
|
||
|
||
4. Created the NetBird OAuth2 provider/app manually (only manual step), then mgmt auto-recovered within 30 s.
|
||
|
||
|
||
### Fix — Phase B (full data migration)
|
||
|
||
Once the mesh was back, the homelab DB was reachable:
|
||
|
||
1. `ssh proxmox ''pct exec 124 -- docker exec authentik-postgresql-1 pg_dump -U authentik authentik''` → gzip.
|
||
|
||
2. Stopped VPS authentik server/worker, dropped + recreated the VPS DB (connect to the `postgres` DB to drop `authentik`),
|
||
restored the dump.
|
||
|
||
3. **Swapped `AUTHENTIK_SECRET_KEY` in `/opt/authentik.env` to the homelab''s key** — Authentik encrypts OAuth2 client
|
||
secrets in the DB with this key; a mismatch makes every client secret unreadable and silently breaks all SSO.
|
||
|
||
4. Restarted Authentik. All users / apps / passwords / groups present.
|
||
|
||
|
||
Scripts left on the VPS: `/opt/migrate-authentik.sh` (re-runnable), `/opt/backup-authentik.sh` (daily cron, 14-day retention).
|
||
|
||
|
||
### Fix — DNS cutover (the long tail)
|
||
|
||
The real reason the browser kept hitting the *old* Authentik even after the VPS one was healthy: split-horizon dnsmasq
|
||
on LXC 124 still had `address=/auth.hubris.network/192.168.8.175`, and the NetBird per-client resolver (`100.122.255.254`)
|
||
cached it. Symptom: the login page console logged `authentik(early): version 2026.2.2` (the old LXC) instead of `2026.5.2`
|
||
(the VPS).
|
||
|
||
|
||
- Updated dnsmasq: `address=/auth.hubris.network/82.165.190.79`, `systemctl restart dnsmasq`.
|
||
|
||
- `netbird down && netbird up` does **not** reliably flush the NetBird resolver cache. Immediate per-client fixes: `/etc/hosts`
|
||
override (mac), `sudo resolvectl flush-caches` (Linux), `sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder`
|
||
(mac).
|
||
|
||
|
||
## Issues hit during the migration (and fixes)
|
||
|
||
|
||
1. **Redirect URI error.** The restored DB had redirect URIs in `REGEX` matching mode; in Authentik 2026.5.x they failed
|
||
to match. Fixed by switching to `STRICT` exact matching (Django ORM, `RedirectURIMatchingMode.STRICT`). Set all four:
|
||
`http://localhost:53000/` (CLI), `https://netbird.hubris.network/{peers,nb-auth,nb-silent-auth}`.
|
||
|
||
2. **Only the password field showed (no username).** NetBird passes `login_hint=<email>` in the OAuth2 URL → Authentik
|
||
pre-identifies and skips the identification stage. Expected behavior; not a bug.
|
||
|
||
3. **"Request has been denied. Unknown error."** Several overlapping causes: wrong password (reset via Django shell),
|
||
reputation lockout after repeated failures (`Reputation.objects.all().delete()` — see [124-authentik](../../../wiki/containers/106-auth-outpost.md)),
|
||
and **broken default expression policies**. The restored DB carried 8 default policies authored in old `return`-style
|
||
syntax incompatible with 2026.5.x''s eval context; `ak apply_blueprints` re-applied the current defaults.
|
||
|
||
4. **Browser ran stale frontend JS.** Console showed `version 2026.2.2` while the backend was `2026.5.2` — because DNS
|
||
still pointed at the old LXC (see DNS cutover above), not a cache issue.
|
||
|
||
5. **WebAuthn devices dead post-migration.** Passkeys are device/origin-bound and don''t survive a host move. Deleted
|
||
all WebAuthn devices via Django ORM; users must re-register MFA.
|
||
|
||
|
||
## Architecture delta
|
||
|
||
|
||
| | Before | After |
|
||
|
||
|---|---|---|
|
||
|
||
| Authentik host | [LXC 124](../../../wiki/containers/106-auth-outpost.md) `192.168.8.180` | VPS `82.165.190.79`, `auth`
|
||
Docker net `172.30.1.0/24` |
|
||
|
||
| Version | `2026.2.2` | `2026.5.2` |
|
||
|
||
| `auth.hubris.network` (LAN) | dnsmasq → `192.168.8.175` (Caddy) | dnsmasq → `82.165.190.79` (VPS traefik) |
|
||
|
||
| `auth.hubris.network` (public) | IONOS wildcard → VPS → mesh → LXC 124 | IONOS wildcard → VPS → local container |
|
||
|
||
| Reachability | mesh-only (split-horizon) | **public** (auth UI now exposed on the open internet) |
|
||
|
||
| Redis | none (docs said 2026.x dropped it) | present (this deploy uses Redis) |
|
||
|
||
| DB engine | sqlite? (`StoreConfig` in mgmt) / Postgres for Authentik | Postgres, separate from netbird |
|
||
|
||
|
||
> **Security note:** the Authentik admin UI is now publicly reachable. Mitigated with a traefik `admin-allowlist` ipAllowList
|
||
middleware on `/if/admin/` (currently `5.61.168.0/24`). Login/flow endpoints are public by necessity. The `auth` Docker
|
||
network isolates Postgres/Redis from the netbird containers.
|
||
|
||
|
||
## Lessons learned
|
||
|
||
|
||
1. **Never put the IdP behind the VPN it authenticates.** The circular dependency is invisible until a restart, then it''s
|
||
total.
|
||
|
||
2. **`depends_on: condition: service_healthy` is the structural fix.** Docker refuses to start mgmt until Authentik is
|
||
healthy, regardless of boot order.
|
||
|
||
3. **Migrating an Authentik DB requires carrying its `AUTHENTIK_SECRET_KEY`.** It encrypts client secrets at rest; a mismatch
|
||
breaks SSO silently with no useful error.
|
||
|
||
4. **Default expression policies break across major Authentik versions.** Always run `ak apply_blueprints` after restoring
|
||
an older DB into a newer Authentik.
|
||
|
||
5. **Split-horizon DNS + a VPN resolver cache = an invisible stale answer.** The NetBird resolver (`100.122.255.254`)
|
||
outlives `netbird down/up`. Removing the dnsmasq override (fall through to public DNS) is the durable fix once a service
|
||
is genuinely public.
|
||
|
||
6. **Redirect URIs: prefer `STRICT` over `REGEX`** in Authentik 2026.x.
|
||
|
||
7. **Frontend/backend version skew shows as "Unknown error."** Confirm which Authentik the browser actually loaded (`version`
|
||
line in the console) before chasing backend logs.
|
||
|
||
8. **WebAuthn/passkeys do not survive host migrations** — delete and re-enroll.
|
||
|
||
|
||
## Forward-auth outpost — RESOLVED 2026-06-01
|
||
|
||
|
||
Forward-auth apps (Paperless, qBittorrent, Artifacto) initially still validated against LXC 124''s *embedded* outpost
|
||
(Caddy → `192.168.8.180:9000`) — split-brain against the frozen DB. Pointing Caddy at `https://auth.hubris.network` instead
|
||
fails: VPS Traefik rewrites `X-Forwarded-Host` → outpost can''t match the app → 404 (tested + reverted).
|
||
|
||
|
||
Fixed with a **dedicated LAN outpost** ([106 — auth-outpost](../../../wiki/containers/106-auth-outpost.md), `192.168.8.6`):
|
||
`goauthentik/proxy` connects outbound to the VPS core and serves forward-auth locally; Caddy → outpost over the LAN, no
|
||
Traefik, header preserved. Outpost `hubris-lan-outpost` carries the 3 proxy providers. Verified with 124-Authentik **stopped**.
|
||
This was Phase 1 of the broader architecture migration (plan: VPS edge / hubris LAN core / Mac Mini redundancy).
|
||
|
||
|
||
### 2026-06-05 — identification stage skip: broken "Trust me" reputation policy
|
||
|
||
|
||
The default authentication flow was skipping the identification stage (showing only a password field with `pending_user:
|
||
""`), making login impossible. Root cause: a **"Trust me" reputation policy** (`threshold: -3`, `negate: true`) attached
|
||
to the identification stage binding (#10). The `negate: true` + broken policy evaluation caused the stage binding to fail
|
||
re-evaluation, and Authentik''s `ReevaluateMarker` skipped the stage entirely.
|
||
|
||
|
||
**Fix:** Deleted from Postgres:
|
||
|
||
```sql
|
||
|
||
DELETE FROM authentik_policies_policybinding WHERE policy_id = ''<uuid>'';
|
||
|
||
DELETE FROM authentik_policies_reputation_reputationpolicy WHERE policy_ptr_id = ''<uuid>'';
|
||
|
||
DELETE FROM authentik_policies_policy WHERE policy_uuid = ''<uuid>'';
|
||
|
||
```
|
||
|
||
The policy was orphaned (no matched type data or had incompatible evaluation). Removing it restores normal two-stage login
|
||
(identification → password).
|
||
|
||
|
||
**Note:** This is different from the 2026-05-31 issue where NetBird''s `login_hint` OAuth2 param intentionally pre-identifies
|
||
users. That was expected behavior; this was a genuine bug.
|
||
|
||
|
||
## Outstanding
|
||
|
||
|
||
- **NetBird client version skew.** Mac client is `0.68.3` vs mgmt `0.71.3`; NetBird-SSH cert exchange fails between them.
|
||
Upgrade all clients toward `0.71.x`.
|
||
|
||
- **republic-laptop** regular SSH publickey auth still rejected (correct `authorized_keys`/perms/SELinux context, yet
|
||
denied — `sshd_config` `AuthorizedKeysFile` suspect). DNS not yet flushed there either.
|
||
|
||
- **muli-laptop** needs `netbird down && netbird up` + `resolvectl flush-caches`.
|
||
|
||
- **VPS port 22** opened for this repair; close once remote access is otherwise stable.
|
||
|
||
- **Decommission LXC 124 Authentik** after a ~2-week dual-run validation. dnsmasq stays on 124 regardless (separate service).
|
||
|
||
- **Reconcile [124-authentik](../../../wiki/containers/106-auth-outpost.md) provider notes** — docs describe a `Public`/PKCE
|
||
provider; the migrated DB carries the `Confidential` `netbird-dashboard` client. Verify which is live and correct the
|
||
page.
|
||
|
||
- **sops-encrypt** the VPS secrets (`/opt/authentik.env`) into the `secrets/` tree.
|
||
|
||
|
||
## Related
|
||
|
||
- [124 — authentik](../../../wiki/containers/106-auth-outpost.md)
|
||
|
||
- [DNS split-horizon](../../../wiki/infrastructure/dns.md)
|
||
|
||
- [Public ingress (VPS traefik)](../../../wiki/infrastructure/ingress.md)
|
||
|
||
- [Mesh migration](../../../wiki/infrastructure/mesh.md)
|
||
|
||
'
|
||
about_slugs:
|
||
- service:caddy
|
||
- service:authentik
|
||
- service:dns
|
||
- service:paperless
|
||
- service:artifacto
|
||
- host:hubris
|
||
tags:
|
||
- investigation
|
||
- slug: investigations/2026-06-01-mac-mini-onboarding
|
||
title: mac-mini onboarding — post-mortem & lessons learned
|
||
date: ''
|
||
status: resolved
|
||
duration: ''
|
||
content: "# mac-mini onboarding — post-mortem & lessons learned\n\nOnboarded the `mac-mini` workstation (macOS Sequoia,\
|
||
\ arm64) into the hubris\nhomelab context system with the `--with-hermes` profile. What follows is a\nchronological recap\
|
||
\ of every hitch, the fix, and the systemic improvements\nneeded so the next workstation takes 5 min instead of an hour.\n\
|
||
\n## Session log\n\n### Step 1 — clone + symlink\n- Manually cloned `git.hubris.network/dtoro/oikos` to `/Users/dtoro/oikos`.\n\
|
||
- Created `/opt/homelab-context` → `/Users/dtoro/oikos` symlink.\n- **Lesson:** bootstrap.sh was designed to do\
|
||
\ this from scratch, but we'd\n already cloned by hand. The bootstrap's `clone exists; pulling` path handled\n it gracefully.\n\
|
||
\n### Step 2 — hostname mismatch\n- `scutil --get LocalHostName` → `Davids-Mac-mini`\n- `hostname -s` → `Mac`\n- Inventory\
|
||
\ file: `hosts/mac-mini.yaml`\n- **Fix:** `sudo scutil --set LocalHostName mac-mini && sudo scutil --set HostName mac-mini`\n\
|
||
- **Lesson:** The bootstrap and `homelab whoami` use different hostname\n resolution. Bootstrap uses `scutil --get LocalHostName`\
|
||
\ (correct on macOS),\n but the `homelab` CLI binary uses `hostname -s`. Both need to match the\n inventory key. On\
|
||
\ a fresh macOS machine, neither does.\n\n### Step 3 — bootstrap dependencies\n- pyyaml was missing → `pip install pyyaml`\n\
|
||
- age and sops were missing → `brew install age sops`\n- Netbird was already installed and connected ✓\n- **Lesson:**\
|
||
\ The bootstrap preflight handles these, but only if running\n `bootstrap.sh` from the start. Since we ran it after manual\
|
||
\ setup, some\n steps (netbird install) were correctly skipped as already-present.\n\n### Step 4 — full bootstrap with\
|
||
\ `--with-mcp --with-hermes`\n- Ran `sudo HOMELAB_GITEA_TOKEN=... bash bootstrap.sh --with-mcp --with-hermes`\n- Age key\
|
||
\ issued ✓\n- Launchd sync timer installed ✓\n- Goose binary installed ✓\n- Hermes CLI linked ✓\n- MCP config merged ✓\n\
|
||
- `refresh-creds` skipped (not yet a recipient) ⚠️\n- Cosmetics: `chown: dtoro: illegal group name` at the end (benign,\
|
||
\ macOS\n group-naming quirk)\n\n### Step 5 — finalize from hubris\n- Ran `homelab client add mac-mini --finalize-pubkey\
|
||
\ <age...> --with-hermes` on\n hubris\n- Push failed: `[rejected] main -> main (fetch first)` — hubris clone was\n stale,\
|
||
\ bootstrap had already pushed from mac-mini\n- **Fix:** `git pull --rebase && git push` on hubris\n- **Lesson:** bootstrap\
|
||
\ pushes remote changes before hubris can finalize,\n creating a race. The `homelab client add --finalize-pubkey` command\
|
||
\ should\n pull before committing/pushing.\n\n### Step 6 — sops couldn't find the age key\n- `homelab secret hello` failed\
|
||
\ because sops looks in\n `/Users/dtoro/.ssh/id_rsa` etc. by default, not `/etc/age/key.txt`\n- The `homelab` CLI re-execs\
|
||
\ via `sudo -E env SOPS_AGE_KEY_FILE=... sops ...`,\n but this requires passwordless sudo and the correct env var passthrough\n\
|
||
- **Fix:**\n 1. Added NOPASSWD sudo rules\n 2. Eventually `SOPS_AGE_KEY` env with the raw key content worked directly\n\
|
||
- **Lesson:** Document the explicit `SOPS_AGE_KEY_FILE` incantation in\n agent-enrollment troubleshooting. New clients\
|
||
\ can't assume `homelab secret`\n works out of the gate — the sudo re-exec chain has permission pitfall.\n\n### Step\
|
||
\ 7 — OpenRouter key was a placeholder\n- `secrets/openrouter-api-key.yaml` contained\n `api_key: PLACEHOLDER_REPLACE_WITH_REAL_OPENROUTER_KEY`\n\
|
||
- User ran `sops` on hubris, but got the same error (age key not found on\n hubris either — `/root/.config/sops/age/keys.txt`\
|
||
\ didn't exist)\n- **Fix:** `SOPS_AGE_KEY_FILE=/etc/age/key.txt sops ...` on hubris.\n Later: the user pasted the real\
|
||
\ key, but the sops file showed\n `sk-or-...5c55` — the literal content was truncated with ellipsis.\n\n### Step 8 —\
|
||
\ editor loaded the wrong data\n- Neovim on the system is configured with `clipboard+=unnamedplus`, which\n points `*`\
|
||
\ and `+` registers to the macOS clipboard manager rather than\n X11. When editing SOPS files, this caused the **system\
|
||
\ clipboard** to be\n pasted instead of the actual ciphertext.\n- This wasn't diagnosed during the session — the sops\
|
||
\ file would load empty\n or show the wrong content because the editor's idea of \"paste\" was\n disconnected from what\
|
||
\ sops expected.\n- **Fix:** Run `sops` with `EDITOR=nano` or another editor that doesn't\n hijack OS clipboards:\n \
|
||
\ ```bash\n EDITOR=nano SOPS_AGE_KEY_FILE=/etc/age/key.txt sops secrets/openrouter-api-key.yaml\n ```\n- **Lesson:**\
|
||
\ Add a strong warning to `hermes-agent.md` / `agent-enrollment.md`:\n macOS neovim with `clipboard+=unnamedplus` silently\
|
||
\ breaks sops editing\n because the paste register reads from the system clipboard instead of the\n sops-managed buffer.\
|
||
\ Use `EDITOR=nano` or `EDITOR=vim` when running sops\n interactively. Alternatively, override the clipboard option with\n\
|
||
\ `EDITOR='nvim -c \"set clipboard=\"'`.\n- Also useful for the troubleshooting table in `agent-enrollment.md` under\
|
||
\ a\n new row: \"sops file loads empty / wrong content on macOS\"\n\n### Step 9 — model doesn't support tool use\n- Goose\
|
||
\ config pinned `nousresearch/hermes-4-405b` via OpenRouter\n- Error: `No endpoints found that support tool use`\n- **Fix:**\
|
||
\ Switched to `deepseek/deepseek-v4-flash` in\n `~/.config/goose/config.yaml`\n- Also updated `operations/hermes-agent.md`\
|
||
\ with the correct model\n- **Lesson:** The default model in `bootstrap.sh` and `hermes-agent.md` was\n never validated\
|
||
\ on OpenRouter for tool-use capability. Need to either:\n (a) Pin a model known to work (`deepseek/deepseek-v4-flash`)\n\
|
||
\ (b) Or have the bootstrap probe OpenRouter at install time\n\n### Step 10 — MCP homelab extension returns 404\n- Goose's\
|
||
\ homelab extension configured as `streamable_http` at\n `https://mcp.hubris.network/mcp` returns HTTP 404\n- The actual\
|
||
\ MCP server runs on SSE (`/sse`), not streamable_http\n- **Known follow-up** from `hermes-agent.md` — server needs migration\n\
|
||
- **Workaround:** The agent works fine via the built-in `developer` extension\n (shell + file) and the `homelab` CLI\n\
|
||
\n### Step 11 — git credential dance\nSeveral layers of issues with git authentication:\n\n1. **Directory permissions:**\
|
||
\ `/etc/homelab-context` was `0700 root` — dtoro\n couldn't access the credential file, and git's `store` helper couldn't\n\
|
||
\ create its `.lock` file.\n - **Fix 1a:** `sudo chmod 755 /etc/homelab-context`\n - **Fix 1b:** `sudo chown -R\
|
||
\ dtoro:staff /etc/homelab-context`\n\n2. **macOS built-in osxkeychain:** Apple's git has `osxkeychain` compiled in\n\
|
||
\ as a default credential helper. Even after configuring the `store` helper,\n git calls osxkeychain after every successful\
|
||
\ `get`, which prompts for\n macOS keychain access.\n - `credential.helper = \"\"` does NOT disable the compiled-in\
|
||
\ default\n - The `store` helper's `.lock` file in `/etc/homelab-context/` also failed\n silently, causing fallback\
|
||
\ behaviour\n - **Fix:** Embed the credential in the remote URL directly:\n `https://dtoro:TOKEN@git.hubris.network/dtoro/oikos.git`\n\
|
||
\ - This bypasses all credential helpers entirely\n\n3. **refresh-creds:** The credential is now managed by the remote\
|
||
\ URL.\n If the PAT is rotated, `homelab refresh-creds` won't update the URL.\n - **Workaround:** Run `git remote\
|
||
\ set-url origin` after `refresh-creds`\n\n### Step 12 — `homelab` CLI hostname detection\n- `homelab whoami` failed with\
|
||
\ `no hosts/Mac.yaml` because the CLI uses\n `hostname -s` (returning `Mac`) while the bootstrap uses\n `scutil --get\
|
||
\ LocalHostName` (which we fixed to `mac-mini`)\n- **Fix:** Set both `LocalHostName` and `HostName` via `scutil`\n- **Lesson:**\
|
||
\ The homelab CLI on macOS should prefer `scutil --get\n LocalHostName` like the bootstrap does, or at least try both\
|
||
\ and pick the\n one that matches an inventory key. This is a CLI bug.\n\n## Improvement backlog\n\n### High priority\n\
|
||
\n| # | Issue | Fix |\n|---|-------|-----|\n| 1 | Homelab CLI uses `hostname -s` on macOS; should use `scutil --get LocalHostName`\
|
||
\ to match bootstrap | Patch `bin/homelab` to try `scutil` first, fall back to `hostname -s` |\n| 2 | `homelab client\
|
||
\ add --finalize-pubkey` on hubris fails to push when remote is ahead (bootstrap pushes first) | `pull --rebase` before\
|
||
\ committing in `cmd_client_add` |\n| 3 | Default Goose model (`nousresearch/hermes-4-405b`) doesn't support tool use\
|
||
\ on OpenRouter | Update `bootstrap.sh` default and `hermes-agent.md` to `deepseek/deepseek-v4-flash` |\n| 4 | MCP server\
|
||
\ needs streamable_http migration | Follow-up #1 from `hermes-agent.md` — FastMCP `transport=\"sse\"` → `transport=\"\
|
||
streamable_http\"` |\n| 5 | `homelab refresh-creds` should also update embedded remote URLs | In `cmd_refresh_creds`,\
|
||
\ after writing the credential file, update any remote URL that has an embedded credential |\n\n### Medium priority\n\n\
|
||
| # | Issue | Fix |\n|---|-------|-----|\n| 6 | macOS neovim with `clipboard+=unnamedplus` breaks sops editing | Add troubleshooting\
|
||
\ row to `agent-enrollment.md`; recommend `EDITOR=nano` |\n| 7 | `/etc/homelab-context` directory with 0700 root permissions\
|
||
\ blocks non-root git | Change to 755 in `bootstrap.sh` on all OSes (or at least macOS) |\n| 8 | Bootstrap `chown` at\
|
||
\ end gives `illegal group name` on macOS | Fix group detection in bootstrap.sh for macOS |\n| 9 | `homelab secret` re-exec\
|
||
\ via sudo chain is fragile — needs passwordless sudo or TTY | Consider `SOPS_AGE_KEY` env-var fallback path in the CLI\
|
||
\ itself |\n\n### Low priority\n\n| # | Issue | Fix |\n|---|-------|-----|\n| 10 | Onboarding this machine revealed no\
|
||
\ `HERMES.md` file in the repo | Create the persona file for the Goose hints symlink |\n| 11 | Git credential `store`\
|
||
\ helper lock mechanism is OS-sensitive and fragile | Document embedded-URL pattern as the preferred approach for workstations\
|
||
\ |\n| 12 | AGENTS.md references `/opt/homelab-context/hosts/<hostname>.yaml` but the canonical path is now a symlink\
|
||
\ | Verify the path resolution edge cases |\n\n## Changelog\n\n### 2026-06-01 — initial post-mortem\nCaptured the full\
|
||
\ onboarding of mac-mini into the hubris homelab context\nsystem, including the hostname mismatch, sudo/sops credential\
|
||
\ chain, model\nselection, git credential gymnastics, and editor clipboard interference.\n\n### 2026-06-02 — MCP server\
|
||
\ migrated to streamable-http\nItem #4 resolved. FastMCP server now uses `transport=\"streamable-http\"` (hyphen,\nnot\
|
||
\ underscore), serving at `https://mcp.hubris.network/mcp`. All doc references\nand client configs (Goose, bootstrap.sh,\
|
||
\ agent-enrollment.md) updated. Service\nre-deployed on apps LXC (105).\n"
|
||
about_slugs:
|
||
- ws:mac-mini
|
||
- host:hubris
|
||
- host:strong
|
||
tags:
|
||
- investigation
|
||
- slug: investigations/2026-06-03-moonlight-sunshine-wifi-jitter
|
||
title: 2026-06-03 — Moonlight/Sunshine game streaming unstable over WiFi
|
||
date: '2026-06-03'
|
||
status: resolved
|
||
duration: ''
|
||
content: '# 2026-06-03 — Moonlight/Sunshine game streaming unstable over WiFi
|
||
|
||
|
||
## Summary
|
||
|
||
|
||
[`ludo-mini`](../../../hosts/strong.yaml) runs Sunshine as the game-streaming server; [`mac-mini`](../../../hosts/mac-mini.yaml)
|
||
runs Moonlight as the client. Despite both machines being on the same physical subnet (192.168.178.0/24), streaming was
|
||
unstable — stuttering, dropouts, and high latency. Root cause: **mac-mini is connected only via WiFi**, while ludo-mini
|
||
is wired Ethernet (2.5 Gbps). WiFi throughput shows 1-second UDP dropouts and high jitter (28 ms stddev), which breaks
|
||
real-time video streaming.
|
||
|
||
|
||
## Timeline
|
||
|
||
|
||
### Pre-existing state
|
||
|
||
- ludo-mini: wired Ethernet (eno1, 2.5 Gbps), 192.168.178.181/24
|
||
|
||
- mac-mini: WiFi (en1, 802.11ac, 234 Mbps Tx rate, 1 stream, -60 dBm RSSI), 192.168.178.99/24
|
||
|
||
- Both on the same subnet via a consumer router at 192.168.178.1
|
||
|
||
- Sunshine configured: bitrate=80 Mbps, encoder=amf (AMD GPU), fec_percentage=5, hevc_mode=2
|
||
|
||
- Moonlight configured: bitrate=57 Mbps, fps=60, localaddr=192.168.178.181 (correct)
|
||
|
||
|
||
### 2026-06-03 — investigation
|
||
|
||
**Network tests (iperf3 between mac-mini ↔ ludo-mini):**
|
||
|
||
- TCP throughput: 49–86 Mbps, average ~65 Mbps — highly variable
|
||
|
||
- UDP throughput: 1-second complete dropout during a 5-second test; 0% loss otherwise
|
||
|
||
- Ping: 4.7 ms avg, but **28 ms stddev**, max **138 ms** — WiFi-level jitter
|
||
|
||
- Netbird P2P connection also showed 7 ms latency over the tunnel (unnecessary given same-subnet direct connection)
|
||
|
||
|
||
**Additional findings:**
|
||
|
||
- mac-mini has a built-in Ethernet port (en0) but it is inactive — no cable connected
|
||
|
||
- mac-mini also has Thunderbolt Ethernet adapters (en5, en6, en7) all inactive
|
||
|
||
- ludo-mini''s en0 shows `speed 2500` (2.5 GbE)
|
||
|
||
- Netbird routes `192.168.8.0/24` via `utun100`, but 192.168.178.0/24 traffic stays on `en1`
|
||
|
||
|
||
## Root cause
|
||
|
||
|
||
mac-mini is on WiFi. WiFi introduces:
|
||
|
||
1. **Throughput variance** — 49–86 Mbps TCP fluctuation
|
||
|
||
2. **Dropouts** — 1-second UDP blackouts from radio retransmissions
|
||
|
||
3. **Jitter** — 28 ms stddev with spikes to 138 ms
|
||
|
||
|
||
Moonlight/Sunshine streams real-time video over UDP. Any jitter spike or dropout causes visual stutter or frame drops.
|
||
The 57 Mbps Moonlight bitrate was too aggressive for the WiFi link''s consistency, even though the average throughput
|
||
is ~65 Mbps.
|
||
|
||
|
||
## Mitigations applied
|
||
|
||
|
||
### mac-mini — Moonlight (via `defaults write`)
|
||
|
||
|
||
| Setting | Before | After | Rationale |
|
||
|
||
||---------|--------|-------|-----------|
|
||
|
||
|| bitrate | 57 Mbps | 25 Mbps | Reduce to a level WiFi can sustain consistently |
|
||
|
||
|| framepacing | off (0) | on (1) | Smooths frame timing over variable latency |
|
||
|
||
|| fps | 30 | 60 | Restored to game-appropriate value |
|
||
|
||
|
||
### mac-mini — Moonlight (current, wired — 2026-06-04)
|
||
|
||
|
||
| Setting | WiFi value | Wired value | Rationale |
|
||
|
||
||---------|-----------|-------------|-----------|
|
||
|
||
|| bitrate | 60 Mbps | 80 Mbps | Wired 1 GbE can sustain comfortably |
|
||
|
||
|| framepacing | on (1) | off (0) | Wired latency is consistent, no smoothing needed |
|
||
|
||
|| fps | 60 | 60 | Unchanged |
|
||
|
||
|
||
### ludo-mini — Sunshine (via SSH)
|
||
|
||
|
||
| Setting | WiFi value | Wired value | Rationale |
|
||
|
||
||---------|-----------|-------------|-----------|
|
||
|
||
|| bitrate | 40 Mbps | 80 Mbps | Wired can handle full bandwidth |
|
||
|
||
|| fec_percentage | 2 | 5 | Restored to default — wired has no dropout concern |
|
||
|
||
|| packetsize | 1024 | 1316 | Restored to default (Ethernet MTU) |
|
||
|
||
|
||
Sunshine restarted after changes; confirmed active.
|
||
|
||
|
||
## Resolution
|
||
|
||
|
||
### 2026-06-04 — mac-mini wired to LAN
|
||
|
||
|
||
mac-mini plugged into Ethernet (en0, 192.168.178.182). This definitively resolves the WiFi jitter that caused streaming
|
||
instability. All WiFi-era conservative settings reverted to appropriate wired values on both ends.
|
||
|
||
|
||
## Open questions
|
||
|
||
|
||
~~1. **Ethernet wiring** — mac-mini has an active en0 port. Plugging it into the same switch/router as ludo-mini would
|
||
eliminate WiFi jitter entirely. This is the definitive fix.~~ **RESOLVED** — mac-mini now wired via en0.
|
||
|
||
2. **WiFi alternatives** — If wiring is impractical, a WiFi 6 (802.11ax) USB adapter or Thunderbolt-to-Ethernet adapter
|
||
would help, but wired Ethernet is the only reliable solution for game streaming.
|
||
|
||
3. **Netbird route** — `192.168.8.0/24` is routed over Netbird `utun100`. If Moonlight ever falls back to that subnet
|
||
(ludo-mini''s 192.168.8.133 LAN IP), traffic tunnels unnecessarily. Moonlight correctly uses 192.168.178.181, so this
|
||
is not currently an issue.
|
||
|
||
|
||
## Changelog
|
||
|
||
|
||
### 2026-06-04 — mac-mini wired to LAN, configs reverted from WiFi mitigations
|
||
|
||
- mac-mini connected via en0 (wired Ethernet, 192.168.178.182), WiFi mitigations no longer needed
|
||
|
||
- Moonlight: bitrate 60→80 Mbps, framepacing disabled
|
||
|
||
- Sunshine: bitrate 40→80 Mbps, fec_percentage 2→5, packetsize restored to default (1316)
|
||
|
||
- First root cause in investigation (Open questions #1) resolved'
|
||
about_slugs:
|
||
- ws:mac-mini
|
||
- host:strong
|
||
tags:
|
||
- investigation
|
||
- slug: investigations/2026-06-06-authentik-session-lifetime
|
||
title: 2026-06-06 — Frequent Authentik login prompts fixed (session duration)
|
||
date: '2026-06-06'
|
||
status: resolved
|
||
duration: ''
|
||
content: "# 2026-06-06 — Frequent Authentik login prompts fixed (session duration)\n\n## Summary\n\nUser needed to re-authenticate\
|
||
\ to Authentik several times per day. Root cause was the Django session being configured as a **session cookie** (cleared\
|
||
\ on browser close) with a short 24-hour lifetime. Fixed by extending both the session cookie lifetime and the user_login\
|
||
\ stage session duration to 30 days.\n\n## Timeline\n\n- **2026-06-06 ~23:00** — User reports \"having to login to authentik\
|
||
\ several times per day\"\n- **~23:10** — Investigation begins: check Authentik OAuth2 provider settings, proxy session\
|
||
\ table, Django session config\n- **~23:20** — Found `SESSION_EXPIRE_AT_BROWSER_CLOSE=True` — the `authentik_session`\
|
||
\ cookie has no Max-Age, cleared on browser close\n- **~23:25** — Found the `user_login` stage has `session_duration=seconds=0`,\
|
||
\ which calls `session.set_expiry(0)` → `_session_expire_at_browser_close=True`\n- **~23:30** — Applied DB fix: updated\
|
||
\ `authentik_stages_user_login_userloginstage` → `session_duration='days=30'`\n- **~23:35** — Added `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30`\
|
||
\ to `/opt/authentik.env` (maps to `sessions.unauthenticated_age` in Authentik config, setting `SESSION_COOKIE_AGE` to\
|
||
\ 30 days)\n- **~23:40** — Recreated container with `docker compose up -d` (restart doesn't re-read env file)\n- **~23:45**\
|
||
\ — Verified both changes took effect\n\n## Root cause\n\nTwo independent but compounding issues:\n\n1. **`SESSION_EXPIRE_AT_BROWSER_CLOSE=True`**\
|
||
\ — hardcoded in `/authentik/root/settings.py`. Makes the `authentik_session` cookie a session cookie (no `Max-Age`),\
|
||
\ cleared when the browser closes. Cannot be changed via env vars or config files (hardcoded in Python source).\n\n2.\
|
||
\ **`user_login` stage `session_duration=seconds=0`** — the default-authentication-login stage called `session.set_expiry(0)`,\
|
||
\ which in Django sets `_session_expire_at_browser_close=True` on the session object, overriding any session cookie lifetime\
|
||
\ set via `SESSION_COOKIE_AGE`.\n\n3. **`sessions.unauthenticated_age=days=1`** (default) — `SESSION_COOKIE_AGE` was 86,400\
|
||
\ seconds (24 hours), so even with the browser left open, the server-side session data expired after 24 hours with no\
|
||
\ way to extend it (`SESSION_SAVE_EVERY_REQUEST=False`).\n\n### Why forward-auth worked but OAuth2 didn't\n\n- **Forward-auth\
|
||
\ (Caddy-gated services)** — The `authentik_proxy_*` cookie is a *persistent* cookie with 24-hour validity, set by the\
|
||
\ Authentik proxy outpost. It survives browser restart. Caddy's forward-auth validates this cookie directly with the outpost\
|
||
\ API — no Django session needed.\n- **OAuth2 (Gitea, Jellyfin, etc.)** — These services redirect to `auth.hubris.network/application/o/authorize/`,\
|
||
\ which checks the Django session (`authentik_session` cookie). If the browser was closed, this cookie is gone → user\
|
||
\ is redirected to the login form.\n\n## Changes applied\n\n### 1. Database — `authentik_stages_user_login_userloginstage`\n\
|
||
\n```sql\nUPDATE authentik_stages_user_login_userloginstage\nSET session_duration = 'days=30'\nWHERE stage_ptr_id = (\n\
|
||
\ SELECT stage_uuid FROM authentik_flows_stage \n WHERE name = 'default-authentication-login'\n);\n```\n\nThis causes\
|
||
\ the login stage to call `session.set_expiry(timedelta(days=30))`, which:\n- Sets `_session_expire_at_browser_close=False`\
|
||
\ for that session\n- Sets the session cookie `Max-Age` to 30 days (persistent cookie)\n- Sets the server-side session\
|
||
\ expiry to 30 days\n\n### 2. Environment — `AUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE`\n\nAdded to `/opt/authentik.env`:\n\
|
||
\n```\nAUTHENTIK_SESSIONS__UNAUTHENTICATED_AGE=days=30\n```\n\nThis maps to config key `sessions.unauthenticated_age`,\
|
||
\ which Authentik's `settings.py` reads via:\n\n```python\nSESSION_COOKIE_AGE = timedelta_from_string(\n CONFIG.get(\"\
|
||
sessions.unauthenticated_age\", \"days=1\")\n).total_seconds()\n```\n\nResult: `SESSION_COOKIE_AGE` = 2,592,000 seconds\
|
||
\ (30 days). This is a fallback for sessions that don't go through the `user_login` stage or when the stage's explicit\
|
||
\ `set_expiry` doesn't apply.\n\n## Verification\n\n```python\n# Inside authentik-server container:\nimport os\nos.environ.setdefault(\"\
|
||
DJANGO_SETTINGS_MODULE\", \"authentik.root.settings\")\nimport django\ndjango.setup()\nfrom django.conf import settings\n\
|
||
print(\"SESSION_COOKIE_AGE:\", settings.SESSION_COOKIE_AGE) # → 2592000.0\nprint(\"SESSION_EXPIRE_AT_BROWSER_CLOSE:\"\
|
||
, settings.SESSION_EXPIRE_AT_BROWSER_CLOSE) # → True\n\nfrom authentik.stages.user_login.models import UserLoginStage\n\
|
||
stage = UserLoginStage.objects.filter(name=\"default-authentication-login\").first()\nprint(\"session_duration:\", stage.session_duration)\
|
||
\ # → \"days=30\"\n```\n\n## Open questions\n\n- `SESSION_COOKIE_SECURE=False` — Should be `True` since Authentik is\
|
||
\ HTTPS-only, but the custom `SessionMiddleware.is_secure()` method dynamically sets the cookie's `secure` flag based\
|
||
\ on the request, so it's fine.\n- 30 days is generous. Could be reduced to 7-14 days if desired. Change the DB value\
|
||
\ and env var accordingly.\n\n## Related\n\n- [Container 106 — auth-outpost](../../wiki/containers/106-auth-outpost.md)\n\
|
||
- [Authentik VPS migration](archive/2026-05-31-authentik-vps-migration.md)\n- [Ingress (VPS Traefik)](../../wiki/infrastructure/ingress.md)\n\
|
||
- `.hermes/plans/2026-06-06_232200-authentik-frequent-login-fix.md` — original plan\n\n## Changelog\n\n### 2026-06-06\
|
||
\ — created\nDocumented the session lifetime investigation, root cause, and applied fixes."
|
||
about_slugs:
|
||
- service:caddy
|
||
- service:authentik
|
||
- service:gitea
|
||
- service:jellyfin
|
||
- host:hubris
|
||
tags:
|
||
- investigation
|
||
- slug: investigations/2026-06-06-caddyfile-truncation
|
||
title: 'Investigation: Caddyfile truncation — all LAN services down (2026-06-06)'
|
||
date: '2026-06-06'
|
||
status: resolved
|
||
duration: ~10 hours (from last known good state ~12:39 UTC to restoration ~22:40 UTC)
|
||
content: '# Investigation: Caddyfile truncation — all LAN services down (2026-06-06)
|
||
|
||
|
||
**Date:** 2026-06-06
|
||
|
||
**Status:** resolved
|
||
|
||
**Duration:** ~10 hours (from last known good state ~12:39 UTC to restoration ~22:40 UTC)
|
||
|
||
|
||
## Symptom
|
||
|
||
|
||
All `*.hubris.network` URLs except `photos.hubris.network` and `auth.hubris.network` returned `tlsv1 alert internal error`
|
||
or TCP timeouts from LAN/mesh clients. `dig @192.168.8.2` and `dig @100.122.255.254` both resolved to `192.168.8.175`
|
||
correctly — DNS was fine. The issue was at the Caddy level.
|
||
|
||
|
||
## Root cause
|
||
|
||
|
||
The Caddyfile on LXC 121 was manually edited directly on the filesystem (not via the `dtoro/caddy-conf` git repo), reducing
|
||
it from 260 lines/30+ site blocks to 43 lines with only 3 photo-related site blocks: `photos.hubris.network`, `prism.hubris.network`,
|
||
and `photos2.hubris.network`.
|
||
|
||
|
||
### Timeline
|
||
|
||
|
||
| Time (UTC+2) | Event |
|
||
|
||
|---|---|
|
||
|
||
| Jun 04 23:43 | Last successful git-push deploy — full Caddyfile (260 lines) |
|
||
|
||
| Jun 06 ~12:00 | Caddyfile manually edited locally, truncating to 3 sites |
|
||
|
||
| Jun 06 12:39 | Deploy webhook triggered → `git pull --ff-only` failed: "Your local changes would be overwritten" |
|
||
|
||
| Jun 06 14:13 | Deploy webhook triggered again → `deploy ok` (the truncated file was committed or merged somehow) |
|
||
|
||
| Jun 06 22:34 | Investigation began |
|
||
|
||
| Jun 06 22:43 | Caddyfile restored from `origin/master`, `systemctl reload caddy` |
|
||
|
||
|
||
### Evidence
|
||
|
||
|
||
- `git diff HEAD -- Caddyfile` on LXC 121: `+3 / -159` lines
|
||
|
||
- Git reflog: HEAD at `32575ce` (fix: sab port 8081→8082), working tree diverged
|
||
|
||
- Backup file `Caddyfile.bak.1780263919`: 225 lines, full original config
|
||
|
||
- `git stash list` shows one auto-stash entry
|
||
|
||
- `origin/master` at `1b977aa`: 260 lines, all site blocks present
|
||
|
||
|
||
### Secondary root cause found during investigation
|
||
|
||
|
||
**elementsynapse (LXC 118)** had `iface eth0 inet dhcp` internally despite `pct set 118 --net0 ... ip=192.168.8.239/24`.
|
||
On DHCP lease renewal, dhclient grabbed `.244` from Technitium''s pool. Caddy''s `reverse_proxy 192.168.8.239:8008` was
|
||
hitting a dead IP.
|
||
|
||
|
||
This is the same class of drift as the June 5th incidents (paperless, HAOS, apps, mule-images). Elementsynapse was missed
|
||
during the 2026-06-02 static-IP migration.
|
||
|
||
|
||
## Fix applied
|
||
|
||
|
||
1. **Caddyfile** → `git checkout --force origin/master -- Caddyfile` + `systemctl reload caddy`
|
||
|
||
2. **elementsynapse** → replaced `iface eth0 inet dhcp` with static, killed dhclient, verified connectivity
|
||
|
||
|
||
## Permanent safeguards (all deployed)
|
||
|
||
|
||
| Safeguard | Location | What it does |
|
||
|
||
|---|---|---|
|
||
|
||
| Site-count guard | `/etc/caddy/scripts/deploy.sh` | Refuses reload if <20 `hubris.network` site blocks |
|
||
|
||
| Dirty-tree auto-stash | `/etc/caddy/scripts/deploy.sh` | Stashes local edits before git pull |
|
||
|
||
| Auto-backup | `/etc/caddy/scripts/deploy.sh` | Saves Caddyfile.bak.<timestamp> before any change, keeps 5 |
|
||
|
||
| Caddy backend health | `/etc/cron.d/caddy-backend-health` on hubris | Runs `check-caddy-backends.sh` every 10 min |
|
||
|
||
| DNS sync | `/etc/cron.d/dns-sync` on LXC 107 | Runs `dns-sync.py` every 10 min (was missing since 2026-06-04) |
|
||
|
||
|
||
## Related
|
||
|
||
|
||
- DHCP drift investigation (previous incident) — not filed as its own investigation; see the [DNS sync fix](../../../.hermes/plans/2026-06-05_170000-prevent-dhcp-ip-drift.md)
|
||
|
||
- [Caddy (121)](../../wiki/containers/121-caddy.md)
|
||
|
||
- [elementsynapse (118)](../../wiki/containers/118-elementsynapse.md)
|
||
|
||
- [dns-sync script](../../../scripts/dns-sync.py)
|
||
|
||
- [check-caddy-backends script](../../../scripts/check-caddy-backends.sh)'
|
||
about_slugs:
|
||
- service:caddy
|
||
- service:dns
|
||
- service:paperless
|
||
- host:hubris
|
||
tags:
|
||
- investigation
|
||
runbooks:
|
||
- slug: client-enrollment
|
||
name: Client enrollment
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: client-enrollment\nrisk_class: config_mutation\ninputs: [hostname, kind, role]\nverification: \"homelab\
|
||
\ doctor (on the new client)\"\ndocs_update_checklist: [hosts_narrative_page_if_lxc_or_vm]\n---\n\n# Client enrollment\n\
|
||
\nGoal: bring a new host (workstation, LXC, VM) into inventory and the\nsecrets model, with mesh membership only where\
|
||
\ it's actually needed.\nThis wraps the existing `homelab client add` flow — see\n[operations/agent-enrollment.md](../../operations/agent-enrollment.md)\
|
||
\ for\nthe full walkthrough; this runbook is the risk/lifecycle framing.\n\n1. On any enrolled client: `homelab client\
|
||
\ add <hostname>` — appends a\n `hosts.<name>:` block to `inventory.yaml` (lifecycle `state: planned`\n → `provisioning`,\
|
||
\ per [oikos/ontology.yaml](../../../oikos/ontology.yaml)),\n commits + pushes.\n2. Netbird join is **optional, not\
|
||
\ a required step** — only needed for\n hosts that must be reachable off-LAN (workstations that roam, e.g.\n `republic-laptop`,\
|
||
\ `mac-mini`). A node reachable on the household LAN\n (192.168.8.0/24 — most LXCs/VMs) doesn't need it: it's already\n\
|
||
\ reachable directly, and off-LAN clients reach it too via hubris's\n routed `192.168.8.0/24` Netbird network resource.\
|
||
\ Skip this step for\n LAN-only nodes; do it (out-of-band, console or setup key) only for\n hosts that need independent\
|
||
\ off-LAN reachability.\n3. On the new host: run `bootstrap.sh` (add `--with-hermes` to also\n enroll the Hermes agent).\
|
||
\ This provisions `/etc/age/key.txt`, the\n sync timer, and prints an age pubkey.\n4. Back on an enrolled client: `homelab\
|
||
\ client add <hostname>\n --finalize-pubkey <age1...>` — sets `age_pubkey`, grants shared\n secrets, re-keys SOPS,\
|
||
\ commits + pushes. This is the\n `provisioning → active` transition.\n5. Verify: `homelab doctor` on the new client\
|
||
\ should show all checks\n green (clone, sync timer, age key, CLI symlink, MCP reachable).\n\nDocs-update checklist:\
|
||
\ if the new host is an LXC/VM, add its narrative\npage under `containers/` or `vms/` and set `doc_page` in its inventory\n\
|
||
entry (host-level cards don't have a `doc_page` field yet — services do;\nnarrative pages are still found via the generated\
|
||
\ `see_also` in\n`hosts/<name>.yaml`).\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: config-change-deploy
|
||
name: Config change + deploy
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: config-change-deploy\nrisk_class: config_mutation\ninputs: [service_name, change_description]\nverification:\
|
||
\ \"curl -sf <service_url> (or homelab service <name> health)\"\ndocs_update_checklist: [doc_page, changelog]\n---\n\n\
|
||
# Config change + deploy\n\nGoal: change a tracked config repo (Caddy, Gitea customizations, an app's\nown repo) and get\
|
||
\ it live, safely.\n\n1. `homelab change preflight <service>` — current health, the service's\n `config_repo`, its risk\
|
||
\ class, and the verification command to run\n after. If risk class requires approval (`config_mutation` or\n `destructive`),\
|
||
\ stop and get operator sign-off before editing — see\n `oikos/policy.yaml`.\n2. Clone/pull the `config_repo` (never\
|
||
\ edit the backend's working tree\n directly — tracked configs change by commit + push, per\n [OIKOS.md](../../OIKOS.md)\
|
||
\ conventions).\n3. Make the change, commit, push to `main`.\n4. The Gitea webhook fires the deploy pipeline for that\
|
||
\ repo (see\n [infrastructure/auto-deploy.md](../../../knowledge/wiki/infrastructure/auto-deploy.md) for\n the exact\
|
||
\ receiver/reload for this service).\n5. Run the preflight's verification command. If it fails, check\n `homelab service\
|
||
\ <name> log` for the reload/restart error.\n6. Record the change: once `oikos/ledger.py` is wired into deploy tooling\n\
|
||
\ (Week 3), this is automatic; until then, note the change and outcome\n in the relevant investigation/plan doc.\n\
|
||
\nDocs-update checklist: update the service's `doc_page` if the change\nalters its behavior, ingress route, or ownership;\
|
||
\ add a changelog entry\nif the page has one.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: docs-lint
|
||
name: Docs lint
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: docs-lint\nrisk_class: read_only\ninputs: [paths]\nverification: \"python3 .agents/skills/docs-lint/lint.py\"\
|
||
\ndocs_update_checklist: []\n---\n\n# Docs lint\n\nCheck committed documentation against the mechanical rules in\n[writing-style.md](../../shared/writing-style.md):\
|
||
\ banned vocabulary and broken relative markdown\nlinks. Prose-voice rules are not machine-checkable — those stay a review\
|
||
\ responsibility.\n\nRun from the repo root:\n\n python3 .agents/skills/docs-lint/lint.py # default: knowledge/\
|
||
\ .agents/ operations/ investigations/ plans/\n python3 .agents/skills/docs-lint/lint.py knowledge/wiki/containers/104-gitea.md\n\
|
||
\nExit code is non-zero when any violation is found, so it can gate a commit. The banned-vocabulary\nlist mirrors `writing-style.md`;\
|
||
\ update both together if the standard changes.\n\n> **Known baseline.** `knowledge/wiki/containers/101-jellyfin.md` links\
|
||
\ into a sibling repo\n> (`devops/homelab-authentik-admin`) that this checkout does not contain — expected, not a bug.\n\
|
||
> Any other broken link is a real regression; investigate before dismissing it as baseline noise.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: incident-investigation
|
||
name: Incident investigation
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: incident-investigation\nrisk_class: read_only\ninputs: [symptom, affected_entity]\nverification: \"\
|
||
n/a — investigation produces a written record, not a state change\"\ndocs_update_checklist: [investigations_entry]\n---\n\
|
||
\n# Incident investigation\n\nGoal: understand what broke and why, before touching anything.\n\n1. `homelab service <name>\
|
||
\ explain` (or `homelab node <name> relations`\n if the affected entity is a host) — get the blast radius and doc\n\
|
||
\ pointer first. Don't start pulling logs blind.\n2. `homelab service <name> health` + `homelab service <name> log`\
|
||
\ (or\n MCP `get_service_status` / `tail_log`) for the affected service.\n3. Walk the blast radius: is a shared dependency\
|
||
\ down (`caddy`, `dns`,\n `authentik`, or the backend host itself)? `homelab node <name>\n relations` shows \"affected\
|
||
\ by\" — check those first.\n4. `homelab apt-audit` if the symptom looks like a dpkg/upgrade\n interaction.\n5. Check\
|
||
\ the change ledger for recent mutations to the affected entity\n or anything upstream of it: `homelab service <name>\
|
||
\ history` (once\n populated) or grep `ledger/*.jsonl`.\n6. Write findings to a new `knowledge/sources/investigations/<date>-<slug>.md`\
|
||
\ — symptom,\n timeline, root cause, fix applied, prevention. This is the durable\n record; don't rely on chat history.\n\
|
||
\nDocs-update checklist: always create the investigation entry. If the\nroot cause was stale/wrong inventory data (a `doc_page`,\
|
||
\ `config_repo`,\nor `backend` that didn't match reality — this happened during Week 1\nkernel work, see the `authentik`\
|
||
\ backend fix), correct `inventory.yaml`\nin the same session.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: lifecycle-activate-node
|
||
name: 'Lifecycle: activate a node'
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: lifecycle-activate-node\nrisk_class: config_mutation\ninputs: [node_name]\nverification: \"homelab\
|
||
\ service <name> health (if it hosts a service); homelab doctor (if it's a client)\"\ndocs_update_checklist: [doc_page_complete]\n\
|
||
transition: \"provisioning -> active\"\n---\n\n# Lifecycle: activate a node\n\nPer [oikos/ontology.yaml](../../../oikos/ontology.yaml).\
|
||
\ Requires: age key\nenrolled if it needs secrets, mesh joined if it needs off-LAN reach,\ningress live if public, health\
|
||
\ check answering, doc page complete,\nledger entry.\n\n1. If the node is a `homelab` client: finish enrollment per\n\
|
||
\ [client-enrollment.md](../client-enrollment/SKILL.md) (`--finalize-pubkey`,\n mesh join, `homelab doctor` green).\n\
|
||
2. If it hosts a public service: add the `services:` entry in\n `inventory.yaml` (backend, url, doc_page, config_repo,\
|
||
\ risk_notes —\n see the Week-1 service contract fields) and wire the Caddy route in\n `dtoro/caddy-conf`.\n3. Confirm\
|
||
\ the health check answers: `homelab service <name> health` or\n a direct `curl`.\n4. Flip `state: provisioning` → `state:\
|
||
\ active` (or delete the `state:`\n field — `active` is the default) in `inventory.yaml`.\n5. Complete the doc page\
|
||
\ (stub → full narrative: role, specs, how it's\n configured, dependencies).\n6. Record the activation: `oikos/ledger.py\
|
||
\ append host:<name> activate\n config_mutation --result ok` (or let the CLI wrapper do this once\n Week 3's runbook\
|
||
\ automation lands).\n\nRegenerate derived data: `python3 mcp/build_host_files.py && python3\noikos/gen-topology.py` so\
|
||
\ `hosts/<name>.yaml`, the topology diagram, and\nthe context card all reflect the new state.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: lifecycle-deprecate-node
|
||
name: 'Lifecycle: deprecate a node'
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: lifecycle-deprecate-node\nrisk_class: config_mutation\ninputs: [node_name, replacement_node_or_reason]\n\
|
||
verification: \"homelab node <name> relations — 'affected by' must be empty before completing\"\ndocs_update_checklist:\
|
||
\ [doc_page_deprecation_note]\ntransition: \"active -> deprecated\"\n---\n\n# Lifecycle: deprecate a node\n\nPer [oikos/ontology.yaml](../../../oikos/ontology.yaml):\
|
||
\ a node keeps running\nbut takes no new dependents. **Completion condition: zero remaining\ninbound `depends-on`/`routes-to`\
|
||
\ edges** — this is a hard gate, not a\nsuggestion; `oikos/policy.yaml` `lifecycle_overrides.deprecated.refuse`\nlists\
|
||
\ `new-inbound-edges` as refused going forward.\n\n1. Set `state: deprecated` on the node.\n2. `homelab node <name> relations`\
|
||
\ — read `affected_by`. Every entry\n there is something still relying on this node.\n3. Migrate or retire each dependent\
|
||
\ one at a time (point its `backend`/\n `config_repo`/ingress route elsewhere, or deprecate it too if it's\n being\
|
||
\ retired alongside).\n4. Re-run `homelab node <name> relations` after each dependent is moved.\n The transition to\
|
||
\ `destroyed` is only safe once `affected_by` is\n empty — check this every time, don't assume from memory.\n5. Note\
|
||
\ the deprecation on the doc page: reason, replacement (if any),\n date.\n\nIf step 2 shows dependents you didn't expect,\
|
||
\ stop and investigate\nbefore proceeding — that's exactly the kind of drift the Week-3 detector\nwill catch automatically,\
|
||
\ but until then this manual check is the gate.\n\nNext (once `affected_by` is empty):\n[lifecycle-destroy-node.md](../lifecycle-destroy-node/SKILL.md).\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: lifecycle-destroy-node
|
||
name: 'Lifecycle: destroy a node'
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: lifecycle-destroy-node\nrisk_class: destructive\ninputs: [node_name]\nverification: \"homelab node\
|
||
\ <name> relations returns unknown-entity; pct list on the backend no longer shows it\"\ndocs_update_checklist: [archaeology_entry,\
|
||
\ containers_index_update]\ntransition: \"deprecated -> destroyed\"\n---\n\n# Lifecycle: destroy a node\n\n**Destructive.**\
|
||
\ Requires operator approval + typed confirmation phrase\nper `oikos/policy.yaml`. Requires (ontology): backups verified,\
|
||
\ secrets\nrecipients removed + re-keyed, ingress/DNS removed, archaeology entry,\nledger entry.\n\n1. Confirm the node\
|
||
\ is `deprecated` with zero `affected_by` edges\n (`homelab node <name> relations`) — do not skip this even if the\n\
|
||
\ deprecation runbook was followed recently; state can drift.\n2. If it's an enrolled client: `homelab client remove\
|
||
\ <name>` — revokes\n the age key, re-keys SOPS, removes the inventory entry. This is\n already destructive-class\
|
||
\ and confirmed in the CLI.\n3. Remove any ingress route (Caddy config repo) and DNS record still\n pointing at it.\n\
|
||
4. Verify backups of anything on it are retained per policy before the\n disk goes away (see `backs-up-to`).\n5. Destroy\
|
||
\ the LXC/VM (`pct destroy` / `qm destroy`).\n6. Move the `hosts.<name>:` block (if any inventory remnant survives\n \
|
||
\ `client remove`, e.g. infra-only LXCs with no age key) into\n inventory.yaml's `archaeology:` section: `pve_id`,\
|
||
\ `destroyed` date,\n `reason`. Add a row to `containers/index.md` \"Recently destroyed\"\n table (kept for human-readable\
|
||
\ browsing alongside the structured\n data).\n7. `oikos/ledger.py append host:<name> destroy destructive --result ok`.\n\
|
||
8. Regenerate: `python3 mcp/build_host_files.py && python3\n oikos/gen-topology.py` — the node drops out of `hosts/*.yaml`\
|
||
\ and\n appears in the topology doc's archaeology table.\n\nIf the destroy fails partway (e.g. secrets revoked but pct\
|
||
\ destroy\nerrors), do not re-run step 2 — `client remove` is not idempotent\nagainst a second revocation attempt on the\
|
||
\ issuance server. Finish the\nremaining steps manually and note the partial state in an investigation.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: lifecycle-migrate-node
|
||
name: 'Lifecycle: migrate a node'
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: lifecycle-migrate-node\nrisk_class: config_mutation\ninputs: [node_name, source_host, target_host]\n\
|
||
verification: \"homelab node <name> relations (re-check blast radius); homelab service <svc> health for every hosted service\"\
|
||
\ndocs_update_checklist: [doc_page_migration_note, inventory_host_and_lan_ip]\ntransition: \"active -> migrating -> active\"\
|
||
\n---\n\n# Lifecycle: migrate a node\n\nModeled on the strong Phase 1+2 migration\n([plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md](../../../.hermes/plans/2026-06-03_110000-library-ssd-migration-to-ludo-mini.md)).\n\
|
||
Requires (ontology): preflight + backup-verified before migrating;\npost-verify + Caddy backends checked + mounts checked\
|
||
\ + docs updated\nbefore returning to `active`.\n\n1. `homelab change preflight <every service the node hosts>` — capture\n\
|
||
\ current health as a baseline.\n2. Verify backups are current for anything with data at rest on the\n node (see `backs-up-to`\
|
||
\ edges once populated).\n3. Set `state: migrating` in `inventory.yaml`.\n4. Perform the migration (pct/qm move, or create-on-target\
|
||
\ +\n data-copy + destroy-source, per the specific case).\n5. Update `inventory.yaml`: new `host:`, `lan_ip`, `mesh`\
|
||
\ addresses for\n the node; update every `services:` entry whose `backend` pointed at\n it if the backend name itself\
|
||
\ changes (usually it doesn't — only the\n `host:`/`lan_ip` on the guest entry moves).\n6. Post-verify: re-run the Week-1\
|
||
\ drift check by hand — confirm Caddy's\n backend IP for each affected service matches the new `lan_ip`\n (automatic\
|
||
\ in Week 3's drift detector), confirm mounts still resolve.\n7. `homelab service <name> health` for every service the\
|
||
\ node hosts.\n8. Set `state: active`. Add a migration note to the node's doc page\n (old host/IP → new, date, phase\
|
||
\ reference) — this repo's convention\n for every past migration (see `containers/101-jellyfin.md`,\n `containers/129-house.md`).\n\
|
||
\nRegenerate: `python3 mcp/build_host_files.py && python3\noikos/gen-topology.py`.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: lifecycle-provision-node
|
||
name: 'Lifecycle: provision a node'
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: lifecycle-provision-node\nrisk_class: config_mutation\ninputs: [node_name, kind, storage_pool]\nverification:\
|
||
\ \"grep 'state: provisioning' hosts/<name>.yaml\"\ndocs_update_checklist: [doc_page_stub]\ntransition: \"planned -> provisioning\"\
|
||
\n---\n\n# Lifecycle: provision a node\n\nPer [oikos/ontology.yaml](../../../oikos/ontology.yaml) `lifecycle.transitions`.\n\
|
||
Policy note: `provisioning` nodes get a lifecycle override —\n`config_mutation` actions downgrade to `reversible_low`\
|
||
\ because nothing\ndepends on the node yet (see `oikos/policy.yaml` `lifecycle_overrides`).\n\nRequires (from ontology):\
|
||
\ inventory entry, IP reserved, storage pool\nchosen, doc page stub.\n\n1. Create the LXC/VM on its target Proxmox host\
|
||
\ (`pct create` /\n `qm create`), choosing the storage pool deliberately — record it as\n the `storage:` field once\
|
||
\ populated (Week 1 schema; not yet backfilled\n for existing nodes).\n2. Add the inventory entry: `homelab client add\
|
||
\ <name>` for anything that\n will run the `homelab` CLI, or a direct `hosts.<name>:` block with\n `state: provisioning`,\
|
||
\ `kind`, `host`, `pve_id`, `lan_ip` for\n infra-only LXCs that won't self-enroll.\n3. Stub the doc page (`containers/<pve_id>-<name>.md`\
|
||
\ or\n `vms/<pve_id>-<name>.md`) — even a one-line \"provisioning, see plan X\"\n is enough to satisfy the transition\
|
||
\ requirement.\n4. Reserve the IP in DNS/DHCP notes if it's a fixed LAN address.\n\nNext: [lifecycle-activate-node.md](../lifecycle-activate-node/SKILL.md).\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: runbook-budget-from-csv
|
||
name: 'Runbook: Budget import from N26 CSV → Yuvomi'
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: budget-import-from-csv\nrisk_class: config_mutation\ninputs: [csv_file]\nreferences: [containers/129-house.md]\n\
|
||
---\n\n# Runbook: Budget import from N26 CSV → Yuvomi\n\nDistil a bank-export CSV into Yuvomi's Budget and Subscriptions\
|
||
\ modules using\nthe `yuvomi-mcp` tools. Run this whenever a new CSV period needs to be\nsummarised into targets and fixed\
|
||
\ costs.\n\n---\n\n## Prerequisites\n\n- `yuvomi-mcp` is running on LXC 129 and connected as an MCP server in Claude.\n\
|
||
- The CSV is an N26 export (columns: Booking Date, Value Date, Partner Name,\n Partner Iban, Type, Payment Reference,\
|
||
\ Account Name, Amount (EUR), …).\n- API token: `homelab secret yuvomi-api-token` (decrypts on any enrolled client).\n\
|
||
- Direct API base: `https://house.hubris.network/api/v1`\n\n---\n\n## API quirks (Yuvomi ≤ 0.77.x)\n\n- **Subscriptions\
|
||
\ live under `/budget/subscriptions`**, NOT `/subscriptions/`.\n A top-level `/subscriptions` route returns 404.\n- `GET\
|
||
\ /budget/subscriptions` → `{ data: { subscriptions: [...], summary: {...} } }`\n- `GET /budget/subscriptions/meta` →\
|
||
\ `{ data: { categories: [...], payment_methods: [...] } }`\n- `POST /budget/subscriptions` → create a subscription (name,\
|
||
\ amount, billing_cycle,\n cycle_interval, next_payment_date, currency, category_id, payment_method_id required)\n- `GET\
|
||
\ /budget/` (no month) → returns only **non-recurring** base entries.\n Use `GET /budget/?month=YYYY-MM` to get all entries\
|
||
\ (recurring + one-time) for a month.\n- `GET /budget/categories` → expense category keys + income category names (German\
|
||
\ keys\n like `\"Erwerbseinkommen\"`, `\"Sozialleistungen\"`, `\"Geschenke & Transfers\"`).\n- Budget entries: `amount`\
|
||
\ positive = income, negative = expense.\n- Recurring entries: set `is_recurring: 1` + `recurrence_interval: \"monthly\"\
|
||
`.\n The `date` field sets the start month.\n- `recurrence_virtual: 1` smooths non-monthly amounts across all months\
|
||
\ in the summary\n (e.g. 55.08 € quarterly → shows as ~18.36 €/month).\n- Custom RRULE strings (`recurrence_rule`) are\
|
||
\ **not accepted** by the API — use\n `cycle_interval` on the subscription instead, or `recurrence_interval` on budget\
|
||
\ entries.\n\n---\n\n## Subscription category IDs (as of 2026-06-26)\n\n| id | name | budget_subcategory_key |\n|---|---|---|\n\
|
||
| 1 | Entertainment | subscription_entertainment |\n| 2 | Productivity | subscription_productivity |\n| 3 | Utilities\
|
||
\ | subscription_utilities |\n| 4 | Health | subscription_health |\n| 5 | Education | subscription_education |\n| 6 |\
|
||
\ Other | subscription_other |\n\n## Payment method IDs\n\n| id | name |\n|---|---|\n| 1 | Credit Card |\n| 2 | Debit\
|
||
\ Card |\n| 3 | PayPal |\n| 6 | Bank Transfer / SEPA |\n| 7 | Other |\n\n---\n\n## Step 1 — Categorise the transactions\n\
|
||
\nSkip these as internal/already-covered:\n- Fixed costs you'll enter as **subscriptions** (Miete, SWM, SYNVIA, Hundefutter,\n\
|
||
\ Netflix, Grover, Rundfunk ARD, KuKita, Lillydoo)\n- Internal transfers (The Joy Pot ↔ Cookie, Hauptkonto, Tagesgeldkonto\
|
||
\ splits)\n- Identified income (Cookie Share, Kindergeld, Pocket Money credits, Distributor)\n- Fun Money pass-throughs\
|
||
\ (in and out same month → net zero)\n\n**Variable expense taxonomy:**\n\n| Category key | Subcategory key | Examples\
|
||
\ |\n|---|---|---|\n| `food` | `groceries` | E-Center, Knuspr, EDEKA, Tegut, VollCorner, Lidl, Netto, REWE, KoRo, Roast\
|
||
\ Market |\n| `food` | `restaurants_bars` | Restaurants, Lieferando, Cafes, Zeit für Brot, Baobab, Höflinger |\n| `personal_health`\
|
||
\ | `beauty_cosmetics` | DM, Rossmann |\n| `personal_health` | `pharmacy` | Apotheke, MVZ Dermatologie |\n| `transport`\
|
||
\ | `apps_taxi` | Uber, RYD GMBH, MVG, Handyparken |\n| `shopping_clothing` | `gifts` | Children products: Schlummersack,\
|
||
\ Catchy Kids, SP EVERY., Dukal, Berger-Lernwelt |\n| `shopping_clothing` | `clothes_shoes` | Zalando, Ernsting's, Schuhmair,\
|
||
\ Thalia, Vinted, Airbnb, Hotel at Booking.com |\n| `shopping_clothing` | `electronics` | Amazon, AMZN Mktp DE |\n| `housing`\
|
||
\ | `renovation_maintenance` | IKEA, Markus Festl, Granit, Sostrene Grene, Mol* tischdecken, Gaertnerei, Dehner |\n| `education`\
|
||
\ | `courses_college` | Kathrin Orlob (PEKiP), Nerina Aupperle |\n| `leisure` | `streaming` | WOW wowtv.de |\n| `financial_other`\
|
||
\ | `bank_fees` | Unidentified PayPal, Ratepay, N26 fees |\n| `Geschenke & Transfers` | *(income)* | One-off incoming\
|
||
\ transfers |\n\n---\n\n## Step 2 — Create subscriptions\n\n```\nget_subscriptions_meta() ← get category_id and payment_method_id\n\
|
||
```\n\n**Standard Cookie household subscriptions (as of 2026-07):**\n\n| Name | Amount | billing_cycle | cycle_interval\
|
||
\ | category_id | payment_method_id |\n|---|---|---|---|---|---|\n| Miete | 1080.00 | monthly | 1 | 6 (Other) | 6 (Bank\
|
||
\ Transfer) |\n| Strom (SWM) | 79.00 | monthly | 1 | 3 (Utilities) | 6 |\n| Internet / TV / Telefon | 29.99 | monthly\
|
||
\ | 1 | 3 (Utilities) | 6 |\n| Hundefutter | 75.00 | monthly | 1 | 6 (Other) | 6 |\n| Netflix | 8.00 | monthly | 1 | 1\
|
||
\ (Entertainment) | 6 |\n| Grover | 16.90 | monthly | 1 | 6 (Other) | 2 (Debit Card) |\n| Rundfunk ARD / ZDF | 55.08 |\
|
||
\ monthly | 3 | 1 (Entertainment) | 6 |\n| KuKita Daycare (Leon) | 503.00 | monthly | 1 | 5 (Education) | 6 |\n| Lillydoo\
|
||
\ diapers | 56.70 | monthly | 2 | 4 (Health) | 3 (PayPal) |\n\nMonthly equivalent total: **1,838.60 €** (Yuvomi applies\
|
||
\ cycle_interval to prorate).\n\n---\n\n## Step 3 — Add recurring income entries\n\n```\nstage_add_budget_entry(\n \
|
||
\ title=\"Kindergeld\",\n amount=55.00,\n category=\"Sozialleistungen\",\n date=\"YYYY-MM-01\",\n is_recurring=True,\n\
|
||
\ recurrence_interval=\"monthly\",\n)\ncommit_pending(pending_id)\n```\n\n**Standard recurring income:**\n\n| Title\
|
||
\ | Amount | category |\n|---|---|---|\n| Kindergeld | +55.00 | Sozialleistungen |\n| Cookie Share | +2650.00 | Erwerbseinkommen\
|
||
\ *(see recommended amount below)* |\n\n---\n\n## Step 4 — Post variable transactions\n\nFor each non-skipped CSV row,\
|
||
\ call `stage_add_budget_entry` with the mapped\ncategory/subcategory and the actual transaction amount and date. Use\
|
||
\ the Partner\nName + Payment Reference as the title (truncate to 100 chars).\n\n---\n\n## Step 5 — Verify\n\n```\nget_budget_summary(\"\
|
||
YYYY-MM\")\nlist_subscriptions()\n```\n\nExpected for a full month with KuKita:\n- Fixed expenses ≥ 1,838 € (subscriptions)\n\
|
||
- Variable expenses ≥ 500 € (groceries alone)\n\n---\n\n## Cookie Share: how much to transfer monthly\n\nCalculated from\
|
||
\ Jan–Jun 2026 data (Cookie account, one-offs stripped):\n\n| | €/month |\n|---|---|\n| **Fixed costs (subscriptions)**\
|
||
\ | **1,839** |\n| Miete | 1,080 |\n| KuKita *(permanent from Jul 2026)* | 503 |\n| Strom + SYNVIA + Rundfunk + Netflix\
|
||
\ + Grover + Hundefutter + Lillydoo | 256 |\n| **Variable (6-month averages)** | **1,032** |\n| Groceries | 595 |\n| Children\
|
||
\ products | 142 |\n| Dining & cafes | 100 |\n| Transport | 66 |\n| Drugstore | 52 |\n| Clothing, Amazon, Pharmacy | 77\
|
||
\ |\n| **Total monthly spend** | **≈ 2,871** |\n| Minus Kindergeld (fixed income) | −55 |\n| Minus Pocket Money (conservative\
|
||
\ ~600 €) | −600 |\n| **→ Recommended Cookie Share** | **≈ 2,650 €** |\n| With 200 € buffer | **≈ 2,850 €** |\n\n**Current\
|
||
\ Cookie Share (Jun 2026): 1,995 € — shortfall ~655 €.**\n\nThe gap was covered by irregular Pocket Money top-ups (avg\
|
||
\ 962 €/mo over 6 months, but\nhighly variable: 121 €–3,000 €). KuKita starting in June is the biggest step-up; raising\n\
|
||
Cookie Share to **2,650 €** makes the budget self-sufficient without relying on top-ups.\n\n---\n\n## Changelog\n\n###\
|
||
\ 2026-06-29 — Corrections from first real import\n- Subscriptions endpoint is `/budget/subscriptions`, NOT `/subscriptions/`\
|
||
\ (404).\n- `recurrence_rule` RRULE strings are rejected by the API; use `cycle_interval` instead.\n- `GET /budget/` (no\
|
||
\ filter) returns only non-recurring entries; use `?month=` for full view.\n- Added Cookie Share recommendation (2,650\
|
||
\ €/month) based on 6-month expense analysis.\n- Added full category taxonomy table.\n\n### 2026-06-29 — Initial runbook\n\
|
||
Created from Jan–Jun 2026 N26 Cookie account analysis.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: runbook-dpkg-interrupted
|
||
name: Runbook — recover from dpkg-interrupted state
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: recover-dpkg-interrupted\nrisk_class: reversible_low\nverification: \"dpkg --audit (should be clean);\
|
||
\ apt-get check\"\n---\n\n# Runbook — recover from dpkg-interrupted state\n\nYou're here because an apt run got killed\
|
||
\ mid-transaction and the target now\nhas packages that are **unpacked but not configured**. Symptoms:\n\n- `apt` refuses\
|
||
\ to do anything new: `Error: dpkg was interrupted, you must\n manually run 'dpkg --configure -a' to correct the problem.`\n\
|
||
- `dpkg --audit` lists packages with header\n `The following packages have been unpacked but not yet configured.`\n-\
|
||
\ `homelab apt-audit` shows `DPKG: DIRTY(N)` for the host.\n\nThe system is still running the **old** binaries (still\
|
||
\ in memory), but the\n**new** binaries are unpacked and waiting for their postinst to run. Two\nworst-case manifestations\
|
||
\ from the 2026-05-21 sweep:\n\n- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was\n still\
|
||
\ serving but the new caddy binary on disk hadn't been wired up.\n- hubris: ssh master died mid-Wave-6 → 135 packages\
|
||
\ unpacked-not-configured,\n including `systemd`, `openssh-server`, `sudo`, `netbird`. The half-\n configured netbird\
|
||
\ daemon dropped the mesh peer, and we got locked out\n until we recovered from the PVE web UI Shell.\n\n**Do not reboot\
|
||
\ until dpkg is clean.** A reboot tries to start the new\nbinaries' services, which may fail because postinst never ran\
|
||
\ (missing users,\nconfig dirs, capabilities, etc.). The system might not come back up cleanly.\n\n## Path A — target\
|
||
\ is still reachable over ssh (preferred)\n\n```\nhomelab ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure\
|
||
\ -a && apt -y -o Dpkg::Options::=--force-confold upgrade'\n```\n\nOr for an LXC by name:\n\n```\nhomelab pct <lxc> exec\
|
||
\ -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'\n\
|
||
```\n\nWhen that returns, confirm:\n\n```\nhomelab apt-audit --target <host>\n```\n\nExpect `DPKG: ok` and the remaining\
|
||
\ `UPGR` count to match what's intentionally\ndeferred (kernel/PVE on hubris, 0 elsewhere).\n\n## Path B — target locked\
|
||
\ out (mesh broken / ssh dead)\n\nMost common for hubris when netbird itself went half-configured: the daemon\ncrashed\
|
||
\ on the new binary, the mesh peer dropped, port 22022 stopped listening,\nand you can't ssh in.\n\n1. Open `https://proxmox.hubris.network`\
|
||
\ in a browser.\n2. Datacenter → node `hubris` → `>_ Shell` (or `_ Console`). That's a root\n shell on hubris served\
|
||
\ by the PVE web UI, independent of the netbird mesh.\n3. Run the recovery one-liner:\n\n```\nDEBIAN_FRONTEND=noninteractive\
|
||
\ dpkg --configure -a \\\n && DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \\\n \
|
||
\ && systemctl restart netbird \\\n && dpkg --audit \\\n && echo RECOVERY_OK\n```\n\n Wait for `RECOVERY_OK`. The\
|
||
\ `systemctl restart netbird` is the bit that\n heals the mesh — once netbird's daemon comes back up clean, your client's\n\
|
||
\ peer state moves from `Connecting` to `Connected` within ~30 seconds and\n the rest of your tooling works again.\n\
|
||
\n4. For an **LXC** that's locked out (less common — LXCs reach the world via\n netbird routed through hubris, so unless\
|
||
\ hubris itself is broken, you can\n still `pct enter` from the hubris shell):\n\n From the PVE web UI shell on hubris:\n\
|
||
\n ```\n pct enter <id>\n DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade\n exit\n ```\n\
|
||
\n## Prevention\n\nThe `homelab apt-upgrade` wrapper launches apt inside a `systemd-run --collect`\nunit on the target,\
|
||
\ so it survives ssh teardown — the failure mode that put\nhubris into this state in the first place is no longer reachable\
|
||
\ through the\nstandard tool. If you absolutely need to run apt manually over ssh, wrap it:\n\n```\nssh <host> systemd-run\
|
||
\ --unit=apt-recovery --collect bash -c 'apt -y upgrade'\n```\n\nThen `systemctl status apt-recovery` from a fresh ssh\
|
||
\ to check progress.\n\n## Related\n\n- [Operations cheatsheet](../../operations/commands.md)\n- [Auto-deploy pipelines](../../../knowledge/wiki/infrastructure/auto-deploy.md)\n\
|
||
- [Hubris host page](../../../knowledge/wiki/hosts/hubris.md)\n\n## Changelog\n\n### 2026-05-21 — initial page\nDocuments\
|
||
\ the dpkg-interrupted recovery path that came out of the\nfleet apt sweep (Wave 6 killed mid-transaction; hubris recovered\
|
||
\ via PVE\nweb Shell).\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|
||
- slug: service-health-check
|
||
name: Service health check
|
||
risk_class: read_only
|
||
entity_type: service
|
||
procedure: {}
|
||
content: "---\nname: service-health-check\nrisk_class: read_only\ninputs: [service_name]\nverification: \"homelab service\
|
||
\ <name> health\"\ndocs_update_checklist: []\n---\n\n# Service health check\n\nGoal: determine whether a service is actually\
|
||
\ healthy, without ad-hoc SSH.\n\n1. `homelab service <name> explain` — read the context card: backend,\n blast radius,\
|
||
\ doc pointer, risk notes.\n2. `homelab service <name> health` — live health probe (HTTP code against\n the service's\
|
||
\ `url`/`endpoint`). Once the Week-3 scheduler ships, this\n reads a cached snapshot by default; pass `--live` to force\
|
||
\ a fresh probe.\n3. If unhealthy, `homelab service <name> log` (or MCP `tail_log`) for the\n last 200 lines.\n4. Cross-check\
|
||
\ blast radius: `homelab node <name> relations` — is this\n entity's own backend host healthy? A downstream failure\
|
||
\ (e.g. `strong`\n down) will show up here before the service's own logs explain anything.\n5. If the fix is a restart:\
|
||
\ classify first (`oikos/policy.yaml` —\n `service-restart` is `reversible_low` unless the service has a\n `service_overrides`\
|
||
\ entry, e.g. `caddy`/`dns` are `config_mutation`).\n Unattended agents may act on `reversible_low` without approval.\n\
|
||
\nDocs-update checklist: none for a pure health check. If the investigation\nreveals stale `risk_notes` or a wrong `doc_page`,\
|
||
\ fix `inventory.yaml` in\nthe same session.\n"
|
||
tags:
|
||
- skill
|
||
- runbook
|