- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
16 KiB
Mesh — Tailscale → Netbird migration
The hubris fleet is migrating from Tailscale to Netbird. Netbird is the target end-state. In-progress as of 2026-04-21.
Current state
- PVE host uses Netbird (
wt0,100.122.38.109/16). Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets*.hubris.network → 192.168.8.175via the system resolver chain. - Netbird mgmt host (
82.165.190.79, FQDNinspiring-ramanujan.netbird.selfhosted, NB IP100.122.165.149) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes the homelab network (192.168.8.0/24) via the PVE peer. This gives the mgmt host LAN access and split-horizon DNS for*.hubris.network. Useful independently of any Authentik integration. - Most LXCs still run Tailscale or use router DNS (
192.168.8.1) / Tailscale MagicDNS (100.100.100.100), both of which return the public IONOS A record*.hubris.network → 82.165.190.79. The VPS only routes hostnames it actually publishes (today,artifacto+blog), so this path is a dead end for any LAN-only service.
ICE / STUN / TURN
Today (post-2026-05-21 migration):
- External STUN servers (Google + Cloudflare) declared under
server.Stunsin/opt/management.json. Embedded STUN is no longer in use. - coturn runs on the VPS host (apt-installed, systemd-managed). TURN URL
turn:netbird.hubris.network:3478?transport=tcpis advertised to peers via mgmt'sTURNConfig.Turnsblock. Long-term credentials atuser=netbird:<pwd from /root/turn-pass.txt>. - For non-symmetric peers, ICE picks direct
srflx/srflx(P2P). For symmetric-NAT peers, ICE falls through to TURN-relay before falling back to the WSS relay (rels://netbird.hubris.network:443).
IONOS port-3478 caveat (load-bearing — undocumented before 2026-05-21):
IONOS upstream filters STUN-class traffic on port 3478 for both UDP and TCP by default. The UDP block was already known (verified 2026-05-10 with tcpdump -ni any udp port 3478). The TCP block was discovered 2026-05-21 when TURN-allocate retries timed out from a symmetric-NAT laptop: TCP handshake completed (carried by kernel-only SYN exchange) but data packets never reached the VPS's ens6 interface.
Resolution: operator must add an inbound firewall exception in the IONOS dashboard (Cloud Panel → Networks → Firewall policy on the VPS) for inbound TCP 3478 (and UDP 3478 if STUN-via-VPS-listener is ever needed; we use external STUN so this is optional). Without that exception, coturn is unreachable from outside even though host nftables + listener look fine.
Verifying TURN works end-to-end from an outside peer:
# python3
import socket, struct, secrets
s = socket.create_connection(("netbird.hubris.network", 3478), timeout=10)
tid = secrets.token_bytes(12)
attrs = struct.pack("!HHI", 0x0019, 4, 17 << 24) # REQUESTED-TRANSPORT = UDP
msg = struct.pack("!HHI", 0x0003, len(attrs), 0x2112A442) + tid + attrs # Allocate
s.sendall(msg)
print(s.recv(4096).hex()) # expect 120 bytes starting 0x0113 (Allocate Error = auth challenge)
A 120-byte 0x0113-typed response means coturn is reachable and responding. An indefinite TCP timeout (after connected from ... prints) means the IONOS rule was reverted or narrowed.
If a peer is still on the WSS relay after this — verify in netbird status -d: Connection type: P2P with srflx/srflx is the cone-NAT happy path; Connection type: P2P with relay/srflx (or relay/prflx) means the peer's network is symmetric-NAT and it's using TURN. Falling back to Relayed (rels://...) only happens if TURN allocate also fails — investigate the IONOS exception first.
Old combined-server note (history, kept for context):
Pre-migration, the bundled netbirdio/netbird-server combined image silently ignored both server.turns: and top-level TURNConfig: YAML, so TURN was non-functional. That's why the architectural migration to the canonical multi-container stack (mgmt + signal + relay + coturn) happened. See the 2026-05-21 changelog entry below and the Netbird-combined-no-TURN finding in homelab memory for the full discovery.
Consequence — every LXC wired to Authentik needs an internal override
Until each LXC is migrated to Netbird, anything that needs to reach auth.hubris.network (Authentik), cloud.hubris.network (Nextcloud), etc., must override the public answer with 192.168.8.175.
Two techniques. Pick by HTTP-client behavior.
A) /etc/hosts override
Works for libc getaddrinfo clients: curl, wget, most Go/Python/Ruby apps, gitea.
- Place the line outside the
# --- BEGIN PVE ---/# --- END PVE ---markers. Proxmox rewrites everything inside that block on every container start. - Belt-and-suspenders:
/etc/systemd/system/hubris-hosts-override.service(oneshot, enabled, idempotent).
B) Local dnsmasq
Required for clients that bypass /etc/hosts. Nextcloud (PHP Guzzle + OC\Http\Client\DnsPinMiddleware) is one — uses dns_get_record(), not getaddrinfo. Likely candidates for the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting.
Recipe:
apt install dnsmasq
cat > /etc/dnsmasq.d/hubris-internal.conf <<EOF
address=/auth.hubris.network/192.168.8.175
server=192.168.8.1
server=1.1.1.1
interface=lo
bind-interfaces
no-hosts
no-resolv
EOF
# Set LXC default nameservers and live resolv.conf
pct set <id> --nameserver "127.0.0.1 192.168.8.1 1.1.1.1"
# Then update /etc/resolv.conf inside the LXC too.
Known overrides applied
| LXC | Technique | Notes |
|---|---|---|
| 104 (gitea) | /etc/hosts + hubris-hosts-override.service |
Standard |
| 114 (nextcloud) | local dnsmasq @ 127.0.0.1 + resolv.conf reorder | Guzzle bypasses /etc/hosts |
| 105 (apps), inner containers | extra_hosts: in compose |
Booklore, mulita, WriteFreely each ship with this |
Adding new LXCs
- Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris.
- When wiring a new app into Authentik:
cat /etc/resolv.confon the target LXC. If nameserver is192.168.8.1or100.100.100.100, add the hosts override. If it's the netbird daemon IP, skip.
Long-term fix
Either:
- Split-horizon DNS at LAN/router level so
*.hubris.network → 192.168.8.175for every LAN client. Eliminates all per-LXC overrides. - Once Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain on the PVE host.
CRITICAL — never docker compose up Portainer-managed stacks
apps (105) runs multiple stacks deployed via the Portainer UI (/var/lib/docker/volumes/portainer_data/_data/compose/<N>/). Running docker compose up -d <svc> from the host shell on these triggers a recreate of OTHER services in the stack — labels drift, compose treats the project as reconciling, containers get destroyed and remade. This wiped Booklore's mariadb data on 2026-04-22 (bind mount ./mariadb/config re-initialized fresh).
Recipe for container-config changes (e.g. adding extra_hosts) on Portainer-managed stacks:
- Edit the compose in Portainer UI → Stacks → → Editor → Update stack. Portainer handles the recreate cleanly with its own state tracking.
- Do NOT edit
/var/lib/docker/volumes/portainer_data/_data/compose/<N>/docker-compose.ymldirectly. - For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the app's built-in OIDC over container-level
extra_hostswhen possible.
Related
- DNS split-horizon
- Authentik (124) — the IdP that triggers most of these overrides
- Nextcloud (114) — example of Technique B
- Gitea (104) — example of Technique A
- Public ingress (VPS traefik) — uses the same mesh as transport
Changelog
2026-05-31 (later) — Authentik moved to the VPS; mesh-dependency for auth eliminated (supersedes the band-aid below)
The earlier same-day fix routed auth.hubris.network through VPS Traefik → Caddy → LXC 124 over the mesh. That restored service but re-created the original fragility: if the mesh is dark when management restarts, the 192.168.8.175 backend is unreachable and management crash-loops again (the "Bootstrap note" in the entry below). That note is now obsolete — Authentik was migrated onto the VPS itself, so OIDC no longer touches the mesh. The auth-authentik → 192.168.8.175 route and its skip-verify transport were removed from /opt/traefik-dynamic.yaml; auth.hubris.network is now served by a local authentik-server container via Traefik Docker-provider labels, and netbird-mgmt has depends_on: authentik-server: condition: service_healthy. The socat / reverse-SSH bootstrap dance is no longer needed. Full detail: 2026-05-31 Authentik VPS migration.
2026-05-31 — Netbird mesh recovered; auth.hubris.network exposed via VPS Traefik
Symptom: netbird-mgmt crash-looped for ~9 days (since 2026-05-21 migration). All peers showed Connecting, management returned 404 (Not Found) for gRPC → EOF on startup.
Root cause: The 2026-05-21 migration configured management.json with OIDCConfigEndpoint: https://auth.hubris.network/.... The public DNS for auth.hubris.network already pointed to 82.165.190.79 (VPS), but the VPS Traefik had no route for that host → management got EOF on every boot, crash-looped.
Fix:
- Added
auth-hubrisrouter to/opt/traefik-dynamic.yaml—Host(auth.hubris.network)withcertResolver: letsencrypt→ serviceauth-authentik. - Service backend:
https://192.168.8.175(Caddy on hubris LAN) via the Netbird mesh. - ServersTransport
skip-verifywithserverName: auth.hubris.network+insecureSkipVerify: trueso Traefik sends the correct SNI to Caddy. - Upgraded VPS Netbird client
0.69.0 → 0.71.3via apt (apt install netbird=0.71.3). - Mesh fully recovered; management connected to peers within ~1 min.
Bootstrap note (if mesh is dark and management must restart): If the WireGuard tunnels are fully dead AND management needs to restart (after e.g. a VPS reboot with no peer handshakes), the Traefik backend 192.168.8.175 will be unreachable and management will crash again. Recovery: temporarily open VPS port 22 via IONOS console (nft insert rule inet hubris-fw input iifname "ens6" tcp dport 22 accept), SSH in, run socat TCP-LISTEN:8443,bind=172.30.0.1,reuseaddr,fork TCP:127.0.0.1:8443 &, then from hubris ssh -f -N -R 127.0.0.1:8443:192.168.8.175:443 root@82.165.190.79 — this bootstraps one management start, after which the mesh self-heals.
Architecture after this change: auth.hubris.network is publicly accessible (HTTPS via VPS Traefik → Caddy on hubris → Authentik LXC 124). External devices authenticating to Netbird now hit this public path. Phase 6 (Authentik as Netbird IdP) is complete and live.
2026-05-21 — VPS migrated combined → vanilla netbird stack with external TURN
The combined netbirdio/netbird-server image was replaced with the canonical multi-container deploy (netbirdio/management:0.71.3 + signal:0.71.3 + relay:0.71.3 + dashboard:latest + host coturn) on /opt/docker-compose.yml. Driver: combined image silently ignored external TURNConfig so symmetric-NAT peers couldn't use TURN.
Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on LXC 124, upgrading mgmt to 0.71.3. The store.db schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old opt_netbird_data volume into the new mgmt_data volume). Pre-cutover backups at /root/netbird-*.tgz on the VPS, ~857 MB, retained for ~7d.
Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478 inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe).
The new Authentik provider for NetBird is Client type: Public (PKCE-only). Confidential would break the dashboard SPA's token exchange. The Device Code grant flow is wired (see containers/124-authentik.md) so interactive netbird up works — --setup-key is no longer required for new peers.
Post-migration JWT-issuer gotcha on existing peers (cost ~30 min to diagnose 2026-05-21):
Existing peers — registered against the old combined image's embedded Dex IdP at https://netbird.hubris.network/oauth2 — cache the OLD expected SSH-JWT issuer in the netbird daemon's in-memory state. After the migration, incoming netbird ssh connections were rejected with:
JWT authentication failed: validate token (
expected issuer=https://netbird.hubris.network/oauth2,
audiences=[netbird-dashboard netbird-cli],
actual issuer=https://auth.hubris.network/application/o/netbird/,
audience=netbird-dashboard
)
Neither systemctl restart netbird nor netbird down && netbird up clears the cache. Root cause: in client/internal/engine_ssh.go, updateSSH() bails out with if e.sshServer != nil { return nil } whenever the SSH server is already running, so mgmt-pushed JWT config updates are silently ignored. Only a full daemon-process tear-down lets the SSH server re-initialize with the new validator config:
sudo systemctl stop netbird
sleep 3
sudo systemctl start netbird
After that, grep -iE "issuer|audience" /var/log/netbird/client.log | tail shows the new Authentik issuer. Run this on every existing peer (PVE host + every LXC + every workstation) once after a future IdP swap.
Username gotcha (related): netbird ssh defaults the remote username to the LOCAL one (e.g. dtoro from the operator's laptop). Hubris + the LXCs only have root, so the JWT is accepted but the session immediately fails with user dtoro not found. Always use the explicit root@ prefix when invoking netbird-ssh manually:
netbird ssh -p 22022 root@proxmox-server.netbird.selfhosted
The homelab CLI handles this automatically via the per-host ssh.user field in inventory.yaml (defaults to root; set explicitly only for workstations whose login user isn't root).
Open follow-up: TURN-over-TLS on TCP 5349 (cert via certbot or extract Traefik's acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient for current usage.
2026-05-10 — ICE direct p2p restored (external STUN swap)
All peers were Connection type: Relayed because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in /opt/config.yaml. After netbird down/up, peers now report P2P with srflx/host candidates. Backup of pre-change config at /opt/config.yaml.bak-20260510-185050. Fixes a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers).
2026-04-28 — wiki entry created
Initial documentation.
2026-04-22 — Booklore mariadb data wiped (lesson recorded)
The "never docker compose up Portainer-managed stacks" rule comes from this. See apps (105).
2026-04-22 — netbird mgmt host joined its own mesh
82.165.190.79 is now a peer (100.122.165.149). LAN access + split-horizon DNS via PVE peer. See VPS hardening.
2026-04-21 — overrides applied to gitea (104) and nextcloud (114)
Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle.