Updates to four pages reflecting the combined → vanilla mgmt+signal+relay+coturn cutover and the IONOS-3478-firewall-exception discovery: * infrastructure/mesh.md — rewrites the ICE/STUN section to cover the new TURN endpoint, the IONOS upstream TCP-3478 filtering (load-bearing, undocumented before today), and the verification probe. New changelog entry covering the migration outcome + Device Code Stage gap. * infrastructure/vps-hardening.md — "At a glance" lists the new 6-service docker stack + host coturn. Firewall section notes the new `iifname ens6 tcp dport 3478 accept` rule plus the IONOS upstream exception. New changelog entry. * containers/124-authentik.md — replaces the "Netbird IdP integration — DEFERRED" section with the LANDED state: Provider details (Public client type — Confidential breaks PKCE on the dashboard SPA), the first-time owner-promotion sqlite recipe, the missing Device Code Stage gap + workaround (setup-keys), and a note that the old 2026-04-22 pre-work Provider/App is now obsolete and safe to delete. Updated changelog (Phase 6 landed). * operations/agent-enrollment.md — new "Getting onto Netbird" subsection explaining the setup-key path (currently the only working flow until Device Code Stage lands) and why direct OIDC from the public internet fails (auth.hubris.network is mesh-only-reachable). Prerequisites table row updated to point at the new section. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
11 KiB
Mesh — Tailscale → Netbird migration
The hubris fleet is migrating from Tailscale to Netbird. Netbird is the target end-state. In-progress as of 2026-04-21.
Current state
- PVE host uses Netbird (
wt0,100.122.38.109/16). Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets*.hubris.network → 192.168.8.175via the system resolver chain. - Netbird mgmt host (
82.165.190.79, FQDNinspiring-ramanujan.netbird.selfhosted, NB IP100.122.165.149) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes the homelab network (192.168.8.0/24) via the PVE peer. This gives the mgmt host LAN access and split-horizon DNS for*.hubris.network. Useful independently of any Authentik integration. - Most LXCs still run Tailscale or use router DNS (
192.168.8.1) / Tailscale MagicDNS (100.100.100.100), both of which return the public IONOS A record*.hubris.network → 82.165.190.79. The VPS only routes hostnames it actually publishes (today,artifacto+blog), so this path is a dead end for any LAN-only service.
ICE / STUN / TURN
Today (post-2026-05-21 migration):
- External STUN servers (Google + Cloudflare) declared under
server.Stunsin/opt/management.json. Embedded STUN is no longer in use. - coturn runs on the VPS host (apt-installed, systemd-managed). TURN URL
turn:netbird.hubris.network:3478?transport=tcpis advertised to peers via mgmt'sTURNConfig.Turnsblock. Long-term credentials atuser=netbird:<pwd from /root/turn-pass.txt>. - For non-symmetric peers, ICE picks direct
srflx/srflx(P2P). For symmetric-NAT peers, ICE falls through to TURN-relay before falling back to the WSS relay (rels://netbird.hubris.network:443).
IONOS port-3478 caveat (load-bearing — undocumented before 2026-05-21):
IONOS upstream filters STUN-class traffic on port 3478 for both UDP and TCP by default. The UDP block was already known (verified 2026-05-10 with tcpdump -ni any udp port 3478). The TCP block was discovered 2026-05-21 when TURN-allocate retries timed out from a symmetric-NAT laptop: TCP handshake completed (carried by kernel-only SYN exchange) but data packets never reached the VPS's ens6 interface.
Resolution: operator must add an inbound firewall exception in the IONOS dashboard (Cloud Panel → Networks → Firewall policy on the VPS) for inbound TCP 3478 (and UDP 3478 if STUN-via-VPS-listener is ever needed; we use external STUN so this is optional). Without that exception, coturn is unreachable from outside even though host nftables + listener look fine.
Verifying TURN works end-to-end from an outside peer:
# python3
import socket, struct, secrets
s = socket.create_connection(("netbird.hubris.network", 3478), timeout=10)
tid = secrets.token_bytes(12)
attrs = struct.pack("!HHI", 0x0019, 4, 17 << 24) # REQUESTED-TRANSPORT = UDP
msg = struct.pack("!HHI", 0x0003, len(attrs), 0x2112A442) + tid + attrs # Allocate
s.sendall(msg)
print(s.recv(4096).hex()) # expect 120 bytes starting 0x0113 (Allocate Error = auth challenge)
A 120-byte 0x0113-typed response means coturn is reachable and responding. An indefinite TCP timeout (after connected from ... prints) means the IONOS rule was reverted or narrowed.
If a peer is still on the WSS relay after this — verify in netbird status -d: Connection type: P2P with srflx/srflx is the cone-NAT happy path; Connection type: P2P with relay/srflx (or relay/prflx) means the peer's network is symmetric-NAT and it's using TURN. Falling back to Relayed (rels://...) only happens if TURN allocate also fails — investigate the IONOS exception first.
Old combined-server note (history, kept for context):
Pre-migration, the bundled netbirdio/netbird-server combined image silently ignored both server.turns: and top-level TURNConfig: YAML, so TURN was non-functional. That's why the architectural migration to the canonical multi-container stack (mgmt + signal + relay + coturn) happened. See the 2026-05-21 changelog entry below and the Netbird-combined-no-TURN finding in homelab memory for the full discovery.
Consequence — every LXC wired to Authentik needs an internal override
Until each LXC is migrated to Netbird, anything that needs to reach auth.hubris.network (Authentik), cloud.hubris.network (Nextcloud), etc., must override the public answer with 192.168.8.175.
Two techniques. Pick by HTTP-client behavior.
A) /etc/hosts override
Works for libc getaddrinfo clients: curl, wget, most Go/Python/Ruby apps, gitea.
- Place the line outside the
# --- BEGIN PVE ---/# --- END PVE ---markers. Proxmox rewrites everything inside that block on every container start. - Belt-and-suspenders:
/etc/systemd/system/hubris-hosts-override.service(oneshot, enabled, idempotent).
B) Local dnsmasq
Required for clients that bypass /etc/hosts. Nextcloud (PHP Guzzle + OC\Http\Client\DnsPinMiddleware) is one — uses dns_get_record(), not getaddrinfo. Likely candidates for the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting.
Recipe:
apt install dnsmasq
cat > /etc/dnsmasq.d/hubris-internal.conf <<EOF
address=/auth.hubris.network/192.168.8.175
server=192.168.8.1
server=1.1.1.1
interface=lo
bind-interfaces
no-hosts
no-resolv
EOF
# Set LXC default nameservers and live resolv.conf
pct set <id> --nameserver "127.0.0.1 192.168.8.1 1.1.1.1"
# Then update /etc/resolv.conf inside the LXC too.
Known overrides applied
| LXC | Technique | Notes |
|---|---|---|
| 104 (gitea) | /etc/hosts + hubris-hosts-override.service |
Standard |
| 114 (nextcloud) | local dnsmasq @ 127.0.0.1 + resolv.conf reorder | Guzzle bypasses /etc/hosts |
| 105 (apps), inner containers | extra_hosts: in compose |
Booklore, mulita, WriteFreely each ship with this |
Adding new LXCs
- Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris.
- When wiring a new app into Authentik:
cat /etc/resolv.confon the target LXC. If nameserver is192.168.8.1or100.100.100.100, add the hosts override. If it's the netbird daemon IP, skip.
Long-term fix
Either:
- Split-horizon DNS at LAN/router level so
*.hubris.network → 192.168.8.175for every LAN client. Eliminates all per-LXC overrides. - Once Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain on the PVE host.
CRITICAL — never docker compose up Portainer-managed stacks
apps (105) runs multiple stacks deployed via the Portainer UI (/var/lib/docker/volumes/portainer_data/_data/compose/<N>/). Running docker compose up -d <svc> from the host shell on these triggers a recreate of OTHER services in the stack — labels drift, compose treats the project as reconciling, containers get destroyed and remade. This wiped Booklore's mariadb data on 2026-04-22 (bind mount ./mariadb/config re-initialized fresh).
Recipe for container-config changes (e.g. adding extra_hosts) on Portainer-managed stacks:
- Edit the compose in Portainer UI → Stacks → → Editor → Update stack. Portainer handles the recreate cleanly with its own state tracking.
- Do NOT edit
/var/lib/docker/volumes/portainer_data/_data/compose/<N>/docker-compose.ymldirectly. - For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the app's built-in OIDC over container-level
extra_hostswhen possible.
Related
- DNS split-horizon
- Authentik (124) — the IdP that triggers most of these overrides
- Nextcloud (114) — example of Technique B
- Gitea (104) — example of Technique A
- Public ingress (VPS traefik) — uses the same mesh as transport
Changelog
2026-05-21 — VPS migrated combined → vanilla netbird stack with external TURN
The combined netbirdio/netbird-server image was replaced with the canonical multi-container deploy (netbirdio/management:0.71.3 + signal:0.71.3 + relay:0.71.3 + dashboard:latest + host coturn) on /opt/docker-compose.yml. Driver: combined image silently ignored external TURNConfig so symmetric-NAT peers couldn't use TURN.
Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on LXC 124, upgrading mgmt to 0.71.3. The store.db schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old opt_netbird_data volume into the new mgmt_data volume). Pre-cutover backups at /root/netbird-*.tgz on the VPS, ~857 MB, retained for ~7d.
Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478 inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe).
The new Authentik provider for NetBird is Client type: Public (PKCE-only). Confidential would break the dashboard SPA's token exchange. Device Code Stage is NOT yet configured in Authentik → netbird up interactive auth flow returns an empty consent screen; new peers must use --setup-key until the stage is added.
Open follow-up: TURN-over-TLS on TCP 5349 (cert via certbot or extract Traefik's acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient for current usage.
2026-05-10 — ICE direct p2p restored (external STUN swap)
All peers were Connection type: Relayed because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in /opt/config.yaml. After netbird down/up, peers now report P2P with srflx/host candidates. Backup of pre-change config at /opt/config.yaml.bak-20260510-185050. Fixes a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers).
2026-04-28 — wiki entry created
Initial documentation.
2026-04-22 — Booklore mariadb data wiped (lesson recorded)
The "never docker compose up Portainer-managed stacks" rule comes from this. See apps (105).
2026-04-22 — netbird mgmt host joined its own mesh
82.165.190.79 is now a peer (100.122.165.149). LAN access + split-horizon DNS via PVE peer. See VPS hardening.
2026-04-21 — overrides applied to gitea (104) and nextcloud (114)
Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle.