Files
oikos/infrastructure/mesh.md
dtoro b42a986cc0 wiki: document 2026-05-21 netbird vanilla migration
Updates to four pages reflecting the combined → vanilla mgmt+signal+relay+coturn
cutover and the IONOS-3478-firewall-exception discovery:

* infrastructure/mesh.md — rewrites the ICE/STUN section to cover the new
  TURN endpoint, the IONOS upstream TCP-3478 filtering (load-bearing,
  undocumented before today), and the verification probe. New changelog
  entry covering the migration outcome + Device Code Stage gap.

* infrastructure/vps-hardening.md — "At a glance" lists the new 6-service
  docker stack + host coturn. Firewall section notes the new
  `iifname ens6 tcp dport 3478 accept` rule plus the IONOS upstream
  exception. New changelog entry.

* containers/124-authentik.md — replaces the "Netbird IdP integration —
  DEFERRED" section with the LANDED state: Provider details (Public
  client type — Confidential breaks PKCE on the dashboard SPA), the
  first-time owner-promotion sqlite recipe, the missing Device Code
  Stage gap + workaround (setup-keys), and a note that the old 2026-04-22
  pre-work Provider/App is now obsolete and safe to delete. Updated
  changelog (Phase 6 landed).

* operations/agent-enrollment.md — new "Getting onto Netbird" subsection
  explaining the setup-key path (currently the only working flow until
  Device Code Stage lands) and why direct OIDC from the public internet
  fails (auth.hubris.network is mesh-only-reachable). Prerequisites table
  row updated to point at the new section.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 13:48:01 +02:00

11 KiB

Mesh — Tailscale → Netbird migration

The hubris fleet is migrating from Tailscale to Netbird. Netbird is the target end-state. In-progress as of 2026-04-21.

Current state

  • PVE host uses Netbird (wt0, 100.122.38.109/16). Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets *.hubris.network → 192.168.8.175 via the system resolver chain.
  • Netbird mgmt host (82.165.190.79, FQDN inspiring-ramanujan.netbird.selfhosted, NB IP 100.122.165.149) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes the homelab network (192.168.8.0/24) via the PVE peer. This gives the mgmt host LAN access and split-horizon DNS for *.hubris.network. Useful independently of any Authentik integration.
  • Most LXCs still run Tailscale or use router DNS (192.168.8.1) / Tailscale MagicDNS (100.100.100.100), both of which return the public IONOS A record *.hubris.network → 82.165.190.79. The VPS only routes hostnames it actually publishes (today, artifacto + blog), so this path is a dead end for any LAN-only service.

ICE / STUN / TURN

Today (post-2026-05-21 migration):

  • External STUN servers (Google + Cloudflare) declared under server.Stuns in /opt/management.json. Embedded STUN is no longer in use.
  • coturn runs on the VPS host (apt-installed, systemd-managed). TURN URL turn:netbird.hubris.network:3478?transport=tcp is advertised to peers via mgmt's TURNConfig.Turns block. Long-term credentials at user=netbird:<pwd from /root/turn-pass.txt>.
  • For non-symmetric peers, ICE picks direct srflx/srflx (P2P). For symmetric-NAT peers, ICE falls through to TURN-relay before falling back to the WSS relay (rels://netbird.hubris.network:443).

IONOS port-3478 caveat (load-bearing — undocumented before 2026-05-21):

IONOS upstream filters STUN-class traffic on port 3478 for both UDP and TCP by default. The UDP block was already known (verified 2026-05-10 with tcpdump -ni any udp port 3478). The TCP block was discovered 2026-05-21 when TURN-allocate retries timed out from a symmetric-NAT laptop: TCP handshake completed (carried by kernel-only SYN exchange) but data packets never reached the VPS's ens6 interface.

Resolution: operator must add an inbound firewall exception in the IONOS dashboard (Cloud Panel → Networks → Firewall policy on the VPS) for inbound TCP 3478 (and UDP 3478 if STUN-via-VPS-listener is ever needed; we use external STUN so this is optional). Without that exception, coturn is unreachable from outside even though host nftables + listener look fine.

Verifying TURN works end-to-end from an outside peer:

# python3
import socket, struct, secrets
s = socket.create_connection(("netbird.hubris.network", 3478), timeout=10)
tid = secrets.token_bytes(12)
attrs = struct.pack("!HHI", 0x0019, 4, 17 << 24)   # REQUESTED-TRANSPORT = UDP
msg = struct.pack("!HHI", 0x0003, len(attrs), 0x2112A442) + tid + attrs  # Allocate
s.sendall(msg)
print(s.recv(4096).hex())  # expect 120 bytes starting 0x0113 (Allocate Error = auth challenge)

A 120-byte 0x0113-typed response means coturn is reachable and responding. An indefinite TCP timeout (after connected from ... prints) means the IONOS rule was reverted or narrowed.

If a peer is still on the WSS relay after this — verify in netbird status -d: Connection type: P2P with srflx/srflx is the cone-NAT happy path; Connection type: P2P with relay/srflx (or relay/prflx) means the peer's network is symmetric-NAT and it's using TURN. Falling back to Relayed (rels://...) only happens if TURN allocate also fails — investigate the IONOS exception first.

Old combined-server note (history, kept for context):

Pre-migration, the bundled netbirdio/netbird-server combined image silently ignored both server.turns: and top-level TURNConfig: YAML, so TURN was non-functional. That's why the architectural migration to the canonical multi-container stack (mgmt + signal + relay + coturn) happened. See the 2026-05-21 changelog entry below and the Netbird-combined-no-TURN finding in homelab memory for the full discovery.

Consequence — every LXC wired to Authentik needs an internal override

Until each LXC is migrated to Netbird, anything that needs to reach auth.hubris.network (Authentik), cloud.hubris.network (Nextcloud), etc., must override the public answer with 192.168.8.175.

Two techniques. Pick by HTTP-client behavior.

A) /etc/hosts override

Works for libc getaddrinfo clients: curl, wget, most Go/Python/Ruby apps, gitea.

  • Place the line outside the # --- BEGIN PVE --- / # --- END PVE --- markers. Proxmox rewrites everything inside that block on every container start.
  • Belt-and-suspenders: /etc/systemd/system/hubris-hosts-override.service (oneshot, enabled, idempotent).

B) Local dnsmasq

Required for clients that bypass /etc/hosts. Nextcloud (PHP Guzzle + OC\Http\Client\DnsPinMiddleware) is one — uses dns_get_record(), not getaddrinfo. Likely candidates for the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting.

Recipe:

apt install dnsmasq

cat > /etc/dnsmasq.d/hubris-internal.conf <<EOF
address=/auth.hubris.network/192.168.8.175
server=192.168.8.1
server=1.1.1.1
interface=lo
bind-interfaces
no-hosts
no-resolv
EOF

# Set LXC default nameservers and live resolv.conf
pct set <id> --nameserver "127.0.0.1 192.168.8.1 1.1.1.1"
# Then update /etc/resolv.conf inside the LXC too.

Known overrides applied

LXC Technique Notes
104 (gitea) /etc/hosts + hubris-hosts-override.service Standard
114 (nextcloud) local dnsmasq @ 127.0.0.1 + resolv.conf reorder Guzzle bypasses /etc/hosts
105 (apps), inner containers extra_hosts: in compose Booklore, mulita, WriteFreely each ship with this

Adding new LXCs

  • Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris.
  • When wiring a new app into Authentik: cat /etc/resolv.conf on the target LXC. If nameserver is 192.168.8.1 or 100.100.100.100, add the hosts override. If it's the netbird daemon IP, skip.

Long-term fix

Either:

  • Split-horizon DNS at LAN/router level so *.hubris.network → 192.168.8.175 for every LAN client. Eliminates all per-LXC overrides.
  • Once Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain on the PVE host.

CRITICAL — never docker compose up Portainer-managed stacks

apps (105) runs multiple stacks deployed via the Portainer UI (/var/lib/docker/volumes/portainer_data/_data/compose/<N>/). Running docker compose up -d <svc> from the host shell on these triggers a recreate of OTHER services in the stack — labels drift, compose treats the project as reconciling, containers get destroyed and remade. This wiped Booklore's mariadb data on 2026-04-22 (bind mount ./mariadb/config re-initialized fresh).

Recipe for container-config changes (e.g. adding extra_hosts) on Portainer-managed stacks:

  1. Edit the compose in Portainer UI → Stacks → → Editor → Update stack. Portainer handles the recreate cleanly with its own state tracking.
  2. Do NOT edit /var/lib/docker/volumes/portainer_data/_data/compose/<N>/docker-compose.yml directly.
  3. For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the app's built-in OIDC over container-level extra_hosts when possible.

Changelog

2026-05-21 — VPS migrated combined → vanilla netbird stack with external TURN

The combined netbirdio/netbird-server image was replaced with the canonical multi-container deploy (netbirdio/management:0.71.3 + signal:0.71.3 + relay:0.71.3 + dashboard:latest + host coturn) on /opt/docker-compose.yml. Driver: combined image silently ignored external TURNConfig so symmetric-NAT peers couldn't use TURN.

Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on LXC 124, upgrading mgmt to 0.71.3. The store.db schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old opt_netbird_data volume into the new mgmt_data volume). Pre-cutover backups at /root/netbird-*.tgz on the VPS, ~857 MB, retained for ~7d.

Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478 inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe).

The new Authentik provider for NetBird is Client type: Public (PKCE-only). Confidential would break the dashboard SPA's token exchange. Device Code Stage is NOT yet configured in Authentik → netbird up interactive auth flow returns an empty consent screen; new peers must use --setup-key until the stage is added.

Open follow-up: TURN-over-TLS on TCP 5349 (cert via certbot or extract Traefik's acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient for current usage.

2026-05-10 — ICE direct p2p restored (external STUN swap)

All peers were Connection type: Relayed because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in /opt/config.yaml. After netbird down/up, peers now report P2P with srflx/host candidates. Backup of pre-change config at /opt/config.yaml.bak-20260510-185050. Fixes a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers).

2026-04-28 — wiki entry created

Initial documentation.

2026-04-22 — Booklore mariadb data wiped (lesson recorded)

The "never docker compose up Portainer-managed stacks" rule comes from this. See apps (105).

2026-04-22 — netbird mgmt host joined its own mesh

82.165.190.79 is now a peer (100.122.165.149). LAN access + split-horizon DNS via PVE peer. See VPS hardening.

2026-04-21 — overrides applied to gitea (104) and nextcloud (114)

Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle.