7.5 KiB
Mesh — Tailscale → Netbird migration
The hubris fleet is migrating from Tailscale to Netbird. Netbird is the target end-state. In-progress as of 2026-04-21.
Current state
- PVE host uses Netbird (
wt0,100.122.38.109/16). Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets*.hubris.network → 192.168.8.175via the system resolver chain. - Netbird mgmt host (
82.165.190.79, FQDNinspiring-ramanujan.netbird.selfhosted, NB IP100.122.165.149) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes the homelab network (192.168.8.0/24) via the PVE peer. This gives the mgmt host LAN access and split-horizon DNS for*.hubris.network. Useful independently of any Authentik integration. - Most LXCs still run Tailscale or use router DNS (
192.168.8.1) / Tailscale MagicDNS (100.100.100.100), both of which return the public IONOS A record*.hubris.network → 82.165.190.79. The VPS only routes hostnames it actually publishes (today,artifacto+blog), so this path is a dead end for any LAN-only service.
ICE / STUN — must use external STUN, not embedded
The bundled netbird-server image runs an embedded STUN listener on UDP 3478. IONOS drops inbound UDP 3478 to the VPS upstream of the host firewall (verified 2026-05-10 via tcpdump -ni any udp port 3478: 0 packets captured during external probes from hubris). Without a reachable STUN server, the management API hands peers a STUN URI nothing can talk to → no srflx candidates → ICE always fails → every peer falls back to the websocket relay (rels://netbird.hubris.network:443). All cross-NAT traffic is then bottlenecked by the relay/VPS bandwidth (observed ~366 kB/s for Nextcloud uploads).
Fix: in /opt/config.yaml on the VPS, declare external STUN servers under server: — this disables the embedded STUN automatically:
server:
stuns:
- uri: "stun:stun.l.google.com:19302"
- uri: "stun:stun1.l.google.com:19302"
- uri: "stun:stun.cloudflare.com:3478"
docker restart netbird-server, then netbird down && netbird up on each peer to force a resync. Verify with netbird status -d — Connection type: should flip from Relayed to P2P for peers that aren't behind double-NAT/CGNAT.
If a peer is still relayed after this, it's a NAT-symmetry problem on its side, not a config bug — would need TURN to fix.
Consequence — every LXC wired to Authentik needs an internal override
Until each LXC is migrated to Netbird, anything that needs to reach auth.hubris.network (Authentik), cloud.hubris.network (Nextcloud), etc., must override the public answer with 192.168.8.175.
Two techniques. Pick by HTTP-client behavior.
A) /etc/hosts override
Works for libc getaddrinfo clients: curl, wget, most Go/Python/Ruby apps, gitea.
- Place the line outside the
# --- BEGIN PVE ---/# --- END PVE ---markers. Proxmox rewrites everything inside that block on every container start. - Belt-and-suspenders:
/etc/systemd/system/hubris-hosts-override.service(oneshot, enabled, idempotent).
B) Local dnsmasq
Required for clients that bypass /etc/hosts. Nextcloud (PHP Guzzle + OC\Http\Client\DnsPinMiddleware) is one — uses dns_get_record(), not getaddrinfo. Likely candidates for the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting.
Recipe:
apt install dnsmasq
cat > /etc/dnsmasq.d/hubris-internal.conf <<EOF
address=/auth.hubris.network/192.168.8.175
server=192.168.8.1
server=1.1.1.1
interface=lo
bind-interfaces
no-hosts
no-resolv
EOF
# Set LXC default nameservers and live resolv.conf
pct set <id> --nameserver "127.0.0.1 192.168.8.1 1.1.1.1"
# Then update /etc/resolv.conf inside the LXC too.
Known overrides applied
| LXC | Technique | Notes |
|---|---|---|
| 104 (gitea) | /etc/hosts + hubris-hosts-override.service |
Standard |
| 114 (nextcloud) | local dnsmasq @ 127.0.0.1 + resolv.conf reorder | Guzzle bypasses /etc/hosts |
| 105 (apps), inner containers | extra_hosts: in compose |
Booklore, mulita, WriteFreely each ship with this |
Adding new LXCs
- Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris.
- When wiring a new app into Authentik:
cat /etc/resolv.confon the target LXC. If nameserver is192.168.8.1or100.100.100.100, add the hosts override. If it's the netbird daemon IP, skip.
Long-term fix
Either:
- Split-horizon DNS at LAN/router level so
*.hubris.network → 192.168.8.175for every LAN client. Eliminates all per-LXC overrides. - Once Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain on the PVE host.
CRITICAL — never docker compose up Portainer-managed stacks
apps (105) runs multiple stacks deployed via the Portainer UI (/var/lib/docker/volumes/portainer_data/_data/compose/<N>/). Running docker compose up -d <svc> from the host shell on these triggers a recreate of OTHER services in the stack — labels drift, compose treats the project as reconciling, containers get destroyed and remade. This wiped Booklore's mariadb data on 2026-04-22 (bind mount ./mariadb/config re-initialized fresh).
Recipe for container-config changes (e.g. adding extra_hosts) on Portainer-managed stacks:
- Edit the compose in Portainer UI → Stacks → → Editor → Update stack. Portainer handles the recreate cleanly with its own state tracking.
- Do NOT edit
/var/lib/docker/volumes/portainer_data/_data/compose/<N>/docker-compose.ymldirectly. - For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the app's built-in OIDC over container-level
extra_hostswhen possible.
Related
- DNS split-horizon
- Authentik (124) — the IdP that triggers most of these overrides
- Nextcloud (114) — example of Technique B
- Gitea (104) — example of Technique A
- Public ingress (VPS traefik) — uses the same mesh as transport
Changelog
2026-05-10 — ICE direct p2p restored (external STUN swap)
All peers were Connection type: Relayed because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in /opt/config.yaml. After netbird down/up, peers now report P2P with srflx/host candidates. Backup of pre-change config at /opt/config.yaml.bak-20260510-185050. Fixes a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers).
2026-04-28 — wiki entry created
Initial documentation.
2026-04-22 — Booklore mariadb data wiped (lesson recorded)
The "never docker compose up Portainer-managed stacks" rule comes from this. See apps (105).
2026-04-22 — netbird mgmt host joined its own mesh
82.165.190.79 is now a peer (100.122.165.149). LAN access + split-horizon DNS via PVE peer. See VPS hardening.
2026-04-21 — overrides applied to gitea (104) and nextcloud (114)
Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle.