Files
oikos/infrastructure/mesh.md
dtoro 21063015c7 homelab CLI: per-host ssh.user + LAN→mesh fallback; wiki for netbird-ssh JWT issuer fix
Three coordinated changes addressing follow-ups from the 2026-05-21 netbird vanilla
migration, plus a related off-LAN ergonomics fix:

bin/homelab:
- New ssh_target(name, force_mesh=False) helper resolves (addr, port, user)
  from inventory, honoring ssh.netbird_port (forces mesh path) and ssh.user
  (default "root"). Falls back to the netbird FQDN when LAN IP fails a
  cached 1.5s TCP probe — helps off-LAN operators on VPN/symmetric-NAT
  paths where the netbird subnet route doesn't reach 192.168.8.0/24.
- New ssh_base() builds the full `ssh ... user@addr` invocation; hubris_ssh()
  is now a back-compat shim. cmd_ssh, cmd_logs, cmd_restart, cmd_nuke
  refactored to use it — no more hardcoded "root@" anywhere.

inventory.yaml:
- New ssh.user convention (root by default, explicit per workstation).
- hubris.ssh.user=root (explicit, documents convention).
- republic-laptop, mac-mini: ssh.user=dtoro. ludo-mini left default (TODO).
- Comment block in the header explains the field + why it exists (netbird-ssh
  defaults to LOCAL username; "user not found" on LXCs is the gotcha).
- hosts/*.yaml regenerated from build_host_files.py.

infrastructure/mesh.md:
- Migration changelog entry updated: Device Code Stage is now configured
  (was "NOT yet" — landed in d41d73f); --setup-key no longer required.
- New subsection documenting the post-migration JWT-issuer cache bug:
  client/internal/engine_ssh.go's updateSSH() bails out when sshServer is
  already running, so systemctl restart and netbird down/up don't refresh
  the SSH JWT validator. Full daemon stop/start is the fix.
- Companion username gotcha (`netbird ssh` defaulting to local username).

operations/agent-enrollment.md:
- Three new troubleshooting rows: JWT-issuer cache, user-not-found, and
  LAN-unreachable-from-mesh-peer (the new homelab CLI behavior).

Verification: ssh_target resolution against the live inventory yields
- hubris → ssh -p 22022 root@proxmox-server.netbird.selfhosted (mesh-forced)
- jellyfin/gitea → ssh root@192.168.8.x (LAN reachable, probe passed)
- republic-laptop/mac-mini → ssh dtoro@<fqdn> (per ssh.user)
- ludo-mini → ssh root@<fqdn> (default)
2026-05-21 21:57:19 +02:00

13 KiB

Mesh — Tailscale → Netbird migration

The hubris fleet is migrating from Tailscale to Netbird. Netbird is the target end-state. In-progress as of 2026-04-21.

Current state

  • PVE host uses Netbird (wt0, 100.122.38.109/16). Its resolver is the local netbird daemon, which forwards to LAN/upstream — so the PVE host gets *.hubris.network → 192.168.8.175 via the system resolver chain.
  • Netbird mgmt host (82.165.190.79, FQDN inspiring-ramanujan.netbird.selfhosted, NB IP 100.122.165.149) is now itself a peer on the mesh (joined 2026-04-22 via setup key, netbird 0.69.0). Routes the homelab network (192.168.8.0/24) via the PVE peer. This gives the mgmt host LAN access and split-horizon DNS for *.hubris.network. Useful independently of any Authentik integration.
  • Most LXCs still run Tailscale or use router DNS (192.168.8.1) / Tailscale MagicDNS (100.100.100.100), both of which return the public IONOS A record *.hubris.network → 82.165.190.79. The VPS only routes hostnames it actually publishes (today, artifacto + blog), so this path is a dead end for any LAN-only service.

ICE / STUN / TURN

Today (post-2026-05-21 migration):

  • External STUN servers (Google + Cloudflare) declared under server.Stuns in /opt/management.json. Embedded STUN is no longer in use.
  • coturn runs on the VPS host (apt-installed, systemd-managed). TURN URL turn:netbird.hubris.network:3478?transport=tcp is advertised to peers via mgmt's TURNConfig.Turns block. Long-term credentials at user=netbird:<pwd from /root/turn-pass.txt>.
  • For non-symmetric peers, ICE picks direct srflx/srflx (P2P). For symmetric-NAT peers, ICE falls through to TURN-relay before falling back to the WSS relay (rels://netbird.hubris.network:443).

IONOS port-3478 caveat (load-bearing — undocumented before 2026-05-21):

IONOS upstream filters STUN-class traffic on port 3478 for both UDP and TCP by default. The UDP block was already known (verified 2026-05-10 with tcpdump -ni any udp port 3478). The TCP block was discovered 2026-05-21 when TURN-allocate retries timed out from a symmetric-NAT laptop: TCP handshake completed (carried by kernel-only SYN exchange) but data packets never reached the VPS's ens6 interface.

Resolution: operator must add an inbound firewall exception in the IONOS dashboard (Cloud Panel → Networks → Firewall policy on the VPS) for inbound TCP 3478 (and UDP 3478 if STUN-via-VPS-listener is ever needed; we use external STUN so this is optional). Without that exception, coturn is unreachable from outside even though host nftables + listener look fine.

Verifying TURN works end-to-end from an outside peer:

# python3
import socket, struct, secrets
s = socket.create_connection(("netbird.hubris.network", 3478), timeout=10)
tid = secrets.token_bytes(12)
attrs = struct.pack("!HHI", 0x0019, 4, 17 << 24)   # REQUESTED-TRANSPORT = UDP
msg = struct.pack("!HHI", 0x0003, len(attrs), 0x2112A442) + tid + attrs  # Allocate
s.sendall(msg)
print(s.recv(4096).hex())  # expect 120 bytes starting 0x0113 (Allocate Error = auth challenge)

A 120-byte 0x0113-typed response means coturn is reachable and responding. An indefinite TCP timeout (after connected from ... prints) means the IONOS rule was reverted or narrowed.

If a peer is still on the WSS relay after this — verify in netbird status -d: Connection type: P2P with srflx/srflx is the cone-NAT happy path; Connection type: P2P with relay/srflx (or relay/prflx) means the peer's network is symmetric-NAT and it's using TURN. Falling back to Relayed (rels://...) only happens if TURN allocate also fails — investigate the IONOS exception first.

Old combined-server note (history, kept for context):

Pre-migration, the bundled netbirdio/netbird-server combined image silently ignored both server.turns: and top-level TURNConfig: YAML, so TURN was non-functional. That's why the architectural migration to the canonical multi-container stack (mgmt + signal + relay + coturn) happened. See the 2026-05-21 changelog entry below and the Netbird-combined-no-TURN finding in homelab memory for the full discovery.

Consequence — every LXC wired to Authentik needs an internal override

Until each LXC is migrated to Netbird, anything that needs to reach auth.hubris.network (Authentik), cloud.hubris.network (Nextcloud), etc., must override the public answer with 192.168.8.175.

Two techniques. Pick by HTTP-client behavior.

A) /etc/hosts override

Works for libc getaddrinfo clients: curl, wget, most Go/Python/Ruby apps, gitea.

  • Place the line outside the # --- BEGIN PVE --- / # --- END PVE --- markers. Proxmox rewrites everything inside that block on every container start.
  • Belt-and-suspenders: /etc/systemd/system/hubris-hosts-override.service (oneshot, enabled, idempotent).

B) Local dnsmasq

Required for clients that bypass /etc/hosts. Nextcloud (PHP Guzzle + OC\Http\Client\DnsPinMiddleware) is one — uses dns_get_record(), not getaddrinfo. Likely candidates for the same treatment: Python/Ruby/Java apps that do explicit DNS resolution before connecting.

Recipe:

apt install dnsmasq

cat > /etc/dnsmasq.d/hubris-internal.conf <<EOF
address=/auth.hubris.network/192.168.8.175
server=192.168.8.1
server=1.1.1.1
interface=lo
bind-interfaces
no-hosts
no-resolv
EOF

# Set LXC default nameservers and live resolv.conf
pct set <id> --nameserver "127.0.0.1 192.168.8.1 1.1.1.1"
# Then update /etc/resolv.conf inside the LXC too.

Known overrides applied

LXC Technique Notes
104 (gitea) /etc/hosts + hubris-hosts-override.service Standard
114 (nextcloud) local dnsmasq @ 127.0.0.1 + resolv.conf reorder Guzzle bypasses /etc/hosts
105 (apps), inner containers extra_hosts: in compose Booklore, mulita, WriteFreely each ship with this

Adding new LXCs

  • Don't add new LXCs to Tailscale; add them to Netbird. Tailscale is being decommissioned on hubris.
  • When wiring a new app into Authentik: cat /etc/resolv.conf on the target LXC. If nameserver is 192.168.8.1 or 100.100.100.100, add the hosts override. If it's the netbird daemon IP, skip.

Long-term fix

Either:

  • Split-horizon DNS at LAN/router level so *.hubris.network → 192.168.8.175 for every LAN client. Eliminates all per-LXC overrides.
  • Once Netbird migration completes, every LXC's resolver becomes the netbird daemon, which already learns hubris.network answers via the system resolver chain on the PVE host.

CRITICAL — never docker compose up Portainer-managed stacks

apps (105) runs multiple stacks deployed via the Portainer UI (/var/lib/docker/volumes/portainer_data/_data/compose/<N>/). Running docker compose up -d <svc> from the host shell on these triggers a recreate of OTHER services in the stack — labels drift, compose treats the project as reconciling, containers get destroyed and remade. This wiped Booklore's mariadb data on 2026-04-22 (bind mount ./mariadb/config re-initialized fresh).

Recipe for container-config changes (e.g. adding extra_hosts) on Portainer-managed stacks:

  1. Edit the compose in Portainer UI → Stacks → → Editor → Update stack. Portainer handles the recreate cleanly with its own state tracking.
  2. Do NOT edit /var/lib/docker/volumes/portainer_data/_data/compose/<N>/docker-compose.yml directly.
  3. For data-preserving DNS overrides on stateful services, prefer Caddy-side handling (forward-auth) or the app's built-in OIDC over container-level extra_hosts when possible.

Changelog

2026-05-21 — VPS migrated combined → vanilla netbird stack with external TURN

The combined netbirdio/netbird-server image was replaced with the canonical multi-container deploy (netbirdio/management:0.71.3 + signal:0.71.3 + relay:0.71.3 + dashboard:latest + host coturn) on /opt/docker-compose.yml. Driver: combined image silently ignored external TURNConfig so symmetric-NAT peers couldn't use TURN.

Same migration also swapped OIDC from the combined image's embedded Dex IdP to Authentik on LXC 124, upgrading mgmt to 0.71.3. The store.db schema auto-migrated cleanly from 0.68.3 (copy-not-move from the old opt_netbird_data volume into the new mgmt_data volume). Pre-cutover backups at /root/netbird-*.tgz on the VPS, ~857 MB, retained for ~7d.

Also during this work: IONOS upstream was found to filter TCP 3478 in addition to UDP 3478. Added a TCP-3478 inbound exception in the IONOS firewall (see ICE/STUN section above for the verification probe).

The new Authentik provider for NetBird is Client type: Public (PKCE-only). Confidential would break the dashboard SPA's token exchange. The Device Code grant flow is wired (see containers/124-authentik.md) so interactive netbird up works — --setup-key is no longer required for new peers.

Post-migration JWT-issuer gotcha on existing peers (cost ~30 min to diagnose 2026-05-21):

Existing peers — registered against the old combined image's embedded Dex IdP at https://netbird.hubris.network/oauth2 — cache the OLD expected SSH-JWT issuer in the netbird daemon's in-memory state. After the migration, incoming netbird ssh connections were rejected with:

JWT authentication failed: validate token (
  expected issuer=https://netbird.hubris.network/oauth2,
            audiences=[netbird-dashboard netbird-cli],
  actual    issuer=https://auth.hubris.network/application/o/netbird/,
            audience=netbird-dashboard
)

Neither systemctl restart netbird nor netbird down && netbird up clears the cache. Root cause: in client/internal/engine_ssh.go, updateSSH() bails out with if e.sshServer != nil { return nil } whenever the SSH server is already running, so mgmt-pushed JWT config updates are silently ignored. Only a full daemon-process tear-down lets the SSH server re-initialize with the new validator config:

sudo systemctl stop netbird
sleep 3
sudo systemctl start netbird

After that, grep -iE "issuer|audience" /var/log/netbird/client.log | tail shows the new Authentik issuer. Run this on every existing peer (PVE host + every LXC + every workstation) once after a future IdP swap.

Username gotcha (related): netbird ssh defaults the remote username to the LOCAL one (e.g. dtoro from the operator's laptop). Hubris + the LXCs only have root, so the JWT is accepted but the session immediately fails with user dtoro not found. Always use the explicit root@ prefix when invoking netbird-ssh manually:

netbird ssh -p 22022 root@proxmox-server.netbird.selfhosted

The homelab CLI handles this automatically via the per-host ssh.user field in inventory.yaml (defaults to root; set explicitly only for workstations whose login user isn't root).

Open follow-up: TURN-over-TLS on TCP 5349 (cert via certbot or extract Traefik's acme.json) for hostile-middlebox networks; plain TCP 3478 is sufficient for current usage.

2026-05-10 — ICE direct p2p restored (external STUN swap)

All peers were Connection type: Relayed because the embedded STUN on the VPS was unreachable from outside (IONOS drops UDP 3478 inbound). Swapped to Google + Cloudflare STUN in /opt/config.yaml. After netbird down/up, peers now report P2P with srflx/host candidates. Backup of pre-change config at /opt/config.yaml.bak-20260510-185050. Fixes a real-world Nextcloud upload bottleneck (~366 kB/s through the relay → LAN-direct on same-network peers).

2026-04-28 — wiki entry created

Initial documentation.

2026-04-22 — Booklore mariadb data wiped (lesson recorded)

The "never docker compose up Portainer-managed stacks" rule comes from this. See apps (105).

2026-04-22 — netbird mgmt host joined its own mesh

82.165.190.79 is now a peer (100.122.165.149). LAN access + split-horizon DNS via PVE peer. See VPS hardening.

2026-04-21 — overrides applied to gitea (104) and nextcloud (114)

Two techniques documented; Nextcloud forced the dnsmasq route because of Guzzle.