After upgrading the Mac client 0.68.3 -> 0.71.3 (matching mgmt) and clearing
its NetBird resolver cache (netbird service restart), the managed-zone
deletion works: with 0 managed-zone records, the Mac resolves all
hubris.network names by forwarding to Technitium (192.168.8.2). iPhone
confirmed on cellular (no LAN path -> proves mesh-forward).
The first deletion "failure" was a misdiagnosis: the old 0.68.3 resolver
cache held stale answers and wouldn't clear on down/up (needs daemon
restart); the Mac's dual LAN+mesh paths muddied it. A direct
dig @100.122.255.254 of an unsynced name had shown forwarding working.
Done:
- Deleted all 23 NetBird managed-zone A-records.
- Removed the */10 dns-sync cron. Kept /opt/dns-sync/sync.py + token +
pre-deletion backup as an emergency-restore tool only.
End state: Technitium is the single DNS source. Mesh peers forward to it
(Core route -> 192.168.8.0/24); LAN/household query it directly. No replica,
no sync. Requires mesh clients on 0.71.x+.
Docs: dns.md + 107-dns.md updated to single-source; subdomain recipe no
longer references the sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Deleting the NetBird managed-zone replica broke mesh-peer DNS: the Mac
(NetBird 0.68.3) could not resolve hubris.network via the home-lab-dns
nameserver group (-> 192.168.8.2) even though it shows "Available" and the
192.168.8.0/24 route is present. Forwarding to the routed-LAN IP does not
actually serve queries on the current client. Restored the managed zone via
dns-sync.py and re-enabled the cron; resolution recovered.
Correction to Phase 2: the route fix delivered roaming-peer *service
connectivity* (the real iPhone win) but did NOT enable DNS forwarding. The
original "NetBird won't forward to Technitium for mesh peers" finding
stands; managed zone + sync are retained as load-bearing.
To finish single-source later: upgrade clients to 0.71.x, or point the
nameserver group at a mesh-native DNS IP (join CT 107 to the mesh).
Docs: dns.md + 107-dns.md corrected to reflect retained managed zone.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Root cause of "NetBird won't forward to Technitium" was NOT a nameserver
bug — it was a missing route. The home-lab-dns nameserver group
(-> 192.168.8.2, domain hubris.network) was applied to all peers, but the
192.168.8.0/24 route (home-lab-network resource) was distributed only to
the Services group. Roaming peers (Core: iphone + laptops) had no route to
192.168.8.2, so forwarding silently failed (Networks: -).
Fix: added Core to the home-lab-network resource distribution via the
NetBird API. Roaming peers now get the subnet route + the already-applied
nameserver forwarding -> *.hubris.network resolves off-LAN. Also grants
roaming devices full homelab service access. No CT 107 mesh-join needed
(original Phase 2 hypothesis obsolete).
Docs: corrected dns.md + 107-dns.md root-cause claims. Managed zone kept
as fallback pending Phase 4.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 1 of the DNS-redundancy cleanup (keep NetBird, collapse toward one
zone) — the safe, mesh-independent half:
- Every homelab LXC now resolves via Technitium (192.168.8.2). Fixed 8
boxes on a dead resolver (.180), the router (.1), or Tailscale MagicDNS
(100.100.100.100): 101,102,104,105,106,114,119,126.
- Removed the redundant /etc/hosts auth/mcp/secrets overrides (Technitium
returns identical-or-better answers); disabled hubris-hosts-override.
- Net effect: on-prem DNS (LXCs + household via Fritz!Box->Technitium) is
now NetBird-independent, so dropping the managed zone later can't break
on-LAN resolution. Phases 2-4 still pending.
Tailscale decommissioned (was legacy/being-phased-out):
- Removed from the 6 LXCs still running it (101,103,104,105,114,119):
logout, disable tailscaled, apt purge, state cleared.
- inventory.yaml: dropped tailscale from accepted + all mesh blocks;
regenerated hosts/*.yaml (also pruned orphan authentik/claudio-bot).
- Tightened secrets-issuance MESH_SUBNETS: removed the now-vestigial
Tailscale CGNAT range 100.64.0.0/10.
- Updated narrative docs (mesh, dns, network, README, AGENTS,
agent-enrollment, homelab-context, 105-apps, 107-dns).
Live infra changed on the fleet + Mac; this commit records the docs/inventory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- scripts/dns-sync.py: reconcile Technitium named A-records -> NetBird managed
zone via API (cron */10 on dns LXC 107). Single authoring source; kills the
manual drift behind the auth/sso/nfs-export saga.
- secrets/netbird-pat.yaml: sops-encrypted NetBird API PAT for the sync.
- dns.md / 107-dns.md: document the sync model + why forward-to-Technitium was
abandoned (NetBird self-IP / nameserver-group quirks).
- Cleanup: removed inert Mac secondary; reverted primary AXFR; home-lab-dns ->
[192.168.8.2] (1/1 Available); deleted vestigial Proxmox Names group.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
hosts/hubris.md:
- Update At a glance network section: vmbr1 uplink (192.168.178.10),
vmbr0 portless internal bridge with 192.168.8.1 alias
- Remove Phase 1 WiFi failover section (wlp3s0 disabled 2026-06-02)
- Changelog: Slate AX retired, SODOLA added, Proxmox as subnet router
containers/121-caddy.md:
- Changelog: caddy.service unit was missing from hubris1 package,
recreated manually; risk of loss on package reinstall noted
containers/107-dns.md:
- Update Who points here: Technitium DHCP hands out .2 as DNS for
homelab clients; Fritz!Box LAN clients still get Fritz!Box DNS
- Add DHCP section documenting the homelab scope (100-240, gw .1)
- Changelog: DHCP enabled 2026-06-02, replaces Slate AX DHCP
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Domain-level forward-auth needs its own external_host domain when the IdP core
and outpost are on different hosts. sso.hubris.network -> Caddy -> LAN outpost.
Includes the redirect_uris-regeneration gotcha. Carry the DNS record into
Technitium in DNS Phase 2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Moved Authentik (2026.2.2 -> 2026.5.2, +Redis, dedicated auth Docker net)
off LXC 124 onto the VPS so netbird-mgmt's OIDC dependency no longer requires
the mesh it authenticates. depends_on: service_healthy makes the deadlock
structurally impossible. Full Postgres DB migrated (users/apps/passwords/groups).
- investigations/2026-05-31-authentik-vps-migration.md: full writeup + lessons
- 124-authentik: migration banner + changelog (now legacy; dnsmasq stays)
- dns: auth.hubris.network -> 82.165.190.79; NetBird resolver cache gotcha
- ingress: auth served by local container via Docker-provider labels (not cert-mirror)
- mesh: follow-up entry superseding the morning band-aid; bootstrap note obsolete
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two changes that together collapse new-workstation onboarding from ~7 steps
to ~2 commands:
* bootstrap.sh:
- Dep-check now AUTO-INSTALLS missing tools (apt/dnf/brew) instead of
printing instructions and exiting. Re-verifies after install.
- New pre-mesh-check block: if netbird isn't installed, installs it
from the netbird apt/dnf repo (or `brew install --cask netbird` on
Darwin), then if mgmt isn't connected, runs `netbird up
--management-url=https://netbird.hubris.network --ssh-jwt-cache-ttl 86400`.
Operator clicks the device-code URL inline. Waits up to ~30s for
Management: Connected before continuing. Skipped on --no-secrets +
--dry-run.
* containers/124-authentik.md: replaces the "KNOWN MISSING — Device Code
Stage" subsection with a working recipe — Authentik 2026.2 routes
/device via a BRAND-level "Device code flow" field, not a provider
field. Documented stage bindings for a `default-device-code-flow`
flow (identification → password → user-login → consent) and the
brand-level binding step.
* operations/agent-enrollment.md: Path B (interactive `netbird up`) is
now the default; Path A (setup-key) demoted to "headless/scripted"
alternative. "Install dependencies" section collapsed into a note
that bootstrap handles it, with the manual recipes kept in a
collapsible <details> block for air-gapped use.
The flow uniquely available to lab owners (single Authentik user today)
still relies on the first-time-owner sqlite promotion documented in
124-authentik.md.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Updates to four pages reflecting the combined → vanilla mgmt+signal+relay+coturn
cutover and the IONOS-3478-firewall-exception discovery:
* infrastructure/mesh.md — rewrites the ICE/STUN section to cover the new
TURN endpoint, the IONOS upstream TCP-3478 filtering (load-bearing,
undocumented before today), and the verification probe. New changelog
entry covering the migration outcome + Device Code Stage gap.
* infrastructure/vps-hardening.md — "At a glance" lists the new 6-service
docker stack + host coturn. Firewall section notes the new
`iifname ens6 tcp dport 3478 accept` rule plus the IONOS upstream
exception. New changelog entry.
* containers/124-authentik.md — replaces the "Netbird IdP integration —
DEFERRED" section with the LANDED state: Provider details (Public
client type — Confidential breaks PKCE on the dashboard SPA), the
first-time owner-promotion sqlite recipe, the missing Device Code
Stage gap + workaround (setup-keys), and a note that the old 2026-04-22
pre-work Provider/App is now obsolete and safe to delete. Updated
changelog (Phase 6 landed).
* operations/agent-enrollment.md — new "Getting onto Netbird" subsection
explaining the setup-key path (currently the only working flow until
Device Code Stage lands) and why direct OIDC from the public internet
fails (auth.hubris.network is mesh-only-reachable). Prerequisites table
row updated to point at the new section.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds infrastructure/homelab-context.md as the architecture reference for
the cross-client context + MCP + secrets-issuance system. Updates:
- 105-apps.md: two new ## Stacks sections (homelab-mcp, secrets-issuance)
with their deploy pipelines + a row each in the public-hostname table;
changelog entry.
- auto-deploy.md: both new pipelines added to the table (one repo, two
webhooks, same push); per-pipeline notes covering the clone-per-service
pattern and the deploy.sh self-restart caveat; changelog entry.
- README.md: link to the new infrastructure page.
Operational walkthrough already lives at operations/agent-enrollment.md;
this commit is the architecture side of the same story.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PhotoPrism rotates its session HMAC key on every container start, so any
auto-deploy that recreated pp-app invalidated in-flight OIDC logins. The
deploy script was force-recreating + image-pulling on each push; pinned
both so pp-app survives a routine code deploy.
Measured cold-cache behaviour: thumbnails ~2ms, HEVC video playback
12-21s TTFB because libx264 transcodes inline and serialised one
ffmpeg at a time. Sibling thumbs unaffected (~2ms during transcode).
With 11/45 .mov files already pre-baked, ~75% of clicks were cold.
Ran `photoprism convert` to bring sidecar coverage to 45/45; previously
cold videos now serve in ~2ms TTFB.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Synapse runs on SQLite (not Postgres — Postgres only hosts the
mautrix bridge dbs). Documented the post-disk-full
event_push_actions cleanup that took @admin'\''s phantom count
from 125 to 4.
Rootfs grew to 16 GiB on 2026-05-15. Added bridges section
(mautrix slack/signal/meta/linkedin/whatsapp), synapse-admin
port, Postgres backend note, and operational tips for future
disk growth.
Container had been stopped since 2026-04-21 and was never re-enabled.
pct destroy 109 --purge cleaned up vm-109-disk-0 on local-lvm and the
config file. /mnt/library/syncthing subtree was already empty at the
time of destruction and is retained as an empty dir (no real data to
migrate or back up).
- README.md, containers/index.md: removed row, moved to "recently
destroyed" table
- hosts/hubris.md: dropped from /mnt/library subtree list, updated
containers/index summary line, added changelog entry
- infrastructure/media-permissions.md: dropped from membership table
and onboarding example, generalised pct-exec gotcha hostname,
added changelog
- vms/100-zimaos.md: dropped from "existing fleet" enumeration
- containers/102-nfs-export.md: dropped from bind-mount sibling list
(7 LXCs now, not 8)
- containers/109-syncthing.md: deleted
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After Files-UI evaluation passed (library renders as folder, thumbnails
work), flipped /etc/exports from ro to rw on LXC 102. Tested: write from
ZimaOS appears on /mnt/library as www-data:media, confirming the
all_squash,anonuid=33,anongid=10000 design works.
Documented two architectural findings discovered this session:
- ZimaOS Drives panel sources from GET /v2/local_storage/storages (read-only
API). Network shares cannot become Drives — Files-as-folder is supported.
- Mesh peers reach ZimaOS via hubris's existing 192.168.8.0/24 netbird subnet
advertisement; no new infra needed, just DNS (Management nameserver group
for hubris.network or per-device /etc/hosts override).
User destroyed the heaper LXC on 2026-05-14. Removed it from the
container index, README quicktable, hubris host doc, and
media-permissions membership table; moved to the "recently destroyed"
archaeology list. /mnt/library/heaper retained (224 MiB).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Reflects state after the 2026-05-14 session: the OpenCLIP vision
classifier and worker-vision are gone, the db image is now postgres:16,
and the one-shot full_refresh.py reaped 7982 orphan thumbnail dirs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ENABLE_METADATA_MANAGEMENT=True + METADATA_SERVER_URL=http://seafile-md-server:8084 in seahub_settings.py. Library owners can now toggle metadata per library and create table/gallery/kanban Views over their files.
Two new services on LXC 125's docker stack:
- seafile-md-server (Pro extended metadata, internal-only on :8084)
- thumbnail-server (Pro accelerated thumbnails, bound to .185:8081)
Caddy now routes /thumbnail/* to the thumbnail-server; everything else stays
on the main seafile container. End-to-end smoke verified: routing works,
per-request permission checks via INNER_SEAHUB_SERVICE_URL=http://seafile
correctly return 403 for cross-user thumbnail requests.
Same-day-as-deploy upgrade: image swapped to seafileltd/seafile-pro-mc:13.0-latest, elasticsearch:8.15.0 added as new compose service for Pro's full-text search. Free Pro tier (<=3 users, no license). Existing data + users survived.
Also documented the Caddy patch stripping IETF resumable-upload headers (Upload-Draft-Interop-Version etc.) so the iOS Seafile Pro 4.0.2 app falls back to plain multipart upload; without it large uploads stalled and cancelled after ~60s.
LXC 125 stood up as a Seafile CE 13.0 docker-compose deployment, behind
files.hubris.network. Authentik OAuth wired up via ak shell. No data
migration — exploration alongside Nextcloud (114).
- Vision fetches NC preview at 640px via new sync helper.
- WORKER_THUMB_SIZES = set(); generate_thumbnails still does pHash.
- All medium.webp purged after verification.
- Bug: SECRET_KEY was missing from workers in compose, so
Fernet decrypt silently failed everywhere outside the backend
container. Phase 3 had been falling back to ExifTool the whole
time. Fix replicates SECRET_KEY to all worker services.
Tries Memories /api/image/info/{fileid} first (OCS-APIRequest header
bypasses CSRF), falls back to ExifTool on 404 / non-NC / errors.
Kept the full date-fallback chain because 35% of the library uses
filename-encoded dates that Memories doesn't recover.
- handle_directory_rename in scan.py covers NC-side renames via the
webhook. Iterates rows in Python (asyncpg int-type quirk on raw
SUBSTRING+LENGTH).
- Existing PATCH /folders/{id} handles mule-side renames; webhook
feedback hits the same helper and is a 0-row no-op (idempotent).
- Folder delete now propagates: webhook handler detects directory
deletes and runs a single UPDATE that discards every Photo under
the path prefix.
- PUT-overwrite of a previously-discarded file now resurrects the
Photo row (is_discarded=false, re-queue extract_metadata).
- Trashbin restore and folder rename remain known gaps (documented).
- Worker-watcher retired; file events come from NC webhook_listeners.
- Thumbnails proxy /index.php/core/preview keyed by photos.nextcloud_fileid.
- backfill_gps auto-trigger killed (was queueing ~60k tasks per deploy).
- Range support added to /api/v1/photos/{id}/original so .mov plays.
- NC cron tightened to */1 for ~60s webhook latency.