Root cause of the provision-time 504s: netbird home-lab-network (192.168.8.0/24)
had no active routing peer — mac-mini routing peer's netbird daemon was down, so
all home-backed public services (artifacto/blog/trmnl) 504'd at the VPS edge.
netbird up on mac-mini restored it; verified trmnl public 200/401, artifacto 200.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
LXC 128 trmnl hosts the TRMNL plugins middleware (dtoro/terminalito), polled by
TRMNL cloud. trmnl-plugins.service on :9851; Caddy block + LE cert; VPS traefik
router trmnl-public + cert mirror. Public path pending VPS<->home netbird route
recovery (was "No networks available" at provision time, artifacto/blog 504 too).
LAN Technitium record + SOPS enrollment + Google/MVG creds pending. Plan -> In Progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- scripts/dns-sync.py: reconcile Technitium named A-records -> NetBird managed
zone via API (cron */10 on dns LXC 107). Single authoring source; kills the
manual drift behind the auth/sso/nfs-export saga.
- secrets/netbird-pat.yaml: sops-encrypted NetBird API PAT for the sync.
- dns.md / 107-dns.md: document the sync model + why forward-to-Technitium was
abandoned (NetBird self-IP / nameserver-group quirks).
- Cleanup: removed inert Mac secondary; reverted primary AXFR; home-lab-dns ->
[192.168.8.2] (1/1 Available); deleted vestigial Proxmox Names group.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
hosts/hubris.md:
- Update At a glance network section: vmbr1 uplink (192.168.178.10),
vmbr0 portless internal bridge with 192.168.8.1 alias
- Remove Phase 1 WiFi failover section (wlp3s0 disabled 2026-06-02)
- Changelog: Slate AX retired, SODOLA added, Proxmox as subnet router
containers/121-caddy.md:
- Changelog: caddy.service unit was missing from hubris1 package,
recreated manually; risk of loss on package reinstall noted
containers/107-dns.md:
- Update Who points here: Technitium DHCP hands out .2 as DNS for
homelab clients; Fritz!Box LAN clients still get Fritz!Box DNS
- Add DHCP section documenting the homelab scope (100-240, gw .1)
- Changelog: DHCP enabled 2026-06-02, replaces Slate AX DHCP
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Domain-level forward-auth needs its own external_host domain when the IdP core
and outpost are on different hosts. sso.hubris.network -> Caddy -> LAN outpost.
Includes the redirect_uris-regeneration gotcha. Carry the DNS record into
Technitium in DNS Phase 2.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Moved Authentik (2026.2.2 -> 2026.5.2, +Redis, dedicated auth Docker net)
off LXC 124 onto the VPS so netbird-mgmt's OIDC dependency no longer requires
the mesh it authenticates. depends_on: service_healthy makes the deadlock
structurally impossible. Full Postgres DB migrated (users/apps/passwords/groups).
- investigations/2026-05-31-authentik-vps-migration.md: full writeup + lessons
- 124-authentik: migration banner + changelog (now legacy; dnsmasq stays)
- dns: auth.hubris.network -> 82.165.190.79; NetBird resolver cache gotcha
- ingress: auth served by local container via Docker-provider labels (not cert-mirror)
- mesh: follow-up entry superseding the morning band-aid; bootstrap note obsolete
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two changes that together collapse new-workstation onboarding from ~7 steps
to ~2 commands:
* bootstrap.sh:
- Dep-check now AUTO-INSTALLS missing tools (apt/dnf/brew) instead of
printing instructions and exiting. Re-verifies after install.
- New pre-mesh-check block: if netbird isn't installed, installs it
from the netbird apt/dnf repo (or `brew install --cask netbird` on
Darwin), then if mgmt isn't connected, runs `netbird up
--management-url=https://netbird.hubris.network --ssh-jwt-cache-ttl 86400`.
Operator clicks the device-code URL inline. Waits up to ~30s for
Management: Connected before continuing. Skipped on --no-secrets +
--dry-run.
* containers/124-authentik.md: replaces the "KNOWN MISSING — Device Code
Stage" subsection with a working recipe — Authentik 2026.2 routes
/device via a BRAND-level "Device code flow" field, not a provider
field. Documented stage bindings for a `default-device-code-flow`
flow (identification → password → user-login → consent) and the
brand-level binding step.
* operations/agent-enrollment.md: Path B (interactive `netbird up`) is
now the default; Path A (setup-key) demoted to "headless/scripted"
alternative. "Install dependencies" section collapsed into a note
that bootstrap handles it, with the manual recipes kept in a
collapsible <details> block for air-gapped use.
The flow uniquely available to lab owners (single Authentik user today)
still relies on the first-time-owner sqlite promotion documented in
124-authentik.md.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Updates to four pages reflecting the combined → vanilla mgmt+signal+relay+coturn
cutover and the IONOS-3478-firewall-exception discovery:
* infrastructure/mesh.md — rewrites the ICE/STUN section to cover the new
TURN endpoint, the IONOS upstream TCP-3478 filtering (load-bearing,
undocumented before today), and the verification probe. New changelog
entry covering the migration outcome + Device Code Stage gap.
* infrastructure/vps-hardening.md — "At a glance" lists the new 6-service
docker stack + host coturn. Firewall section notes the new
`iifname ens6 tcp dport 3478 accept` rule plus the IONOS upstream
exception. New changelog entry.
* containers/124-authentik.md — replaces the "Netbird IdP integration —
DEFERRED" section with the LANDED state: Provider details (Public
client type — Confidential breaks PKCE on the dashboard SPA), the
first-time owner-promotion sqlite recipe, the missing Device Code
Stage gap + workaround (setup-keys), and a note that the old 2026-04-22
pre-work Provider/App is now obsolete and safe to delete. Updated
changelog (Phase 6 landed).
* operations/agent-enrollment.md — new "Getting onto Netbird" subsection
explaining the setup-key path (currently the only working flow until
Device Code Stage lands) and why direct OIDC from the public internet
fails (auth.hubris.network is mesh-only-reachable). Prerequisites table
row updated to point at the new section.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds infrastructure/homelab-context.md as the architecture reference for
the cross-client context + MCP + secrets-issuance system. Updates:
- 105-apps.md: two new ## Stacks sections (homelab-mcp, secrets-issuance)
with their deploy pipelines + a row each in the public-hostname table;
changelog entry.
- auto-deploy.md: both new pipelines added to the table (one repo, two
webhooks, same push); per-pipeline notes covering the clone-per-service
pattern and the deploy.sh self-restart caveat; changelog entry.
- README.md: link to the new infrastructure page.
Operational walkthrough already lives at operations/agent-enrollment.md;
this commit is the architecture side of the same story.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PhotoPrism rotates its session HMAC key on every container start, so any
auto-deploy that recreated pp-app invalidated in-flight OIDC logins. The
deploy script was force-recreating + image-pulling on each push; pinned
both so pp-app survives a routine code deploy.
Measured cold-cache behaviour: thumbnails ~2ms, HEVC video playback
12-21s TTFB because libx264 transcodes inline and serialised one
ffmpeg at a time. Sibling thumbs unaffected (~2ms during transcode).
With 11/45 .mov files already pre-baked, ~75% of clicks were cold.
Ran `photoprism convert` to bring sidecar coverage to 45/45; previously
cold videos now serve in ~2ms TTFB.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Synapse runs on SQLite (not Postgres — Postgres only hosts the
mautrix bridge dbs). Documented the post-disk-full
event_push_actions cleanup that took @admin'\''s phantom count
from 125 to 4.
Rootfs grew to 16 GiB on 2026-05-15. Added bridges section
(mautrix slack/signal/meta/linkedin/whatsapp), synapse-admin
port, Postgres backend note, and operational tips for future
disk growth.
Container had been stopped since 2026-04-21 and was never re-enabled.
pct destroy 109 --purge cleaned up vm-109-disk-0 on local-lvm and the
config file. /mnt/library/syncthing subtree was already empty at the
time of destruction and is retained as an empty dir (no real data to
migrate or back up).
- README.md, containers/index.md: removed row, moved to "recently
destroyed" table
- hosts/hubris.md: dropped from /mnt/library subtree list, updated
containers/index summary line, added changelog entry
- infrastructure/media-permissions.md: dropped from membership table
and onboarding example, generalised pct-exec gotcha hostname,
added changelog
- vms/100-zimaos.md: dropped from "existing fleet" enumeration
- containers/102-nfs-export.md: dropped from bind-mount sibling list
(7 LXCs now, not 8)
- containers/109-syncthing.md: deleted
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After Files-UI evaluation passed (library renders as folder, thumbnails
work), flipped /etc/exports from ro to rw on LXC 102. Tested: write from
ZimaOS appears on /mnt/library as www-data:media, confirming the
all_squash,anonuid=33,anongid=10000 design works.
Documented two architectural findings discovered this session:
- ZimaOS Drives panel sources from GET /v2/local_storage/storages (read-only
API). Network shares cannot become Drives — Files-as-folder is supported.
- Mesh peers reach ZimaOS via hubris's existing 192.168.8.0/24 netbird subnet
advertisement; no new infra needed, just DNS (Management nameserver group
for hubris.network or per-device /etc/hosts override).
User destroyed the heaper LXC on 2026-05-14. Removed it from the
container index, README quicktable, hubris host doc, and
media-permissions membership table; moved to the "recently destroyed"
archaeology list. /mnt/library/heaper retained (224 MiB).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Reflects state after the 2026-05-14 session: the OpenCLIP vision
classifier and worker-vision are gone, the db image is now postgres:16,
and the one-shot full_refresh.py reaped 7982 orphan thumbnail dirs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ENABLE_METADATA_MANAGEMENT=True + METADATA_SERVER_URL=http://seafile-md-server:8084 in seahub_settings.py. Library owners can now toggle metadata per library and create table/gallery/kanban Views over their files.
Two new services on LXC 125's docker stack:
- seafile-md-server (Pro extended metadata, internal-only on :8084)
- thumbnail-server (Pro accelerated thumbnails, bound to .185:8081)
Caddy now routes /thumbnail/* to the thumbnail-server; everything else stays
on the main seafile container. End-to-end smoke verified: routing works,
per-request permission checks via INNER_SEAHUB_SERVICE_URL=http://seafile
correctly return 403 for cross-user thumbnail requests.
Same-day-as-deploy upgrade: image swapped to seafileltd/seafile-pro-mc:13.0-latest, elasticsearch:8.15.0 added as new compose service for Pro's full-text search. Free Pro tier (<=3 users, no license). Existing data + users survived.
Also documented the Caddy patch stripping IETF resumable-upload headers (Upload-Draft-Interop-Version etc.) so the iOS Seafile Pro 4.0.2 app falls back to plain multipart upload; without it large uploads stalled and cancelled after ~60s.
LXC 125 stood up as a Seafile CE 13.0 docker-compose deployment, behind
files.hubris.network. Authentik OAuth wired up via ak shell. No data
migration — exploration alongside Nextcloud (114).
- Vision fetches NC preview at 640px via new sync helper.
- WORKER_THUMB_SIZES = set(); generate_thumbnails still does pHash.
- All medium.webp purged after verification.
- Bug: SECRET_KEY was missing from workers in compose, so
Fernet decrypt silently failed everywhere outside the backend
container. Phase 3 had been falling back to ExifTool the whole
time. Fix replicates SECRET_KEY to all worker services.
Tries Memories /api/image/info/{fileid} first (OCS-APIRequest header
bypasses CSRF), falls back to ExifTool on 404 / non-NC / errors.
Kept the full date-fallback chain because 35% of the library uses
filename-encoded dates that Memories doesn't recover.
- handle_directory_rename in scan.py covers NC-side renames via the
webhook. Iterates rows in Python (asyncpg int-type quirk on raw
SUBSTRING+LENGTH).
- Existing PATCH /folders/{id} handles mule-side renames; webhook
feedback hits the same helper and is a 0-row no-op (idempotent).
- Folder delete now propagates: webhook handler detects directory
deletes and runs a single UPDATE that discards every Photo under
the path prefix.
- PUT-overwrite of a previously-discarded file now resurrects the
Photo row (is_discarded=false, re-queue extract_metadata).
- Trashbin restore and folder rename remain known gaps (documented).
- Worker-watcher retired; file events come from NC webhook_listeners.
- Thumbnails proxy /index.php/core/preview keyed by photos.nextcloud_fileid.
- backfill_gps auto-trigger killed (was queueing ~60k tasks per deploy).
- Range support added to /api/v1/photos/{id}/original so .mov plays.
- NC cron tightened to */1 for ~60s webhook latency.