Files
oikos/investigations/2026-06-06-caddyfile-truncation.md
dtoro 4efddb8bed docs: fix pre-existing broken links surfaced by docs-lint
Problem: docs-lint (added in the wiki-hq reorg) surfaced 126 broken relative
links that predated this session — a container rename, incident/plan docs
that moved into archive/done subfolders without their inbound links being
updated, and a handful of relative-depth bugs in files nested under
containers/archive/ and plans/done/.

Fixes applied, by category:
- 124-authentik.md -> 106-auth-outpost.md (container was renamed; ~40 refs).
- investigations/{2026-04-21-hubris-crash-loop,2026-05-31-authentik-vps-migration}.md
  -> archive/ prefix (both moved to investigations/archive/ previously).
- plans/{2026-06-01-slate-ax-to-sodola-migration,2026-06-04_130000-deprecate-claudio-bot,
  2026-06-25-yuvomi-deployment}.md -> plans/done/ prefix.
- Depth bugs in files nested one level deeper than their siblings assumed
  (investigations/archive/*, knowledge/wiki/containers/archive/*,
  plans/done/*) — corrected relative-path depth.
- Destroyed containers with no surviving page (126-plato) delinked to the
  containers/index.md archaeology row instead of a 404.
- ludo-mini.yaml -> strong.yaml (host was renamed, same physical machine).
- netbird-vps.md (no narrative page exists) -> netbird-vps.yaml (substrate
  record, matching the existing convention for hosts without a wiki page).
- runbook-dpkg-interrupted.md refs -> .agents/skills/runbook-dpkg-interrupted/SKILL.md
  (missed in the phase-4 runbook move because the referencing files used a
  bare filename, not a runbooks/ prefix).
- One dangling forward-reference to a never-written investigation delinked
  to the actual incident record it was describing.

Left alone: two links in knowledge/wiki/containers/101-jellyfin.md into
devops/homelab-authentik-admin/ — an intentional reference to a sibling repo,
not present in this checkout.

Verification: broken-link count 126 -> 2 (real remainder is the cross-repo
reference above); gen-topology.py --check still exit 0; build_host_files.py
still idempotent; all inventory.yaml doc_page targets still resolve.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 17:53:35 +02:00

3.4 KiB

Investigation: Caddyfile truncation — all LAN services down (2026-06-06)

Date: 2026-06-06 Status: resolved Duration: ~10 hours (from last known good state ~12:39 UTC to restoration ~22:40 UTC)

Symptom

All *.hubris.network URLs except photos.hubris.network and auth.hubris.network returned tlsv1 alert internal error or TCP timeouts from LAN/mesh clients. dig @192.168.8.2 and dig @100.122.255.254 both resolved to 192.168.8.175 correctly — DNS was fine. The issue was at the Caddy level.

Root cause

The Caddyfile on LXC 121 was manually edited directly on the filesystem (not via the dtoro/caddy-conf git repo), reducing it from 260 lines/30+ site blocks to 43 lines with only 3 photo-related site blocks: photos.hubris.network, prism.hubris.network, and photos2.hubris.network.

Timeline

Time (UTC+2) Event
Jun 04 23:43 Last successful git-push deploy — full Caddyfile (260 lines)
Jun 06 ~12:00 Caddyfile manually edited locally, truncating to 3 sites
Jun 06 12:39 Deploy webhook triggered → git pull --ff-only failed: "Your local changes would be overwritten"
Jun 06 14:13 Deploy webhook triggered again → deploy ok (the truncated file was committed or merged somehow)
Jun 06 22:34 Investigation began
Jun 06 22:43 Caddyfile restored from origin/master, systemctl reload caddy

Evidence

  • git diff HEAD -- Caddyfile on LXC 121: +3 / -159 lines
  • Git reflog: HEAD at 32575ce (fix: sab port 8081→8082), working tree diverged
  • Backup file Caddyfile.bak.1780263919: 225 lines, full original config
  • git stash list shows one auto-stash entry
  • origin/master at 1b977aa: 260 lines, all site blocks present

Secondary root cause found during investigation

elementsynapse (LXC 118) had iface eth0 inet dhcp internally despite pct set 118 --net0 ... ip=192.168.8.239/24. On DHCP lease renewal, dhclient grabbed .244 from Technitium's pool. Caddy's reverse_proxy 192.168.8.239:8008 was hitting a dead IP.

This is the same class of drift as the June 5th incidents (paperless, HAOS, apps, mule-images). Elementsynapse was missed during the 2026-06-02 static-IP migration.

Fix applied

  1. Caddyfilegit checkout --force origin/master -- Caddyfile + systemctl reload caddy
  2. elementsynapse → replaced iface eth0 inet dhcp with static, killed dhclient, verified connectivity

Permanent safeguards (all deployed)

Safeguard Location What it does
Site-count guard /etc/caddy/scripts/deploy.sh Refuses reload if <20 hubris.network site blocks
Dirty-tree auto-stash /etc/caddy/scripts/deploy.sh Stashes local edits before git pull
Auto-backup /etc/caddy/scripts/deploy.sh Saves Caddyfile.bak. before any change, keeps 5
Caddy backend health /etc/cron.d/caddy-backend-health on hubris Runs check-caddy-backends.sh every 10 min
DNS sync /etc/cron.d/dns-sync on LXC 107 Runs dns-sync.py every 10 min (was missing since 2026-06-04)