- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
3.4 KiB
Investigation: Caddyfile truncation — all LAN services down (2026-06-06)
Date: 2026-06-06 Status: resolved Duration: ~10 hours (from last known good state ~12:39 UTC to restoration ~22:40 UTC)
Symptom
All *.hubris.network URLs except photos.hubris.network and auth.hubris.network returned tlsv1 alert internal error or TCP timeouts from LAN/mesh clients. dig @192.168.8.2 and dig @100.122.255.254 both resolved to 192.168.8.175 correctly — DNS was fine. The issue was at the Caddy level.
Root cause
The Caddyfile on LXC 121 was manually edited directly on the filesystem (not via the dtoro/caddy-conf git repo), reducing it from 260 lines/30+ site blocks to 43 lines with only 3 photo-related site blocks: photos.hubris.network, prism.hubris.network, and photos2.hubris.network.
Timeline
| Time (UTC+2) | Event |
|---|---|
| Jun 04 23:43 | Last successful git-push deploy — full Caddyfile (260 lines) |
| Jun 06 ~12:00 | Caddyfile manually edited locally, truncating to 3 sites |
| Jun 06 12:39 | Deploy webhook triggered → git pull --ff-only failed: "Your local changes would be overwritten" |
| Jun 06 14:13 | Deploy webhook triggered again → deploy ok (the truncated file was committed or merged somehow) |
| Jun 06 22:34 | Investigation began |
| Jun 06 22:43 | Caddyfile restored from origin/master, systemctl reload caddy |
Evidence
git diff HEAD -- Caddyfileon LXC 121:+3 / -159lines- Git reflog: HEAD at
32575ce(fix: sab port 8081→8082), working tree diverged - Backup file
Caddyfile.bak.1780263919: 225 lines, full original config git stash listshows one auto-stash entryorigin/masterat1b977aa: 260 lines, all site blocks present
Secondary root cause found during investigation
elementsynapse (LXC 118) had iface eth0 inet dhcp internally despite pct set 118 --net0 ... ip=192.168.8.239/24. On DHCP lease renewal, dhclient grabbed .244 from Technitium's pool. Caddy's reverse_proxy 192.168.8.239:8008 was hitting a dead IP.
This is the same class of drift as the June 5th incidents (paperless, HAOS, apps, mule-images). Elementsynapse was missed during the 2026-06-02 static-IP migration.
Fix applied
- Caddyfile →
git checkout --force origin/master -- Caddyfile+systemctl reload caddy - elementsynapse → replaced
iface eth0 inet dhcpwith static, killed dhclient, verified connectivity
Permanent safeguards (all deployed)
| Safeguard | Location | What it does |
|---|---|---|
| Site-count guard | /etc/caddy/scripts/deploy.sh |
Refuses reload if <20 hubris.network site blocks |
| Dirty-tree auto-stash | /etc/caddy/scripts/deploy.sh |
Stashes local edits before git pull |
| Auto-backup | /etc/caddy/scripts/deploy.sh |
Saves Caddyfile.bak. before any change, keeps 5 |
| Caddy backend health | /etc/cron.d/caddy-backend-health on hubris |
Runs check-caddy-backends.sh every 10 min |
| DNS sync | /etc/cron.d/dns-sync on LXC 107 |
Runs dns-sync.py every 10 min (was missing since 2026-06-04) |
Related
- DHCP drift investigation (previous incident) — not filed as its own investigation; see the DNS sync fix
- Caddy (121)
- elementsynapse (118)
- dns-sync script
- check-caddy-backends script