- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
6.1 KiB
Plan: Fix Caddyfile truncation + prevent recurring outages
Date: 2026-06-06 Slug: caddyfile-truncation-permanent-fix
Goal
Restore all *.hubris.network services that went offline when the Caddyfile on LXC 121 was truncated to only 3 photo-related site blocks, and implement automated safeguards to prevent this class of outage from recurring.
Root cause
The Caddyfile at /etc/caddy/Caddyfile on LXC 121 was manually edited locally (not via the dtoro/caddy-conf git repo), overwriting ~260 lines (30+ site blocks + forward-auth infrastructure) with only 43 lines covering photos.hubris.network, prism.hubris.network, and a manually-added photos2.hubris.network.
Evidence:
git diff HEAD -- Caddyfileshows+3 / -159lines diff — all other blocks deleted- Git reflog shows HEAD at
32575ce(fix: sab... port 8081→8082), but working tree diverges - Deploy webhook log: Jun 06 12:39 —
deploy failed: git pull(dirty tree blocks merge) - Backup file
Caddyfile.bak.1780263919(225 lines) confirms the full original was intact before truncation origin/masterat1b977aais the authoritative source — 260 lines, all blocks present
Why "third time this week":
| Incident | Date | Cause |
|---|---|---|
| 1 | Jun 02 | DHCP IP drift — paperless (130→243), HAOS (101→241) |
| 2 | Jun 05 | More DHCP drift — apps (205), mule-images (136 overridden by dhclient) |
| 3 | Jun 06 | Caddyfile truncated — unrelated to IPs, much worse |
The Caddyfile truncation is the most severe: it took down all LAN services except photos.hubris.network and auth.hubris.network (VPS-hosted).
Immediate fix
Step 1: Restore Caddyfile from origin/master and reload
On LXC 121:
cd /etc/caddy
# Stash any local changes
git stash
# Reset to origin/master
git checkout --force origin/master -- Caddyfile
# Caddyfile now has all 30+ sites
caddy validate --config /etc/caddy/Caddyfile
systemctl reload caddy
This restores all service blocks including: media, git, paperless, books, home, cloud, matrix, proxmox, docker, jellyseerr, qbit, sab, blog, auth, artifacto, plato, zimaos, mcp, secrets, sso + authentik forward-auth infrastructure.
Step 2: Add photos2.hubris.network via git (if still needed)
The photos2.hubris.network block was manually added locally and is NOT in origin/master. If the user wants to keep it, submit a PR/commit to the dtoro/caddy-conf repo.
Step 3: Verify
- From any LAN/mesh client:
curl -sk https://media.hubris.network/→ 200 - Run
bash /opt/homelab-context/scripts/check-caddy-backends.shfrom hubris → all targets reachable - Flush mac-mini DNS:
sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder
Permanent safeguards
Layer 1: Caddyfile integrity check (deploy hook)
Add a site-count validation to the deploy script (/etc/caddy/scripts/deploy.sh):
# Count site blocks (lines matching *.hubris.network {)
SITE_COUNT=$(grep -c '^[a-z].*hubris.network {' Caddyfile)
if [ "$SITE_COUNT" -lt 20 ]; then
echo "[deploy] ERROR: Only $SITE_COUNT sites found (expected 20+). Refusing to reload."
exit 1
fi
This catches any future truncation before caddy reload runs.
Layer 2: Caddyfile backup on deploy
Add to deploy script before git pull:
cp Caddyfile "Caddyfile.bak.$(date +%s)"
Keep last 3 backups, auto-rotate.
Layer 3: Dirty-tree handling in deploy webhook
The deploy webhook currently hard-fails when the working tree is dirty. Change the receiver script to handle this gracefully:
cd /etc/caddy
# If dirty, stash local changes
if ! git diff --quiet; then
echo "[deploy] Working tree dirty — stashing"
git stash push -m "auto-stash by deploy webhook $(date)"
fi
git pull --ff-only
This prevents the webhook from blocking on future local edits.
Layer 4: Scheduled Caddyfile health check
Add a homelab cron job that runs check-caddy-backends.sh every 10 minutes and notifies if any Caddy backend is unreachable.
# In homelab context: cronjob
schedule: "*/10 * * * *"
script: /opt/homelab-context/scripts/check-caddy-backends.sh
Layer 5: DNS sync cron (fix already-deployed sync)
The dns-sync.py on LXC 107 at /opt/dns-sync/sync.py is installed but has no crontab — the sync never runs automatically. The NetBird managed DNS zone has drifted from Technitium. Add a systemd timer or crontab:
echo "*/10 * * * * root python3 /opt/dns-sync/sync.py >> /var/log/dns-sync.log 2>&1" > /etc/cron.d/dns-sync
Files likely to change
| File | Change |
|---|---|
/etc/caddy/Caddyfile on LXC 121 |
Restore from origin/master |
/etc/caddy/scripts/deploy.sh on LXC 121 |
Add site-count validation + backup + dirty-tree handling |
caddy-conf git repo |
PR with deploy.sh improvements + photos2 (if wanted) |
cronjob in Hermes |
Schedule check-caddy-backends.sh |
/etc/cron.d/dns-sync on LXC 107 |
New — add dns-sync cron |
Verification
- All
*.hubris.networkURLs load from mac-mini:media,git,paperless,cloud,home,proxmox, etc. check-caddy-backends.shexits 0 on hubrissystemctl status caddyshows active on LXC 121dns-syncruns and writes to/var/log/dns-sync.log
Risks / Tradeoffs
- Restoring from origin/master overwrites photos2.hubris.network — recreate it via proper git commit
- Caddy staging ACME certs for prism/photos2: The
tls dns ionosdirective uses staging env (acme-staging-v02.api.letsencrypt.org), which fails DNS propagation check (VPS port 53 unreachable from LXC). Once restored, these two subdomains will have the same issue. Move them to production IONOS DNS-01 by removing the staging CA directive or setting the correctacme_issuerin Caddyfile. - Dirty-tree stash could lose edits — mitigated by
git stash push --message+ backup file creation before stash
Open questions
- Keep
photos2.hubris.network? If yes, add via proper git push. prism.hubris.networkandphotos2certs fail on staging ACME — set productionacme_issuerin Caddyfile?- Should
check-caddy-backends.shrun as a homelab cron job or as a regular cron on LXC 121?