archive/ contained the old narrative wiki (superseded by DB as source of truth), hermes-plans, oikos-cards, ledger, secrets-issuance, and SOPS backups — all Python-era artifacts with no ongoing value. Updated all cross-references in: - AGENTS.md, README.md - .agents/operations/commands.md (point to docs/infrastructure/) - .agents/shared/llm-wiki.md, page-templates.md - .agents/domains/knowledge/schema.md, operations/schema.md - .agents/skills/*/SKILL.md - docs/infrastructure/*.md (removed archive link targets) - docs-lint/SKILL.md known-baseline note
4.3 KiB
name, risk_class, verification
| name | risk_class | verification |
|---|---|---|
| recover-dpkg-interrupted | reversible_low | dpkg --audit (should be clean); apt-get check |
Runbook — recover from dpkg-interrupted state
You're here because an apt run got killed mid-transaction and the target now has packages that are unpacked but not configured. Symptoms:
aptrefuses to do anything new:Error: dpkg was interrupted, you must manually run 'dpkg --configure -a' to correct the problem.dpkg --auditlists packages with headerThe following packages have been unpacked but not yet configured.dpkg --auditon the host directly shows unpacked-not-configured packages (there's no fleet-wide audit tool anymore — check per-host).
The system is still running the old binaries (still in memory), but the new binaries are unpacked and waiting for their postinst to run. Two worst-case manifestations from the 2026-05-21 sweep:
- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was still serving but the new caddy binary on disk hadn't been wired up.
- hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured,
including
systemd,openssh-server,sudo,netbird. The half- configured netbird daemon dropped the mesh peer, and we got locked out until we recovered from the PVE web UI Shell.
Do not reboot until dpkg is clean. A reboot tries to start the new binaries' services, which may fail because postinst never ran (missing users, config dirs, capabilities, etc.). The system might not come back up cleanly.
Path A — target is still reachable over ssh (preferred)
ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
Or for an LXC by name (via the MCP run tool, or directly on the Proxmox
host):
pct exec <lxc> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
When that returns, confirm:
ssh <host> -- dpkg --audit
Expect DPKG: ok and the remaining UPGR count to match what's intentionally
deferred (kernel/PVE on hubris, 0 elsewhere).
Path B — target locked out (mesh broken / ssh dead)
Most common for hubris when netbird itself went half-configured: the daemon crashed on the new binary, the mesh peer dropped, port 22022 stopped listening, and you can't ssh in.
- Open
https://proxmox.hubris.networkin a browser. - Datacenter → node
hubris→>_ Shell(or_ Console). That's a root shell on hubris served by the PVE web UI, independent of the netbird mesh. - Run the recovery one-liner:
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
&& DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
&& systemctl restart netbird \
&& dpkg --audit \
&& echo RECOVERY_OK
Wait for RECOVERY_OK. The systemctl restart netbird is the bit that
heals the mesh — once netbird's daemon comes back up clean, your client's
peer state moves from Connecting to Connected within ~30 seconds and
the rest of your tooling works again.
-
For an LXC that's locked out (less common — LXCs reach the world via netbird routed through hubris, so unless hubris itself is broken, you can still
pct enterfrom the hubris shell):From the PVE web UI shell on hubris:
pct enter <id> DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade exit
Prevention
The old homelab apt-upgrade wrapper (retired along with the rest of the
homelab CLI) used to launch apt inside a systemd-run --collect unit on
the target so it survived ssh teardown — that's the failure mode that put
hubris into this state in the first place. There's no fleet-wide wrapper
anymore; if you run apt manually over ssh, wrap it yourself the same way:
ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'
Then systemctl status apt-recovery from a fresh ssh to check progress.
Related
- Operations cheatsheet
- Auto-deploy pipelines
- Hubris host page
Changelog
2026-05-21 — initial page
Documents the dpkg-interrupted recovery path that came out of the fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE web Shell).