Files
oikos/.agents/skills/runbook-dpkg-interrupted/SKILL.md
dtoro 986937799a archive: remove entire archive/ directory and all references
archive/ contained the old narrative wiki (superseded by DB as source
of truth), hermes-plans, oikos-cards, ledger, secrets-issuance, and
SOPS backups — all Python-era artifacts with no ongoing value.

Updated all cross-references in:
- AGENTS.md, README.md
- .agents/operations/commands.md (point to docs/infrastructure/)
- .agents/shared/llm-wiki.md, page-templates.md
- .agents/domains/knowledge/schema.md, operations/schema.md
- .agents/skills/*/SKILL.md
- docs/infrastructure/*.md (removed archive link targets)
- docs-lint/SKILL.md known-baseline note
2026-08-16 11:33:26 +02:00

118 lines
4.3 KiB
Markdown

---
name: recover-dpkg-interrupted
risk_class: reversible_low
verification: "dpkg --audit (should be clean); apt-get check"
---
# Runbook — recover from dpkg-interrupted state
You're here because an apt run got killed mid-transaction and the target now
has packages that are **unpacked but not configured**. Symptoms:
- `apt` refuses to do anything new: `Error: dpkg was interrupted, you must
manually run 'dpkg --configure -a' to correct the problem.`
- `dpkg --audit` lists packages with header
`The following packages have been unpacked but not yet configured.`
- `dpkg --audit` on the host directly shows unpacked-not-configured packages
(there's no fleet-wide audit tool anymore — check per-host).
The system is still running the **old** binaries (still in memory), but the
**new** binaries are unpacked and waiting for their postinst to run. Two
worst-case manifestations from the 2026-05-21 sweep:
- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was
still serving but the new caddy binary on disk hadn't been wired up.
- hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured,
including `systemd`, `openssh-server`, `sudo`, `netbird`. The half-
configured netbird daemon dropped the mesh peer, and we got locked out
until we recovered from the PVE web UI Shell.
**Do not reboot until dpkg is clean.** A reboot tries to start the new
binaries' services, which may fail because postinst never ran (missing users,
config dirs, capabilities, etc.). The system might not come back up cleanly.
## Path A — target is still reachable over ssh (preferred)
```
ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
```
Or for an LXC by name (via the MCP `run` tool, or directly on the Proxmox
host):
```
pct exec <lxc> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
```
When that returns, confirm:
```
ssh <host> -- dpkg --audit
```
Expect `DPKG: ok` and the remaining `UPGR` count to match what's intentionally
deferred (kernel/PVE on hubris, 0 elsewhere).
## Path B — target locked out (mesh broken / ssh dead)
Most common for hubris when netbird itself went half-configured: the daemon
crashed on the new binary, the mesh peer dropped, port 22022 stopped listening,
and you can't ssh in.
1. Open `https://proxmox.hubris.network` in a browser.
2. Datacenter → node `hubris` → `>_ Shell` (or `_ Console`). That's a root
shell on hubris served by the PVE web UI, independent of the netbird mesh.
3. Run the recovery one-liner:
```
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
&& DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
&& systemctl restart netbird \
&& dpkg --audit \
&& echo RECOVERY_OK
```
Wait for `RECOVERY_OK`. The `systemctl restart netbird` is the bit that
heals the mesh — once netbird's daemon comes back up clean, your client's
peer state moves from `Connecting` to `Connected` within ~30 seconds and
the rest of your tooling works again.
4. For an **LXC** that's locked out (less common — LXCs reach the world via
netbird routed through hubris, so unless hubris itself is broken, you can
still `pct enter` from the hubris shell):
From the PVE web UI shell on hubris:
```
pct enter <id>
DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
exit
```
## Prevention
The old `homelab apt-upgrade` wrapper (retired along with the rest of the
`homelab` CLI) used to launch apt inside a `systemd-run --collect` unit on
the target so it survived ssh teardown — that's the failure mode that put
hubris into this state in the first place. There's no fleet-wide wrapper
anymore; if you run apt manually over ssh, wrap it yourself the same way:
```
ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'
```
Then `systemctl status apt-recovery` from a fresh ssh to check progress.
## Related
- [Operations cheatsheet](../../operations/commands.md)
- [Auto-deploy pipelines](../../../docs/infrastructure/auto-deploy.md)
- Hubris host page
## Changelog
### 2026-05-21 — initial page
Documents the dpkg-interrupted recovery path that came out of the
fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE
web Shell).