Files
oikos/.agents/skills/runbook-dpkg-interrupted/SKILL.md
dtoro 986937799a archive: remove entire archive/ directory and all references
archive/ contained the old narrative wiki (superseded by DB as source
of truth), hermes-plans, oikos-cards, ledger, secrets-issuance, and
SOPS backups — all Python-era artifacts with no ongoing value.

Updated all cross-references in:
- AGENTS.md, README.md
- .agents/operations/commands.md (point to docs/infrastructure/)
- .agents/shared/llm-wiki.md, page-templates.md
- .agents/domains/knowledge/schema.md, operations/schema.md
- .agents/skills/*/SKILL.md
- docs/infrastructure/*.md (removed archive link targets)
- docs-lint/SKILL.md known-baseline note
2026-08-16 11:33:26 +02:00

4.3 KiB

name, risk_class, verification
name risk_class verification
recover-dpkg-interrupted reversible_low dpkg --audit (should be clean); apt-get check

Runbook — recover from dpkg-interrupted state

You're here because an apt run got killed mid-transaction and the target now has packages that are unpacked but not configured. Symptoms:

  • apt refuses to do anything new: Error: dpkg was interrupted, you must manually run 'dpkg --configure -a' to correct the problem.
  • dpkg --audit lists packages with header The following packages have been unpacked but not yet configured.
  • dpkg --audit on the host directly shows unpacked-not-configured packages (there's no fleet-wide audit tool anymore — check per-host).

The system is still running the old binaries (still in memory), but the new binaries are unpacked and waiting for their postinst to run. Two worst-case manifestations from the 2026-05-21 sweep:

  • LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was still serving but the new caddy binary on disk hadn't been wired up.
  • hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured, including systemd, openssh-server, sudo, netbird. The half- configured netbird daemon dropped the mesh peer, and we got locked out until we recovered from the PVE web UI Shell.

Do not reboot until dpkg is clean. A reboot tries to start the new binaries' services, which may fail because postinst never ran (missing users, config dirs, capabilities, etc.). The system might not come back up cleanly.

Path A — target is still reachable over ssh (preferred)

ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'

Or for an LXC by name (via the MCP run tool, or directly on the Proxmox host):

pct exec <lxc> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'

When that returns, confirm:

ssh <host> -- dpkg --audit

Expect DPKG: ok and the remaining UPGR count to match what's intentionally deferred (kernel/PVE on hubris, 0 elsewhere).

Path B — target locked out (mesh broken / ssh dead)

Most common for hubris when netbird itself went half-configured: the daemon crashed on the new binary, the mesh peer dropped, port 22022 stopped listening, and you can't ssh in.

  1. Open https://proxmox.hubris.network in a browser.
  2. Datacenter → node hubris>_ Shell (or _ Console). That's a root shell on hubris served by the PVE web UI, independent of the netbird mesh.
  3. Run the recovery one-liner:
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
  && DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
  && systemctl restart netbird \
  && dpkg --audit \
  && echo RECOVERY_OK

Wait for RECOVERY_OK. The systemctl restart netbird is the bit that heals the mesh — once netbird's daemon comes back up clean, your client's peer state moves from Connecting to Connected within ~30 seconds and the rest of your tooling works again.

  1. For an LXC that's locked out (less common — LXCs reach the world via netbird routed through hubris, so unless hubris itself is broken, you can still pct enter from the hubris shell):

    From the PVE web UI shell on hubris:

    pct enter <id>
    DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
    exit
    

Prevention

The old homelab apt-upgrade wrapper (retired along with the rest of the homelab CLI) used to launch apt inside a systemd-run --collect unit on the target so it survived ssh teardown — that's the failure mode that put hubris into this state in the first place. There's no fleet-wide wrapper anymore; if you run apt manually over ssh, wrap it yourself the same way:

ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'

Then systemctl status apt-recovery from a fresh ssh to check progress.

Changelog

2026-05-21 — initial page

Documents the dpkg-interrupted recovery path that came out of the fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE web Shell).