Files
oikos/runbooks/runbook-dpkg-interrupted.md
dtoro fd35b48c8d Phase 1-4: full doc reorg
Phase 1 — fix stale state after strong migration (Phase 1+2, 2026-07-05)
  - README: corrected IPs (jellyfin 206→246, arriman 132→245, etc.),
    added missing containers (128 trmnl, 129 house, 133 seanime, 134 romm,
    124 authentik), updated last-refreshed date, added strong host context
  - containers/101-jellyfin.md: IP 206→246, host hubris→strong, mount
    /mnt/library→/mnt/media_local, GPU 760M→680M+RX7600, privilege→priv
  - containers/118-elementsynapse.md: IP 239→242, added Host: strong
  - containers/122-arriman.md: IP 132→245, mount→/mnt/media_local, added Host
  - containers/129-house.md: IP 212→244, added Host: strong
  - containers/130-grimmory.md: IP 213→247, mount→/mnt/media_local, added Host
  - containers/121-caddy.md: fixed site list (books→grimmory, removed auth→VPS,
    added house, roms, teddy, trmnl)
  - hosts/strong.md: updated At-a-glance to reflect 7 LXCs hosted
  - containers/123-claudio-bot.md, 127-mule-photos-new.md: archived to
    containers/archive/ (were destroyed LXCs with living pages)
  - inventory.yaml: verified correct — no changes needed

Phase 2 — structural cleanup
  - infrastructure/index.md: one-page overview of all cross-cutting systems
  - runbooks/: moved runbook-budget-from-csv.md and runbook-dpkg-interrupted.md
    from operations/ with YAML frontmatter added
  - plans/done/: moved 4 completed plans out of active view; updated index
  - vms/index.md: added VM index page

Phase 3 — navigation & discoverability
  - GLOSSARY.md: term definitions (Authentik, Caddy, LXC, VAAPI, etc.)
  - README: added table of contents, links to glossary + infrastructure index
  - investigations/: archived 2 resolved cases (crash-loop, authentik-migration)
    to investigations/archive/; updated index with active vs archived sections

Phase 4 — ongoing discipline
  - CONTRIBUTING.md: documented same-session update rule with explicit checklist
  - README: replaced full LXC table with summary + link to containers/index.md
    (single source of truth; de-duplication)
2026-07-06 00:46:27 +02:00

4.1 KiB

name, risk_class, verification
name risk_class verification
recover-dpkg-interrupted reversible_low dpkg --audit (should be clean); apt-get check

Runbook — recover from dpkg-interrupted state

You're here because an apt run got killed mid-transaction and the target now has packages that are unpacked but not configured. Symptoms:

  • apt refuses to do anything new: Error: dpkg was interrupted, you must manually run 'dpkg --configure -a' to correct the problem.
  • dpkg --audit lists packages with header The following packages have been unpacked but not yet configured.
  • homelab apt-audit shows DPKG: DIRTY(N) for the host.

The system is still running the old binaries (still in memory), but the new binaries are unpacked and waiting for their postinst to run. Two worst-case manifestations from the 2026-05-21 sweep:

  • LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was still serving but the new caddy binary on disk hadn't been wired up.
  • hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured, including systemd, openssh-server, sudo, netbird. The half- configured netbird daemon dropped the mesh peer, and we got locked out until we recovered from the PVE web UI Shell.

Do not reboot until dpkg is clean. A reboot tries to start the new binaries' services, which may fail because postinst never ran (missing users, config dirs, capabilities, etc.). The system might not come back up cleanly.

Path A — target is still reachable over ssh (preferred)

homelab ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'

Or for an LXC by name:

homelab pct <lxc> exec -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'

When that returns, confirm:

homelab apt-audit --target <host>

Expect DPKG: ok and the remaining UPGR count to match what's intentionally deferred (kernel/PVE on hubris, 0 elsewhere).

Path B — target locked out (mesh broken / ssh dead)

Most common for hubris when netbird itself went half-configured: the daemon crashed on the new binary, the mesh peer dropped, port 22022 stopped listening, and you can't ssh in.

  1. Open https://proxmox.hubris.network in a browser.
  2. Datacenter → node hubris>_ Shell (or _ Console). That's a root shell on hubris served by the PVE web UI, independent of the netbird mesh.
  3. Run the recovery one-liner:
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
  && DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
  && systemctl restart netbird \
  && dpkg --audit \
  && echo RECOVERY_OK

Wait for RECOVERY_OK. The systemctl restart netbird is the bit that heals the mesh — once netbird's daemon comes back up clean, your client's peer state moves from Connecting to Connected within ~30 seconds and the rest of your tooling works again.

  1. For an LXC that's locked out (less common — LXCs reach the world via netbird routed through hubris, so unless hubris itself is broken, you can still pct enter from the hubris shell):

    From the PVE web UI shell on hubris:

    pct enter <id>
    DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
    exit
    

Prevention

The homelab apt-upgrade wrapper launches apt inside a systemd-run --collect unit on the target, so it survives ssh teardown — the failure mode that put hubris into this state in the first place is no longer reachable through the standard tool. If you absolutely need to run apt manually over ssh, wrap it:

ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'

Then systemctl status apt-recovery from a fresh ssh to check progress.

Changelog

2026-05-21 — initial page

Documents the dpkg-interrupted recovery path that came out of the fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE web Shell).