Phase 1 — fix stale state after strong migration (Phase 1+2, 2026-07-05)
- README: corrected IPs (jellyfin 206→246, arriman 132→245, etc.),
added missing containers (128 trmnl, 129 house, 133 seanime, 134 romm,
124 authentik), updated last-refreshed date, added strong host context
- containers/101-jellyfin.md: IP 206→246, host hubris→strong, mount
/mnt/library→/mnt/media_local, GPU 760M→680M+RX7600, privilege→priv
- containers/118-elementsynapse.md: IP 239→242, added Host: strong
- containers/122-arriman.md: IP 132→245, mount→/mnt/media_local, added Host
- containers/129-house.md: IP 212→244, added Host: strong
- containers/130-grimmory.md: IP 213→247, mount→/mnt/media_local, added Host
- containers/121-caddy.md: fixed site list (books→grimmory, removed auth→VPS,
added house, roms, teddy, trmnl)
- hosts/strong.md: updated At-a-glance to reflect 7 LXCs hosted
- containers/123-claudio-bot.md, 127-mule-photos-new.md: archived to
containers/archive/ (were destroyed LXCs with living pages)
- inventory.yaml: verified correct — no changes needed
Phase 2 — structural cleanup
- infrastructure/index.md: one-page overview of all cross-cutting systems
- runbooks/: moved runbook-budget-from-csv.md and runbook-dpkg-interrupted.md
from operations/ with YAML frontmatter added
- plans/done/: moved 4 completed plans out of active view; updated index
- vms/index.md: added VM index page
Phase 3 — navigation & discoverability
- GLOSSARY.md: term definitions (Authentik, Caddy, LXC, VAAPI, etc.)
- README: added table of contents, links to glossary + infrastructure index
- investigations/: archived 2 resolved cases (crash-loop, authentik-migration)
to investigations/archive/; updated index with active vs archived sections
Phase 4 — ongoing discipline
- CONTRIBUTING.md: documented same-session update rule with explicit checklist
- README: replaced full LXC table with summary + link to containers/index.md
(single source of truth; de-duplication)
115 lines
4.1 KiB
Markdown
115 lines
4.1 KiB
Markdown
---
|
|
name: recover-dpkg-interrupted
|
|
risk_class: reversible_low
|
|
verification: "dpkg --audit (should be clean); apt-get check"
|
|
---
|
|
|
|
# Runbook — recover from dpkg-interrupted state
|
|
|
|
You're here because an apt run got killed mid-transaction and the target now
|
|
has packages that are **unpacked but not configured**. Symptoms:
|
|
|
|
- `apt` refuses to do anything new: `Error: dpkg was interrupted, you must
|
|
manually run 'dpkg --configure -a' to correct the problem.`
|
|
- `dpkg --audit` lists packages with header
|
|
`The following packages have been unpacked but not yet configured.`
|
|
- `homelab apt-audit` shows `DPKG: DIRTY(N)` for the host.
|
|
|
|
The system is still running the **old** binaries (still in memory), but the
|
|
**new** binaries are unpacked and waiting for their postinst to run. Two
|
|
worst-case manifestations from the 2026-05-21 sweep:
|
|
|
|
- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was
|
|
still serving but the new caddy binary on disk hadn't been wired up.
|
|
- hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured,
|
|
including `systemd`, `openssh-server`, `sudo`, `netbird`. The half-
|
|
configured netbird daemon dropped the mesh peer, and we got locked out
|
|
until we recovered from the PVE web UI Shell.
|
|
|
|
**Do not reboot until dpkg is clean.** A reboot tries to start the new
|
|
binaries' services, which may fail because postinst never ran (missing users,
|
|
config dirs, capabilities, etc.). The system might not come back up cleanly.
|
|
|
|
## Path A — target is still reachable over ssh (preferred)
|
|
|
|
```
|
|
homelab ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
|
|
```
|
|
|
|
Or for an LXC by name:
|
|
|
|
```
|
|
homelab pct <lxc> exec -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
|
|
```
|
|
|
|
When that returns, confirm:
|
|
|
|
```
|
|
homelab apt-audit --target <host>
|
|
```
|
|
|
|
Expect `DPKG: ok` and the remaining `UPGR` count to match what's intentionally
|
|
deferred (kernel/PVE on hubris, 0 elsewhere).
|
|
|
|
## Path B — target locked out (mesh broken / ssh dead)
|
|
|
|
Most common for hubris when netbird itself went half-configured: the daemon
|
|
crashed on the new binary, the mesh peer dropped, port 22022 stopped listening,
|
|
and you can't ssh in.
|
|
|
|
1. Open `https://proxmox.hubris.network` in a browser.
|
|
2. Datacenter → node `hubris` → `>_ Shell` (or `_ Console`). That's a root
|
|
shell on hubris served by the PVE web UI, independent of the netbird mesh.
|
|
3. Run the recovery one-liner:
|
|
|
|
```
|
|
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
|
|
&& DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
|
|
&& systemctl restart netbird \
|
|
&& dpkg --audit \
|
|
&& echo RECOVERY_OK
|
|
```
|
|
|
|
Wait for `RECOVERY_OK`. The `systemctl restart netbird` is the bit that
|
|
heals the mesh — once netbird's daemon comes back up clean, your client's
|
|
peer state moves from `Connecting` to `Connected` within ~30 seconds and
|
|
the rest of your tooling works again.
|
|
|
|
4. For an **LXC** that's locked out (less common — LXCs reach the world via
|
|
netbird routed through hubris, so unless hubris itself is broken, you can
|
|
still `pct enter` from the hubris shell):
|
|
|
|
From the PVE web UI shell on hubris:
|
|
|
|
```
|
|
pct enter <id>
|
|
DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
|
|
exit
|
|
```
|
|
|
|
## Prevention
|
|
|
|
The `homelab apt-upgrade` wrapper launches apt inside a `systemd-run --collect`
|
|
unit on the target, so it survives ssh teardown — the failure mode that put
|
|
hubris into this state in the first place is no longer reachable through the
|
|
standard tool. If you absolutely need to run apt manually over ssh, wrap it:
|
|
|
|
```
|
|
ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'
|
|
```
|
|
|
|
Then `systemctl status apt-recovery` from a fresh ssh to check progress.
|
|
|
|
## Related
|
|
|
|
- [Operations cheatsheet](commands.md)
|
|
- [Auto-deploy pipelines](../infrastructure/auto-deploy.md)
|
|
- [Hubris host page](../hosts/hubris.md)
|
|
|
|
## Changelog
|
|
|
|
### 2026-05-21 — initial page
|
|
Documents the dpkg-interrupted recovery path that came out of the
|
|
fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE
|
|
web Shell).
|