Bundles the documentation slice of the apt-sweep backlog:
* operations/runbook-dpkg-interrupted.md (NEW) — Path A (ssh-reachable
recovery) + Path B (PVE web Shell when the netbird mesh broke
alongside the dpkg state, as happened during Wave 6 on hubris).
Closes B2.
* operations/commands.md — new "Fleet apt operations" section
documenting `homelab apt-audit` and `homelab apt-upgrade`
(--status / --safe / --force). Adds the dpkg-interrupted runbook to
Related.
* operations/agent-enrollment.md —
- new "Claude Code permissions for fleet ops" section with the
`permissions.allow` snippet (`Bash(ssh -p 22022 *)`,
`Bash(homelab *)`) for `~/.claude/settings.json`. Closes A3.
- two new Troubleshooting rows: chat-mode `!` sudo no-tty gotcha
(G2) and the cosmetic netbird DNS-probe warning.
* infrastructure/auto-deploy.md — new "Custom-built binaries that
overlap apt-managed paths" section describing the two acceptable
patterns (epoch-versioned .deb à la caddy 1:2.11.3-hubris1; or
apt-mark hold) and the discovery path via `homelab apt-audit`'s
NONAPT column. Closes D3.
Remaining backlog after this commit: A4 (upstream OpenSSH/netbird mux
bug), D1 (apt-mark hold caddy in caddy-conf bootstrap — superseded
in practice by the epoch .deb), E1/E2 (LXC DNS fallback for
tailscale-managed resolv.conf), F1/F2 (vzdump fallback doc; F3 already
shipped).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
109 lines
4.0 KiB
Markdown
109 lines
4.0 KiB
Markdown
# Runbook — recover from dpkg-interrupted state
|
|
|
|
You're here because an apt run got killed mid-transaction and the target now
|
|
has packages that are **unpacked but not configured**. Symptoms:
|
|
|
|
- `apt` refuses to do anything new: `Error: dpkg was interrupted, you must
|
|
manually run 'dpkg --configure -a' to correct the problem.`
|
|
- `dpkg --audit` lists packages with header
|
|
`The following packages have been unpacked but not yet configured.`
|
|
- `homelab apt-audit` shows `DPKG: DIRTY(N)` for the host.
|
|
|
|
The system is still running the **old** binaries (still in memory), but the
|
|
**new** binaries are unpacked and waiting for their postinst to run. Two
|
|
worst-case manifestations from the 2026-05-21 sweep:
|
|
|
|
- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was
|
|
still serving but the new caddy binary on disk hadn't been wired up.
|
|
- hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured,
|
|
including `systemd`, `openssh-server`, `sudo`, `netbird`. The half-
|
|
configured netbird daemon dropped the mesh peer, and we got locked out
|
|
until we recovered from the PVE web UI Shell.
|
|
|
|
**Do not reboot until dpkg is clean.** A reboot tries to start the new
|
|
binaries' services, which may fail because postinst never ran (missing users,
|
|
config dirs, capabilities, etc.). The system might not come back up cleanly.
|
|
|
|
## Path A — target is still reachable over ssh (preferred)
|
|
|
|
```
|
|
homelab ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
|
|
```
|
|
|
|
Or for an LXC by name:
|
|
|
|
```
|
|
homelab pct <lxc> exec -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
|
|
```
|
|
|
|
When that returns, confirm:
|
|
|
|
```
|
|
homelab apt-audit --target <host>
|
|
```
|
|
|
|
Expect `DPKG: ok` and the remaining `UPGR` count to match what's intentionally
|
|
deferred (kernel/PVE on hubris, 0 elsewhere).
|
|
|
|
## Path B — target locked out (mesh broken / ssh dead)
|
|
|
|
Most common for hubris when netbird itself went half-configured: the daemon
|
|
crashed on the new binary, the mesh peer dropped, port 22022 stopped listening,
|
|
and you can't ssh in.
|
|
|
|
1. Open `https://proxmox.hubris.network` in a browser.
|
|
2. Datacenter → node `hubris` → `>_ Shell` (or `_ Console`). That's a root
|
|
shell on hubris served by the PVE web UI, independent of the netbird mesh.
|
|
3. Run the recovery one-liner:
|
|
|
|
```
|
|
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
|
|
&& DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
|
|
&& systemctl restart netbird \
|
|
&& dpkg --audit \
|
|
&& echo RECOVERY_OK
|
|
```
|
|
|
|
Wait for `RECOVERY_OK`. The `systemctl restart netbird` is the bit that
|
|
heals the mesh — once netbird's daemon comes back up clean, your client's
|
|
peer state moves from `Connecting` to `Connected` within ~30 seconds and
|
|
the rest of your tooling works again.
|
|
|
|
4. For an **LXC** that's locked out (less common — LXCs reach the world via
|
|
netbird routed through hubris, so unless hubris itself is broken, you can
|
|
still `pct enter` from the hubris shell):
|
|
|
|
From the PVE web UI shell on hubris:
|
|
|
|
```
|
|
pct enter <id>
|
|
DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
|
|
exit
|
|
```
|
|
|
|
## Prevention
|
|
|
|
The `homelab apt-upgrade` wrapper launches apt inside a `systemd-run --collect`
|
|
unit on the target, so it survives ssh teardown — the failure mode that put
|
|
hubris into this state in the first place is no longer reachable through the
|
|
standard tool. If you absolutely need to run apt manually over ssh, wrap it:
|
|
|
|
```
|
|
ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'
|
|
```
|
|
|
|
Then `systemctl status apt-recovery` from a fresh ssh to check progress.
|
|
|
|
## Related
|
|
|
|
- [Operations cheatsheet](commands.md)
|
|
- [Auto-deploy pipelines](../infrastructure/auto-deploy.md)
|
|
- [Hubris host page](../hosts/hubris.md)
|
|
|
|
## Changelog
|
|
|
|
### 2026-05-21 — initial page
|
|
Documents the dpkg-interrupted recovery path that came out of the
|
|
fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE
|
|
web Shell).
|