docs: dpkg-interrupted runbook + apt-fleet ops + non-apt-binary pattern + Claude Code settings + chat-sudo gotcha
Bundles the documentation slice of the apt-sweep backlog:
* operations/runbook-dpkg-interrupted.md (NEW) — Path A (ssh-reachable
recovery) + Path B (PVE web Shell when the netbird mesh broke
alongside the dpkg state, as happened during Wave 6 on hubris).
Closes B2.
* operations/commands.md — new "Fleet apt operations" section
documenting `homelab apt-audit` and `homelab apt-upgrade`
(--status / --safe / --force). Adds the dpkg-interrupted runbook to
Related.
* operations/agent-enrollment.md —
- new "Claude Code permissions for fleet ops" section with the
`permissions.allow` snippet (`Bash(ssh -p 22022 *)`,
`Bash(homelab *)`) for `~/.claude/settings.json`. Closes A3.
- two new Troubleshooting rows: chat-mode `!` sudo no-tty gotcha
(G2) and the cosmetic netbird DNS-probe warning.
* infrastructure/auto-deploy.md — new "Custom-built binaries that
overlap apt-managed paths" section describing the two acceptable
patterns (epoch-versioned .deb à la caddy 1:2.11.3-hubris1; or
apt-mark hold) and the discovery path via `homelab apt-audit`'s
NONAPT column. Closes D3.
Remaining backlog after this commit: A4 (upstream OpenSSH/netbird mux
bug), D1 (apt-mark hold caddy in caddy-conf bootstrap — superseded
in practice by the epoch .deb), E1/E2 (LXC DNS fallback for
tailscale-managed resolv.conf), F1/F2 (vzdump fallback doc; F3 already
shipped).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
108
operations/runbook-dpkg-interrupted.md
Normal file
108
operations/runbook-dpkg-interrupted.md
Normal file
@@ -0,0 +1,108 @@
|
||||
# Runbook — recover from dpkg-interrupted state
|
||||
|
||||
You're here because an apt run got killed mid-transaction and the target now
|
||||
has packages that are **unpacked but not configured**. Symptoms:
|
||||
|
||||
- `apt` refuses to do anything new: `Error: dpkg was interrupted, you must
|
||||
manually run 'dpkg --configure -a' to correct the problem.`
|
||||
- `dpkg --audit` lists packages with header
|
||||
`The following packages have been unpacked but not yet configured.`
|
||||
- `homelab apt-audit` shows `DPKG: DIRTY(N)` for the host.
|
||||
|
||||
The system is still running the **old** binaries (still in memory), but the
|
||||
**new** binaries are unpacked and waiting for their postinst to run. Two
|
||||
worst-case manifestations from the 2026-05-21 sweep:
|
||||
|
||||
- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was
|
||||
still serving but the new caddy binary on disk hadn't been wired up.
|
||||
- hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured,
|
||||
including `systemd`, `openssh-server`, `sudo`, `netbird`. The half-
|
||||
configured netbird daemon dropped the mesh peer, and we got locked out
|
||||
until we recovered from the PVE web UI Shell.
|
||||
|
||||
**Do not reboot until dpkg is clean.** A reboot tries to start the new
|
||||
binaries' services, which may fail because postinst never ran (missing users,
|
||||
config dirs, capabilities, etc.). The system might not come back up cleanly.
|
||||
|
||||
## Path A — target is still reachable over ssh (preferred)
|
||||
|
||||
```
|
||||
homelab ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
|
||||
```
|
||||
|
||||
Or for an LXC by name:
|
||||
|
||||
```
|
||||
homelab pct <lxc> exec -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
|
||||
```
|
||||
|
||||
When that returns, confirm:
|
||||
|
||||
```
|
||||
homelab apt-audit --target <host>
|
||||
```
|
||||
|
||||
Expect `DPKG: ok` and the remaining `UPGR` count to match what's intentionally
|
||||
deferred (kernel/PVE on hubris, 0 elsewhere).
|
||||
|
||||
## Path B — target locked out (mesh broken / ssh dead)
|
||||
|
||||
Most common for hubris when netbird itself went half-configured: the daemon
|
||||
crashed on the new binary, the mesh peer dropped, port 22022 stopped listening,
|
||||
and you can't ssh in.
|
||||
|
||||
1. Open `https://proxmox.hubris.network` in a browser.
|
||||
2. Datacenter → node `hubris` → `>_ Shell` (or `_ Console`). That's a root
|
||||
shell on hubris served by the PVE web UI, independent of the netbird mesh.
|
||||
3. Run the recovery one-liner:
|
||||
|
||||
```
|
||||
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
|
||||
&& DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
|
||||
&& systemctl restart netbird \
|
||||
&& dpkg --audit \
|
||||
&& echo RECOVERY_OK
|
||||
```
|
||||
|
||||
Wait for `RECOVERY_OK`. The `systemctl restart netbird` is the bit that
|
||||
heals the mesh — once netbird's daemon comes back up clean, your client's
|
||||
peer state moves from `Connecting` to `Connected` within ~30 seconds and
|
||||
the rest of your tooling works again.
|
||||
|
||||
4. For an **LXC** that's locked out (less common — LXCs reach the world via
|
||||
netbird routed through hubris, so unless hubris itself is broken, you can
|
||||
still `pct enter` from the hubris shell):
|
||||
|
||||
From the PVE web UI shell on hubris:
|
||||
|
||||
```
|
||||
pct enter <id>
|
||||
DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
|
||||
exit
|
||||
```
|
||||
|
||||
## Prevention
|
||||
|
||||
The `homelab apt-upgrade` wrapper launches apt inside a `systemd-run --collect`
|
||||
unit on the target, so it survives ssh teardown — the failure mode that put
|
||||
hubris into this state in the first place is no longer reachable through the
|
||||
standard tool. If you absolutely need to run apt manually over ssh, wrap it:
|
||||
|
||||
```
|
||||
ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'
|
||||
```
|
||||
|
||||
Then `systemctl status apt-recovery` from a fresh ssh to check progress.
|
||||
|
||||
## Related
|
||||
|
||||
- [Operations cheatsheet](commands.md)
|
||||
- [Auto-deploy pipelines](../infrastructure/auto-deploy.md)
|
||||
- [Hubris host page](../hosts/hubris.md)
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-05-21 — initial page
|
||||
Documents the dpkg-interrupted recovery path that came out of the
|
||||
fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE
|
||||
web Shell).
|
||||
Reference in New Issue
Block a user