Files
oikos/operations/runbook-dpkg-interrupted.md
dtoro 7a062bdb6b docs: dpkg-interrupted runbook + apt-fleet ops + non-apt-binary pattern + Claude Code settings + chat-sudo gotcha
Bundles the documentation slice of the apt-sweep backlog:

* operations/runbook-dpkg-interrupted.md (NEW) — Path A (ssh-reachable
  recovery) + Path B (PVE web Shell when the netbird mesh broke
  alongside the dpkg state, as happened during Wave 6 on hubris).
  Closes B2.

* operations/commands.md — new "Fleet apt operations" section
  documenting `homelab apt-audit` and `homelab apt-upgrade`
  (--status / --safe / --force). Adds the dpkg-interrupted runbook to
  Related.

* operations/agent-enrollment.md —
  - new "Claude Code permissions for fleet ops" section with the
    `permissions.allow` snippet (`Bash(ssh -p 22022 *)`,
    `Bash(homelab *)`) for `~/.claude/settings.json`. Closes A3.
  - two new Troubleshooting rows: chat-mode `!` sudo no-tty gotcha
    (G2) and the cosmetic netbird DNS-probe warning.

* infrastructure/auto-deploy.md — new "Custom-built binaries that
  overlap apt-managed paths" section describing the two acceptable
  patterns (epoch-versioned .deb à la caddy 1:2.11.3-hubris1; or
  apt-mark hold) and the discovery path via `homelab apt-audit`'s
  NONAPT column. Closes D3.

Remaining backlog after this commit: A4 (upstream OpenSSH/netbird mux
bug), D1 (apt-mark hold caddy in caddy-conf bootstrap — superseded
in practice by the epoch .deb), E1/E2 (LXC DNS fallback for
tailscale-managed resolv.conf), F1/F2 (vzdump fallback doc; F3 already
shipped).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 09:32:06 +02:00

109 lines
4.0 KiB
Markdown

# Runbook — recover from dpkg-interrupted state
You're here because an apt run got killed mid-transaction and the target now
has packages that are **unpacked but not configured**. Symptoms:
- `apt` refuses to do anything new: `Error: dpkg was interrupted, you must
manually run 'dpkg --configure -a' to correct the problem.`
- `dpkg --audit` lists packages with header
`The following packages have been unpacked but not yet configured.`
- `homelab apt-audit` shows `DPKG: DIRTY(N)` for the host.
The system is still running the **old** binaries (still in memory), but the
**new** binaries are unpacked and waiting for their postinst to run. Two
worst-case manifestations from the 2026-05-21 sweep:
- LXC 121 caddy: leftover state from a prior aborted apt run; caddy itself was
still serving but the new caddy binary on disk hadn't been wired up.
- hubris: ssh master died mid-Wave-6 → 135 packages unpacked-not-configured,
including `systemd`, `openssh-server`, `sudo`, `netbird`. The half-
configured netbird daemon dropped the mesh peer, and we got locked out
until we recovered from the PVE web UI Shell.
**Do not reboot until dpkg is clean.** A reboot tries to start the new
binaries' services, which may fail because postinst never ran (missing users,
config dirs, capabilities, etc.). The system might not come back up cleanly.
## Path A — target is still reachable over ssh (preferred)
```
homelab ssh <host> -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
```
Or for an LXC by name:
```
homelab pct <lxc> exec -- bash -c 'DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y -o Dpkg::Options::=--force-confold upgrade'
```
When that returns, confirm:
```
homelab apt-audit --target <host>
```
Expect `DPKG: ok` and the remaining `UPGR` count to match what's intentionally
deferred (kernel/PVE on hubris, 0 elsewhere).
## Path B — target locked out (mesh broken / ssh dead)
Most common for hubris when netbird itself went half-configured: the daemon
crashed on the new binary, the mesh peer dropped, port 22022 stopped listening,
and you can't ssh in.
1. Open `https://proxmox.hubris.network` in a browser.
2. Datacenter → node `hubris` → `>_ Shell` (or `_ Console`). That's a root
shell on hubris served by the PVE web UI, independent of the netbird mesh.
3. Run the recovery one-liner:
```
DEBIAN_FRONTEND=noninteractive dpkg --configure -a \
&& DEBIAN_FRONTEND=noninteractive apt -y -o Dpkg::Options::=--force-confold upgrade \
&& systemctl restart netbird \
&& dpkg --audit \
&& echo RECOVERY_OK
```
Wait for `RECOVERY_OK`. The `systemctl restart netbird` is the bit that
heals the mesh — once netbird's daemon comes back up clean, your client's
peer state moves from `Connecting` to `Connected` within ~30 seconds and
the rest of your tooling works again.
4. For an **LXC** that's locked out (less common — LXCs reach the world via
netbird routed through hubris, so unless hubris itself is broken, you can
still `pct enter` from the hubris shell):
From the PVE web UI shell on hubris:
```
pct enter <id>
DEBIAN_FRONTEND=noninteractive dpkg --configure -a && apt -y upgrade
exit
```
## Prevention
The `homelab apt-upgrade` wrapper launches apt inside a `systemd-run --collect`
unit on the target, so it survives ssh teardown — the failure mode that put
hubris into this state in the first place is no longer reachable through the
standard tool. If you absolutely need to run apt manually over ssh, wrap it:
```
ssh <host> systemd-run --unit=apt-recovery --collect bash -c 'apt -y upgrade'
```
Then `systemctl status apt-recovery` from a fresh ssh to check progress.
## Related
- [Operations cheatsheet](commands.md)
- [Auto-deploy pipelines](../infrastructure/auto-deploy.md)
- [Hubris host page](../hosts/hubris.md)
## Changelog
### 2026-05-21 — initial page
Documents the dpkg-interrupted recovery path that came out of the
fleet apt sweep (Wave 6 killed mid-transaction; hubris recovered via PVE
web Shell).