Files
oikos/operations/commands.md
dtoro 8a6422bd7d docs: move narrative wiki under knowledge/wiki/ (phase 3)
Problem: node and cross-cutting narratives lived at the repo root
(containers/, vms/, infrastructure/, host .md files), interleaved with the
machine-readable substrate.

Change:
- Move containers/ -> knowledge/wiki/containers/, vms/ -> knowledge/wiki/vms/,
  infrastructure/ -> knowledge/wiki/infrastructure/, hosts/{hubris,strong}.md ->
  knowledge/wiki/hosts/, infrastructure/references/ -> knowledge/sources/references/,
  GLOSSARY.md -> knowledge/GLOSSARY.md.
- Add knowledge/{index.md,log.md,sources/index.md} scaffolding.
- Rewrite all relative links repo-wide via a path-resolving mapper (inbound +
  outbound + between-moved-files), including .hermes/, runbooks, operations,
  investigations, plans, README, AGENTS.
- Repoint inventory.yaml doc_page fields and regenerate hosts/*.yaml (which
  embed doc_page); update oikos/gen-topology.py output path, candidate doc
  paths, and footer links; update code-comment doc paths.

Substrate untouched in place: inventory.yaml, hosts/*.yaml (regenerated,
idempotent), oikos/ code, mcp/, secrets/, bin/.

Verification:
- Logical broken-link set identical to pre-move baseline (net 128 -> 127; the
  topology regen fixed one, introduced none). Remaining are pre-existing refs
  to destroyed/archived nodes, out of scope for this move.
- gen-topology.py --check exit 0 (in sync); cards carry knowledge/wiki/ doc paths.
- build_host_files.py idempotent; all inventory doc_page targets resolve.
- MCP contract verified: get_page/search_docs/get_changelog resolve moved pages.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-06 14:35:23 +02:00

6.4 KiB
Raw Blame History

Operations cheatsheet

Run from the hubris host as root. When working from /root on Linux you're already on hubris — don't ssh hubris / ping hubris.

Proxmox CLI

Command Use
pct list / qm list List LXC containers / VMs
pct config <id> / qm config <id> Container / VM config
pct exec <id> -- <cmd> Run command inside an LXC without entering it (no initgroups — see media permissions)
pct enter <id> Shell into a container
pct start <id> / pct stop <id> Boot / halt a container
pvesm status Storage pools status
pvesh get /nodes --output-format json Node summary as JSON
pvesh get /nodes/hubris/lxc/<id>/status/current Live container status
pvesh get /cluster/resources --type vm --output-format json Bulk per-LXC CPU/mem/disk (used by the homelab-health-watchdog Hermes cron — see monitoring; the old claudio-monitor this once fed is deprecated)
pveversion PVE version
journalctl -u pve-cluster -n 100 PVE service logs

Storage

  • Shared mount: /mnt/library (ext4 on lvmthin library).
  • Bind into a container: pct set <id> -mp<N> /mnt/library/<sub>,mp=/data
  • For the standard whole-tree mount: pct set <id> -mp0 /mnt/library,mp=/mnt/library. See media permissions for the GID-10000 onboarding recipe.

Reverse proxy

  • Caddyfile: /etc/caddy/Caddyfile on LXC 121.
  • CRITICAL: This file is tracked in dtoro/caddy-conf (https://git.hubris.network/dtoro/caddy-conf). Never edit it directly on the LXC — commit + push to the repo instead. Caddy auto-deploys on push (see auto-deploy). If you edit directly, the change will be lost on the next pull and agents won't know about it.
  • Hot reload: pct exec 121 -- systemctl reload caddy.
  • Validate: pct exec 121 -- caddy validate --config /etc/caddy/Caddyfile.
  • Git workflow shortcut: pct exec 121 -- "cd /etc/caddy && git add Caddyfile && git commit -m '...' && git push".

DNS

  • Split-horizon authority: Technitium DNS on dns (107) at 192.168.8.2:53. Web UI at http://192.168.8.2. (Formerly dnsmasq on the now-destroyed LXC 124 — decommissioned 2026-06-04.)
  • Add/edit records in the Technitium UI; the NetBird managed zone sync (scripts/dns-sync.py cron on 107) picks changes up within ~10 minutes.
  • Verify: dig @192.168.8.2 +short <host>.hubris.network.
  • See DNS.

Web access

  • https://proxmox.hubris.network or https://192.168.8.77:8006 — Proxmox UI

Telemetry quick checks

  • ras-mc-ctl --summary — summary of any RAS events (memory / PCIe AER / thermal) since boot
  • ras-mc-ctl --errors — full event log
  • cat /sys/devices/system/cpu/cpu0/cpufreq/energy_performance_preference — should be balance_power
  • cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor — should be powersave
  • ls /sys/fs/pstore/ /var/lib/systemd/pstore/ — panic traces from a previous crash (empty for pure hardware hangs — see investigation)

Fleet apt operations

Two homelab subcommands wrap the common patterns; both fan out to hubris + every LXC.

Command What it does
homelab apt-audit [--target HOST] Per-host table: dpkg-interrupted state, holds, upgradable count, non-apt binaries in system paths, DNS health. Exits nonzero if any host has dpkg-interrupted state.
homelab apt-upgrade --target HOST Launch apt update && apt upgrade inside a transient systemd-run --collect unit on the target. Survives ssh teardown. Apt configured with Acquire::Retries=3 + ForceIPv4=true.
homelab apt-upgrade --all Same, fanned out across the standard targets.
homelab apt-upgrade ... --status Show running unit + tail /var/log/homelab-apt-upgrade.log on each target.
homelab apt-upgrade ... --safe Take a pre-upgrade snapshot per LXC first (pct snapshotvzdump fallback for bind-mounted LXCs). Refuses if any snapshot fails unless --force.
homelab apt-upgrade ... --force Skip both the dpkg-audit gate and snapshot-failure refusal.

PVE/kernel deferral on hubris: homelab apt-upgrade --target hubris will try every upgrade, including kernel + pve-*. To skip those, apt-mark hold the relevant packages on hubris first; homelab apt-audit shows held packages so you can confirm.

Oikos (agent OS layer)

See OIKOS.md for the operating model. Quick reference:

Command What it does
homelab service <name> explain|health|docs|log|actions|history Service Console v0 — context card, cached health (--live to force a probe), docs, logs, safe actions + risk class, ledger history
homelab node <name> relations Ontology blast-radius query: what this host/service impacts, is affected by, and its full transitive blast radius
homelab change preflight <service> Dry-run report before mutating: risk class, current health, config repo, verification command
homelab decide <action> <entity> Decision classifier: risk × blast radius × confidence → auto-act or escalate
homelab signal list|raise|ack|resolve|mute The attention layer — pending updates, thresholds, drift, anything needing attention
homelab approval request|list|reply|check Escalate-route grants (Matrix-delivered via Hermes, or the Oikos Console's /approvals page)
homelab restart <service> [--approval-id <id>] --approval-id is required whenever the service's risk class needs approval (e.g. caddy, dns) — refuses mechanically without a valid grant

Oikos Console (read-mostly dashboard): oikos.hubris.network once deployed — see oikos/console/deploy/README.md.