Files
oikos/hosts/hubris.md

10 KiB
Raw Blame History

hubris — Proxmox host

Single-node Proxmox VE running 1 VM and 14 LXC containers (one of which is currently stopped). The whole homelab.

At a glance

  • Role: Proxmox VE 9.1.2 hypervisor (kernel 6.14.11-4-pve)
  • Hardware: GMKtec NucBox M6 Ultra — AMD Ryzen 5 7640HS (Phoenix APU), 12 vCPU / ~28 GiB RAM, 2× Samsung 990 EVO Plus NVMe (one SSD primary, one for library LVM). 2× Realtek RTL8125 NICs (r8169).
  • BIOS: 1.02 (2025-08-06) — vendor not on LVFS, no automated update path. See investigations.
  • LAN (primary): 192.168.8.77/24 on bridge vmbr0 (slave: eno1), gateway 192.168.8.1. Default route metric 0.
  • WiFi (failover): 192.168.8.141/24 on wlp3s0 (MediaTek MT7922, AX), DHCP from the same router. Default route metric 200. See Phase 1 WiFi failover below.
  • Mesh: Netbird wt0 100.122.38.109/16. Resolver: 100.122.38.109 (the local netbird daemon, which forwards to LAN/upstream and learns *.hubris.network answers via that path). See mesh.
  • UI: https://proxmox.hubris.network (via caddy) or https://192.168.8.77:8006.

Storage

Pool Type Size Use
local dir ~95G ISOs, templates, /etc, configs
local-lvm lvmthin 856G LXC/VM rootfs
library lvmthin 3.7T Backs /mnt/library ext4 (mounted as /dev/mapper/library-library)

/mnt/library holds the shared media + data pool: anime, audiobooks, books, comics, documents, downloads, heaper, homecloud, images, marimo, movies, music, notes, podcasts, repos, roms, sophia, syncthing. Bind-mounted into every container that needs it. Permissions standard: media GID 10000.

Tenants

VMs

LXC containers

See containers/index. 14 active, 1 stopped.

Boot-time tuning (load-bearing)

  • cpu-epp.service (enabled) sets scaling_governor=powersave + EPP=balance_power at boot — drops idle Tctl from ~95 °C → ~60 °C, clocks from 4.4 GHz pinned → 1.12 GHz idle. In amd-pstate=active mode the governor MUST be powersave, not performance, for EPP to apply. Unit ordering: After=sysinit.target + Before=pve-guests.service so it runs before the guest fleet starts (the original After=multi-user.target left the hottest boot window on performance). Fixed 2026-04-22.
  • Crash capture: /etc/sysctl.d/60-crash-capture.conf (panic on oops/hardlockup/softlockup/rcu, auto-reboot 10 s), /etc/modprobe.d/softdog.conf (soft_panic=1 soft_margin=60), /etc/systemd/system.conf.d/watchdog.conf (RuntimeWatchdogSec=15s). Pstore traces collected to /var/lib/systemd/pstore/ by systemd-pstore.service. Caveat: silent CPU lockups leave pstore empty.
  • rasdaemon (Debian pkg) collects MCE / memory / PCIe AER / thermal events to /var/lib/rasdaemon/ras-mc_event.db. Query with ras-mc-ctl --summary / --errors. (mcelog is retired in Debian 13 — don't go looking for it.)

Phase 1 WiFi failover

Host is dual-homed on LAN (eno1/vmbr0) and WiFi (wlp3s0) so management/SSH stay reachable when LAN drops. Guests are not yet failed over — the LXC fleet remains on vmbr0/eno1. Phase 2 will migrate guest networking off the bridge so the homelab survives full LAN loss.

  • WiFi creds in /etc/wpa_supplicant/wpa_supplicant-wlp3s0.conf (hashed PSK, mode 600). SSID lives in /etc/network/interfaces as wpa-conf.
  • Both interfaces sit on the same 192.168.8.0/24; cross-talk avoided with arp_ignore=1 + arp_announce=2 on eno1/vmbr0/wlp3s0 (set via post-up in /etc/network/interfaces).
  • A second default route at metric 200 is added on wlp3s0 (post-up). LAN wins while up.
  • Carrier-based failover: vmbr0's carrier follows the LXC veth members, so it stays 1 even when eno1 loses link. ignore_routes_with_linkdown is therefore not enough on its own. wan-failover.service (/usr/local/sbin/wan-failover.sh) watches /sys/class/net/eno1/carrier via ip monitor link and removes/restores the vmbr0 default route on transitions. Logs to journalctl -t wan-failover.
  • Failover verified 2026-04-28: ip link set eno1 down → outbound HTTP keeps working via WiFi; ip link set eno1 up → vmbr0 default restored.
  • Reachable on 192.168.8.77 (LAN) and 192.168.8.141 (WiFi); SSH works on either.

Host services owned by external repos

What Repo Path on host
claudio-monitor (5-min watchdog) dtoro/claudio-monitor /opt/claudio-monitor
backup-library (restic) dtoro/backup-library /opt/backup-library (disabled)
hubris-public-cert-sync (script + systemd unit) /usr/local/bin/hubris-public-cert-sync.sh, daily timer

See monitoring, backups, ingress.

Quirks

  • /etc/pve is fuse — normal for the Proxmox cluster filesystem, even on a single-node install.
  • ZFS is not in use; storage is LVM-thin + ext4.
  • Two Realtek 8125 NICs use the in-tree r8169 driver, not the OOT r8125.
  • Hardware is thermally marginal. NVMe sensors live near warn temp under load. Thermal pads installed on the SSDs 2026-04-23; host relocated to a better-ventilated spot 2026-04-29.
  • Non-ECC RAM: silent memory faults are possible; suspect DIMMs if cpu-epp is on and crashes still happen.
  • BIOS update path requires a Windows-To-Go USB (no LVFS, no Linux flasher).

Authorized SSH keys (root)

  • root@hubris (self, RSA) — local
  • d.toro.v@pm.me (ed25519) — user's iMac, added 2026-04-22

OpenSSH on 0.0.0.0:22. Netbird's built-in SSH server is on 100.122.38.109:22022 and bypasses authorized_keys (OIDC/browser). See SSH access for the dual-server gotcha.

Changelog

2026-04-29 — relocated to better-ventilated spot

User physically moved the host to a new location with improved airflow. Post-move idle baseline (45 min uptime, light load): k10temp Tctl 47.2 °C, amdgpu edge 42 °C, nvme0 composite 34.9 °C / sensor1 32.9 °C, nvme1 composite 38.9 °C / sensor1 52.9 °C, DRAM 3435.5 °C, ACPI zone 4749 °C. Compares well against the 2026-04-23 thermal-pad steady-state (nvme0 sensor1 6061 °C). Watch the lifetime NVMe warning-time counter over the coming days for confirmation. See investigation.

2026-04-28 — Phase 1 WiFi failover

Host now dual-homed: LAN 192.168.8.77 (primary) + WiFi 192.168.8.141 (failover, metric 200) on the GL-AXT1800-714-5G AP. Installed wpasupplicant+iw; added wlp3s0 stanza to /etc/network/interfaces with wpa-conf; ARP isolation sysctls in post-up. Built wan-failover.service to remove the vmbr0 default route on eno1 carrier loss, since the bridge's carrier doesn't follow eno1 (the LXC veths keep it 1). LXC/VM guests are still LAN-only — Phase 2 will migrate them.

2026-04-28 — wiki started

This wiki created. Live state at this date: 14 LXCs running (109 syncthing stopped), 1 VM, kernel 6.14.11-4-pve, uptime 3 d 0 h post drive-removal A/B test. Compared to memory snapshot from a week ago, destroyed: LXC 100 (yunohost arr), 106 (flaresolverr), 107 (marimo), 110 (photoprism), 111 (karakeep), 112 (immich), 115 (reticulum). 100 + 106 destroyed per the planned 2026-04-21 *arr migration retention; the others removed since.

2026-04-23 — SSD cooling + thermal pads installed

Thermal pads on both NVMe drives. Steady-state nvme0 composite 47 °C / sensor1 6061 °C, nvme1 3840 °C. Zero new warning-time minutes after install. Watch the lifetime warning-time counter going forward, not absolute sensor1. See investigation.

2026-04-22 — drive removal A/B test

Removed external USB backup drive (Silicon Motion 090c:2320). Disabled the four backup-library*.timer units, commented the fstab entry. Goal: confirm whether the drive + UAS interaction on the AMD USB4 PCIe tunnel is the dominant root cause of the silent hard-locks. Pre-drive uptime was 33 days; with drive, repeated crashes despite UAS blacklist + mount-on-demand. Result so far: 3+ days uptime — the drive looks like the primary contributor; cpu-epp remains as belt-and-suspenders thermal protection. See investigation.

2026-04-22 — cpu-epp.service ordering bug fixed

Was After=multi-user.target + WantedBy=multi-user.target — queued behind pve-guests.service, so the hottest boot window (20+ guests starting on performance) preceded EPP application. Now After=sysinit.target + Before=pve-guests.service.

2026-04-21 — crash-capture + RAS telemetry enabled

60-crash-capture.conf, softdog soft_panic=1, RuntimeWatchdog 15 s. rasdaemon installed and enabled. Pure silicon hangs still leave no trace; this catches everything else.

2026-04-21 — cpu-epp.service deployed

Pinned governor=powersave, EPP=balance_power at boot. Stopped the host idling at ~95 °C with everything pinned at 4.4 GHz. First fix in the crash-loop incident.