Files
oikos/.hermes/plans/2026-07-05_strong-migration-assessment.md

16 KiB
Raw Permalink Blame History

Assessment: Which nodes can move to strong

Executive summary

hubris is memory-starved: 28 GiB RAM, 71.9 GiB allocated across 18 LXC + 2 VM (2.5× overcommit), 10 GiB swap in active use. strong sits completely empty — 28 GiB RAM, 25 GiB free, 0 guests, 2.7 TiB unused storage. The single most effective decongestion move is to shift guests off hubris onto strong.

This document assesses every guest for move-readiness, grouped by constraints (library dependency, GPU, core-infra status), and proposes a phased migration that does not require the physical library-SSD move (the blocker of the original plan) — library access from strong is provided via NFS from hubris.


Current resource state (live, 2026-07-05)

hubris — overloaded

Resource Capacity Allocated (all guests) Actual use Status
RAM 28 GiB 70.2 GiB (2.5× overcommit) 18 GiB used + 10 GiB swap ⚠️ heavy swap pressure
CPU 12 vCPU (6c/12t) 43 vCPU (3.6× overcommit) ~43% scaling MHz OK (shares)
local-lvm 856 GiB 582 GiB allocated (29% thin) OK
library (lvmthin) 3.7 TiB 1.2 TiB used (34%) OK, 2.3 TiB free

strong — empty, ready

Resource Capacity Used Status
RAM 28 GiB 2.3 GiB (host only) 25 GiB free
CPU 16 vCPU (8c/16t) idle (0.08 load) 100% free
local-lvm 856 GiB 0 empty
ludo-lvm 1.8 TiB 0 empty
Guests 0 LXC, 0 VM nothing running

Network topology constraint

Fritz!Box (192.168.178.1)
  └── SODOLA 2.5G switch
       ├── hubris eno1 → vmbr1 (192.168.178.10) → vmbr0 (192.168.8.0/24)
       │     └── all 20 guests on 192.168.8.x
       └── strong vmbr0 (192.168.178.181)
             └── no internal bridge yet, guests would be on 192.168.178.x

strong reaches 192.168.8.0/24 via the Fritz static route through hubris. Guests on strong get 192.168.178.x IPs unless we add an internal bridge on strong (Phase 0 prerequisite — see below).


Per-guest assessment

Tier 1 — Move immediately (no library dependency, no core-infra)

These guests mount no /mnt/library and are not part of the core infrastructure spine (caddy/dns/auth/mcp). They are the easiest wins.

ID Name Cores RAM Library? GPU? Notes
118 elementsynapse 2 4 GiB Matrix homeserver. Public via Caddy (matrix.hubris.network). Only change: Caddy backend IP. Easiest move in the fleet.
129 house 2 3 GiB Yuvomi family planner (Docker). No public Caddy route yet (uses VPS traefik directly). Self-contained.

Combined RAM freed from hubris: 7 GiB. No NFS, no library, no GPU.

Tier 2 — Move with library NFS (high resource consumers)

These are the heaviest guests and the original migration plan's primary targets. They mount /mnt/library and two use the iGPU. Moving them requires an NFS export from hubris → strong (reverse of the original plan's direction, since the physical SSD hasn't moved).

ID Name Cores RAM Library? GPU? I/O profile Notes
120 mule-images 6 12 GiB mp0 iGPU Write-heavy (photo processing) #1 RAM consumer. Strong has Radeon 680M iGPU (VAAPI works).
122 arriman 4 8 GiB mp0 Write-heavy (downloads) *arr stack + qbit + sab. Mounts library for download writes.
101 jellyfin 4 8 GiB mp0 iGPU Read-heavy sequential Media streaming + transcode. Strong 680M handles VAAPI.

Combined RAM freed: 28 GiB. This alone would eliminate hubris's swap pressure entirely.

Tier 3 — Could move, low urgency

ID Name Cores RAM Library? Notes
130 grimmory 1 2 GiB mp0 Book library (Docker). Migrated from apps LXC recently.
131 teddycloud 1 1 GiB mp0 New (not in inventory.yaml yet).
132 rclone 1 2 GiB mp0 (ro) Backup container. Read-only library mount.
128 trmnl 1 768 MiB TRMNL middleware. No library. Could move but tiny.
119 sophia 2 1 GiB mp0 Workshop. Light use.

Stay on hubris (core infrastructure)

ID Name Cores RAM Why it stays
121 caddy 1 512 MiB Reverse proxy — terminates all *.hubris.network. Must stay on hubris for LAN-side reachability. Needs backend IP updates when guests move.
107 dns 1 1 GiB Technitium DNS, split-horizon. Core.
106 auth-outpost 1 512 MiB Authentik SSO enforcement. Core.
105 apps 2 4 GiB homelab MCP + secrets-issuance + artifacto. Core infra. Mounts library.
104 gitea 1 1 GiB Git server. NFS would hurt git lock/stat perf. Mounts library (bare repos).
103 paperless 2 3 GiB Document archive. Moderate I/O, OCR writes. Mounts library.
102 nfs-export 1 512 MiB Exports library to zimaos via NFS. Must stay with the physical library.
100 zimaos 4 8 GiB (VM) NAS frontend eval. Already NFS-mounts library from 102.
108 haos 2 4 GiB (VM) Home Assistant OS. Hardware access, low latency.

Constraints & prerequisites

1. Network — strong needs an internal bridge (Phase 0)

strong currently has only vmbr0 on 192.168.178.0/24. Guests created there get household-LAN IPs, not homelab-subnet IPs. Two options:

  • Option A (recommended): Add vmbr1 on strong as a portless internal bridge with a 192.168.8.x/24 address (e.g. 192.168.8.3). Route between strong's vmbr0 and vmbr1 the same way hubris does. Guests go on vmbr1 and get 192.168.8.x IPs — transparent to Caddy, DNS, and inter-LXC refs. Requires adding a static route on Fritz (or relying on hubris's existing route — strong would need IP forwarding + a route to 192.168.8.0/24 via vmbr1).

  • Option B (simpler, messier): Put guests on 192.168.178.x directly. Caddy can still reach them (hubris routes to 192.168.178.0/24). But DNS records, inter-LXC references, and firewall rules all assume 192.168.8.x. More config churn per guest.

2. Storage — rootfs migration (no shared storage)

local-lvm is per-node (not shared). Moving an LXC requires either:

  • vzdump → restore on strong (clean, but needs temp disk space + downtime)
  • rsync the rootfs to a new LXC on strong (faster for large rootfs like 120's 100G)
  • pct migrate only works with shared storage — not applicable here

For VMs (100, 108): qm migrate also needs shared storage. Not moving VMs.

3. Library access — NFS from hubris to strong

Since the physical library SSD is still on hubris, strong's guests that need /mnt/library must NFS-mount it from hubris. Options:

  • Export from hubris host directly (simplest): add /mnt/library to /etc/exports on hubris with the same squash params as LXC 102 (rw,all_squash,anonuid=33,anongid=10000,no_subtree_check). Mount on strong at /mnt/library. Strong's guests bind-mount it just like hubris's guests do.

  • Use existing nfs-export LXC 102: strong NFS-mounts from 192.168.8.200 (LXC 102). This already has the right squash config. Less host-level change. This is the path of least resistance.

4. GPU — iGPU passthrough on strong

strong has a Ryzen 7 PRO 6850U with Radeon 680M iGPU. For jellyfin (VAAPI transcoding) and mule-images (photo processing), we need:

  • /dev/dri/renderD128 passed to the LXC (lxc.cgroup2.devices.allow + lxc.mount.entry or Proxmox's dev0: passthrough)
  • video / render group membership inside the container
  • Confirm amdgpu driver loads on strong's host kernel (it should — same APU family)

5. Quorum — 2-node cluster, no QDevice

Moving guests to strong does NOT fix the quorum issue but reduces blast radius: if hubris reboots (its known thermal instability), the guests on strong keep running independently. Consider adding a QDevice as a separate follow-up — it's orthogonal to this migration.


Revised migration phases

The original plan's NFS-over-LAN approach has been superseded. Instead, media library data moves to ludo-lvm on strong so migrated guests access it as a local ext4 mount. Data is split by origin:

hubris (stays):   library SSD (3.7T, 1.2T used)
                     └── /mnt/library/{documents,images,cloud,homecloud,notes,repos,sophia}
                         ↑ user-generated content (docs, photos, cloud sync, notes, repos, workshop)

strong (moves):   ludo-lvm (1.8T, 0 used at start)
                     └── /mnt/media_local ← 1.5T thin volume
                            └── {downloads,movies,music,tv,anime,books}
                            ↑ non-user-generated content (media arr stack, book library)
Category Stays on hubris Moves to strong
Media downloads (25G), movies (51G), music (29G), tv (30G), anime (206G)
Books books (2.6G)
Docs/Photos documents (249M), images (4K)
Cloud sync cloud (287G), homecloud (367G)
Personal notes (6.7M), repos (84M), sophia (151G)
Total ~805G ~344G

ludo-lvm (1.8T) fits all media + books with ~1.15T headroom for growth. hubris library SSD (3.7T, 1.2T used) retains the user-generated content. Both sides keep their data local — no cross-node NFS needed for daily I/O.


Phase 2a — Prepare ludo-lvm on strong

  1. Create a ext4 filesystem on ludo-lvm for media:
    lvcreate -n media -L 1.5T ludo-lvm
    mkfs.ext4 /dev/ludo-lvm/media
    
  2. Mount at /mnt/media_local on strong, add to /etc/fstab
  3. rsync media directories from hubris → strong:
    rsync -av --progress /mnt/library/{movies,tv,anime,downloads,music,books} strong:/mnt/media_local/
    

Phase 2b — Migrate arriman (122) to strong

  1. Stop arriman on hubris, dump rootfs (24G)
  2. Restore on strong with IP 192.168.8.245/28 on vmbr1
  3. Mount /mnt/media_local/mnt/library via mp0 (downloads land locally)
  4. Update Caddy: jellyseerr/qbit/sab backends → new IP
  5. Update inventory.yaml

Phase 2c — Migrate jellyfin (101) to strong

  1. Stop jellyfin on hubris, dump rootfs (16G)
  2. Restore on strong with IP 192.168.8.246/28 on vmbr1
  3. Pass /dev/dri/renderD128 + /dev/dri/card0 (Radeon 680M + RX 7600)
  4. Mount /mnt/media_local/mnt/library via mp0 (media reads locally)
  5. Update Caddy: media.hubris.network → new IP
  6. Reinstall sso-inject.js in web dir (lost on every apt upgrade)
  7. Test VAAPI transcoding, SSO login, media playback

Phase 2d — Migrate grimmory (130) to strong

  1. Stop grimmory on hubris, dump rootfs (16G)
  2. Restore on strong with IP 192.168.8.247/28 on vmbr1
  3. Mount /mnt/media_local/mnt/library via mp0 (books read locally)
  4. Update Caddy: books.hubris.network → new IP
  5. Update inventory.yaml
  6. Test: book browsing, calibre-web access

No NFS export needed

With the data split by origin, hubris guests that only need user-generated content (documents, images, cloud, repos, sophia) still access them from the original library SSD — no cross-node NFS required. The two sides are independent.

Result after Phase 2: hubris frees 26 GiB RAM (4 migrated guests) + 344G of library I/O burden. Strong becomes the media/books powerhouse.


Phase 3 — Migrate mule-images (120) to strong

Move photo management (12 GiB RAM, 6 cores, iGPU) last because it needs:

  • /mnt/library access (now NFS from strong — already set up in Phase 2d)
  • /dev/dri/renderD128 (Radeon 680M — confirm VAAPI compatibility first)

Steps:

  1. Stop mule-images on hubris, rsync the 100G rootfs to strong (faster than vzdump)
  2. Restore on strong with IP on vmbr1
  3. Pass Radeon 680M iGPU
  4. Reconfigure library paths → /mnt/media_local (or keep NFS mount)
  5. Update Caddy: photos.hubris.network → new IP
  6. Test photo import + processing pipeline

Phase 4 — Tier 3 moves (optional)

Migrate grimmory (130), teddycloud (131), rclone (132), trmnl (128), sophia (119) as needed — each frees 12 GiB. Not urgent; do when convenient.


Phase 5 — Follow-up

  • QDevice: add a tiebreaker for 2-node quorum
  • Gaming VM: strong's 6850U has enough cores alongside migrated LXCs
  • Hubris library cleanup: after all guests are confirmed working, decide whether to keep the original library SSD as backup or repurpose it

Resource math after Phase 3 (all Tier 1 + 2 moved)

hubris strong
Guests 11 LXC + 2 VM 5 LXC
RAM allocated ~25 GiB ~45 GiB
RAM capacity 28 GiB 28 GiB
Overcommit 0.9× (under-committed) 1.6× (manageable)
Library disk Local ext4 (3.7T) → NFS client Local ext4 on ludo-lvm (1.8T)
GPU Radeon 760M (idle) Radeon 680M (jellyfin + mule-images)

strong becomes the media/library powerhouse. hubris becomes a lean core-infra node (DNS, auth, git, docs, caddy, HA).



Risk register

Risk Impact Mitigation
NFS latency for library reads (jellyfin, arriman) Media playback stutter, slow downloads Test iperf between strong↔hubris first. If 2.5G link, NFS throughput is fine (~1 Gbit/s).
GPU passthrough on strong (680M vs 760M) Transcode quality/compat differences Both are AMD VAAPI — same driver stack. Test vainfo inside LXC before going live.
Caddy backend IP churn Service outage if IP wrong Update Caddyfile in git repo (caddy-conf), test each route before destroying old LXC.
vzdump/restore downtime Service unavailable during migration Schedule off-hours. Use rsync for large rootfs (120's 100G) to minimize freeze window.
2-node quorum still fragile If hubris goes down, strong /etc/pve goes read-only Guests keep running. Add QDevice as follow-up.
Library data integrity during NFS transition Permission drift NFS all_squash,anonuid=33,anongid=10000 matches existing LXC 102 config. Verify with ls -la /mnt/library after mount.

Open questions for operator

  1. Internal bridge on strong: proceed with vmbr1 on 192.168.8.3/24 (Option A), or use 192.168.178.x guest IPs (Option B)?
  2. Migration method: vzdump/restore (clean, downtime) vs rsync rootfs (faster for large disks, needs manual config copy)?
  3. Phase 1 priority: move elementsynapse + house first (quick wins), or go straight to Phase 2 (mule-images/jellyfin/arriman) for maximum relief?
  4. Should we add a QDevice now before moving anything, to protect management plane during the migration?

Changelog

2026-07-05 — Phase 2d complete (grimmory migrated; media NFS to zimaos)

grimmory (130) → 192.168.8.247 on strong. Rsync'd /books (2.6G) to ludo-lvm. LXC 102 (nfs-export) now mounts strong's NFS at /mnt/media and exports it as a second share alongside /mnt/library. Zimaos mounts both: /media/library (hubris user-generated) and /media/media (strong media+books). See hosts/strong.md changelog.

2026-07-05 — Phase 2 complete (arriman + jellyfin migrated; library on ludo-lvm)

arriman (122) → 192.168.8.245, jellyfin (101) → 192.168.8.246. Created 1.5T thin volume on ludo-lvm, rsync'd 363G of media data. Both containers use local ext4 mount — no NFS. Jellyfin has 680M + RX 7600 GPU passthrough. Caddy backends updated. See hosts/strong.md changelog.

2026-07-05 — Phase 1 complete (elementsynapse + house migrated to strong)

Both Tier 1 guests moved: elementsynapse (118) → 192.168.8.242, house (129) → 192.168.8.244. Strong now has vmbr1 at 192.168.8.241/28. Hubris has proxy ARP + /32 routes for strong guest range. DHCP scope narrowed to 192.168.8.100-239 to avoid conflicts. Teddycloud (LXC 131) given static IP 192.168.8.150 due to IP conflict with previous DHCP allocation at 192.168.8.243. See hosts/strong.md changelog for full steps.

2026-07-05 — assessment created

Built from live pct config + pvesm status + free -h data pulled from both nodes. Supersedes the storage-migration framing of the original library-SSD plan — this assessment treats the SSD move as optional and focuses on guest relocation via NFS.