16 KiB
Assessment: Which nodes can move to strong
Executive summary
hubris is memory-starved: 28 GiB RAM, 71.9 GiB allocated across 18 LXC + 2 VM (2.5× overcommit), 10 GiB swap in active use. strong sits completely empty — 28 GiB RAM, 25 GiB free, 0 guests, 2.7 TiB unused storage. The single most effective decongestion move is to shift guests off hubris onto strong.
This document assesses every guest for move-readiness, grouped by constraints (library dependency, GPU, core-infra status), and proposes a phased migration that does not require the physical library-SSD move (the blocker of the original plan) — library access from strong is provided via NFS from hubris.
Current resource state (live, 2026-07-05)
hubris — overloaded
| Resource | Capacity | Allocated (all guests) | Actual use | Status |
|---|---|---|---|---|
| RAM | 28 GiB | 70.2 GiB (2.5× overcommit) | 18 GiB used + 10 GiB swap | ⚠️ heavy swap pressure |
| CPU | 12 vCPU (6c/12t) | 43 vCPU (3.6× overcommit) | ~43% scaling MHz | OK (shares) |
| local-lvm | 856 GiB | 582 GiB allocated (29% thin) | — | OK |
| library (lvmthin) | 3.7 TiB | — | 1.2 TiB used (34%) | OK, 2.3 TiB free |
strong — empty, ready
| Resource | Capacity | Used | Status |
|---|---|---|---|
| RAM | 28 GiB | 2.3 GiB (host only) | 25 GiB free |
| CPU | 16 vCPU (8c/16t) | idle (0.08 load) | 100% free |
| local-lvm | 856 GiB | 0 | empty |
| ludo-lvm | 1.8 TiB | 0 | empty |
| Guests | — | 0 LXC, 0 VM | nothing running |
Network topology constraint
Fritz!Box (192.168.178.1)
└── SODOLA 2.5G switch
├── hubris eno1 → vmbr1 (192.168.178.10) → vmbr0 (192.168.8.0/24)
│ └── all 20 guests on 192.168.8.x
└── strong vmbr0 (192.168.178.181)
└── no internal bridge yet, guests would be on 192.168.178.x
strong reaches 192.168.8.0/24 via the Fritz static route through hubris.
Guests on strong get 192.168.178.x IPs unless we add an internal bridge
on strong (Phase 0 prerequisite — see below).
Per-guest assessment
Tier 1 — Move immediately (no library dependency, no core-infra)
These guests mount no /mnt/library and are not part of the core
infrastructure spine (caddy/dns/auth/mcp). They are the easiest wins.
| ID | Name | Cores | RAM | Library? | GPU? | Notes |
|---|---|---|---|---|---|---|
| 118 | elementsynapse | 2 | 4 GiB | ❌ | ❌ | Matrix homeserver. Public via Caddy (matrix.hubris.network). Only change: Caddy backend IP. Easiest move in the fleet. |
| 129 | house | 2 | 3 GiB | ❌ | ❌ | Yuvomi family planner (Docker). No public Caddy route yet (uses VPS traefik directly). Self-contained. |
Combined RAM freed from hubris: 7 GiB. No NFS, no library, no GPU.
Tier 2 — Move with library NFS (high resource consumers)
These are the heaviest guests and the original migration plan's primary
targets. They mount /mnt/library and two use the iGPU. Moving them
requires an NFS export from hubris → strong (reverse of the original
plan's direction, since the physical SSD hasn't moved).
| ID | Name | Cores | RAM | Library? | GPU? | I/O profile | Notes |
|---|---|---|---|---|---|---|---|
| 120 | mule-images | 6 | 12 GiB | ✅ mp0 | ✅ iGPU | Write-heavy (photo processing) | #1 RAM consumer. Strong has Radeon 680M iGPU (VAAPI works). |
| 122 | arriman | 4 | 8 GiB | ✅ mp0 | ❌ | Write-heavy (downloads) | *arr stack + qbit + sab. Mounts library for download writes. |
| 101 | jellyfin | 4 | 8 GiB | ✅ mp0 | ✅ iGPU | Read-heavy sequential | Media streaming + transcode. Strong 680M handles VAAPI. |
Combined RAM freed: 28 GiB. This alone would eliminate hubris's swap pressure entirely.
Tier 3 — Could move, low urgency
| ID | Name | Cores | RAM | Library? | Notes |
|---|---|---|---|---|---|
| 130 | grimmory | 1 | 2 GiB | ✅ mp0 | Book library (Docker). Migrated from apps LXC recently. |
| 131 | teddycloud | 1 | 1 GiB | ✅ mp0 | New (not in inventory.yaml yet). |
| 132 | rclone | 1 | 2 GiB | ✅ mp0 (ro) | Backup container. Read-only library mount. |
| 128 | trmnl | 1 | 768 MiB | ❌ | TRMNL middleware. No library. Could move but tiny. |
| 119 | sophia | 2 | 1 GiB | ✅ mp0 | Workshop. Light use. |
Stay on hubris (core infrastructure)
| ID | Name | Cores | RAM | Why it stays |
|---|---|---|---|---|
| 121 | caddy | 1 | 512 MiB | Reverse proxy — terminates all *.hubris.network. Must stay on hubris for LAN-side reachability. Needs backend IP updates when guests move. |
| 107 | dns | 1 | 1 GiB | Technitium DNS, split-horizon. Core. |
| 106 | auth-outpost | 1 | 512 MiB | Authentik SSO enforcement. Core. |
| 105 | apps | 2 | 4 GiB | homelab MCP + secrets-issuance + artifacto. Core infra. Mounts library. |
| 104 | gitea | 1 | 1 GiB | Git server. NFS would hurt git lock/stat perf. Mounts library (bare repos). |
| 103 | paperless | 2 | 3 GiB | Document archive. Moderate I/O, OCR writes. Mounts library. |
| 102 | nfs-export | 1 | 512 MiB | Exports library to zimaos via NFS. Must stay with the physical library. |
| 100 | zimaos | 4 | 8 GiB (VM) | NAS frontend eval. Already NFS-mounts library from 102. |
| 108 | haos | 2 | 4 GiB (VM) | Home Assistant OS. Hardware access, low latency. |
Constraints & prerequisites
1. Network — strong needs an internal bridge (Phase 0)
strong currently has only vmbr0 on 192.168.178.0/24. Guests created there
get household-LAN IPs, not homelab-subnet IPs. Two options:
-
Option A (recommended): Add
vmbr1on strong as a portless internal bridge with a192.168.8.x/24address (e.g.192.168.8.3). Route between strong'svmbr0andvmbr1the same way hubris does. Guests go onvmbr1and get192.168.8.xIPs — transparent to Caddy, DNS, and inter-LXC refs. Requires adding a static route on Fritz (or relying on hubris's existing route — strong would need IP forwarding + a route to 192.168.8.0/24 via vmbr1). -
Option B (simpler, messier): Put guests on
192.168.178.xdirectly. Caddy can still reach them (hubris routes to 192.168.178.0/24). But DNS records, inter-LXC references, and firewall rules all assume192.168.8.x. More config churn per guest.
2. Storage — rootfs migration (no shared storage)
local-lvm is per-node (not shared). Moving an LXC requires either:
vzdump→ restore on strong (clean, but needs temp disk space + downtime)rsyncthe rootfs to a new LXC on strong (faster for large rootfs like 120's 100G)pct migrateonly works with shared storage — not applicable here
For VMs (100, 108): qm migrate also needs shared storage. Not moving VMs.
3. Library access — NFS from hubris to strong
Since the physical library SSD is still on hubris, strong's guests that need
/mnt/library must NFS-mount it from hubris. Options:
-
Export from hubris host directly (simplest): add
/mnt/libraryto/etc/exportson hubris with the same squash params as LXC 102 (rw,all_squash,anonuid=33,anongid=10000,no_subtree_check). Mount on strong at/mnt/library. Strong's guests bind-mount it just like hubris's guests do. -
Use existing nfs-export LXC 102: strong NFS-mounts from
192.168.8.200(LXC 102). This already has the right squash config. Less host-level change. This is the path of least resistance.
4. GPU — iGPU passthrough on strong
strong has a Ryzen 7 PRO 6850U with Radeon 680M iGPU. For jellyfin (VAAPI transcoding) and mule-images (photo processing), we need:
/dev/dri/renderD128passed to the LXC (lxc.cgroup2.devices.allow+lxc.mount.entryor Proxmox'sdev0:passthrough)video/rendergroup membership inside the container- Confirm
amdgpudriver loads on strong's host kernel (it should — same APU family)
5. Quorum — 2-node cluster, no QDevice
Moving guests to strong does NOT fix the quorum issue but reduces blast radius: if hubris reboots (its known thermal instability), the guests on strong keep running independently. Consider adding a QDevice as a separate follow-up — it's orthogonal to this migration.
Revised migration phases
The original plan's NFS-over-LAN approach has been superseded. Instead, media library data moves to ludo-lvm on strong so migrated guests access it as a local ext4 mount. Data is split by origin:
hubris (stays): library SSD (3.7T, 1.2T used)
└── /mnt/library/{documents,images,cloud,homecloud,notes,repos,sophia}
↑ user-generated content (docs, photos, cloud sync, notes, repos, workshop)
strong (moves): ludo-lvm (1.8T, 0 used at start)
└── /mnt/media_local ← 1.5T thin volume
└── {downloads,movies,music,tv,anime,books}
↑ non-user-generated content (media arr stack, book library)
| Category | Stays on hubris | Moves to strong |
|---|---|---|
| Media | — | downloads (25G), movies (51G), music (29G), tv (30G), anime (206G) |
| Books | — | books (2.6G) |
| Docs/Photos | documents (249M), images (4K) | — |
| Cloud sync | cloud (287G), homecloud (367G) | — |
| Personal | notes (6.7M), repos (84M), sophia (151G) | — |
| Total | ~805G | ~344G |
ludo-lvm (1.8T) fits all media + books with ~1.15T headroom for growth. hubris library SSD (3.7T, 1.2T used) retains the user-generated content. Both sides keep their data local — no cross-node NFS needed for daily I/O.
Phase 2a — Prepare ludo-lvm on strong
- Create a ext4 filesystem on ludo-lvm for media:
lvcreate -n media -L 1.5T ludo-lvm mkfs.ext4 /dev/ludo-lvm/media - Mount at
/mnt/media_localon strong, add to/etc/fstab - rsync media directories from hubris → strong:
rsync -av --progress /mnt/library/{movies,tv,anime,downloads,music,books} strong:/mnt/media_local/
Phase 2b — Migrate arriman (122) to strong
- Stop arriman on hubris, dump rootfs (24G)
- Restore on strong with IP
192.168.8.245/28on vmbr1 - Mount
/mnt/media_local→/mnt/libraryvia mp0 (downloads land locally) - Update Caddy: jellyseerr/qbit/sab backends → new IP
- Update inventory.yaml
Phase 2c — Migrate jellyfin (101) to strong
- Stop jellyfin on hubris, dump rootfs (16G)
- Restore on strong with IP
192.168.8.246/28on vmbr1 - Pass
/dev/dri/renderD128+/dev/dri/card0(Radeon 680M + RX 7600) - Mount
/mnt/media_local→/mnt/libraryvia mp0 (media reads locally) - Update Caddy:
media.hubris.network→ new IP - Reinstall
sso-inject.jsin web dir (lost on every apt upgrade) - Test VAAPI transcoding, SSO login, media playback
Phase 2d — Migrate grimmory (130) to strong
- Stop grimmory on hubris, dump rootfs (16G)
- Restore on strong with IP
192.168.8.247/28on vmbr1 - Mount
/mnt/media_local→/mnt/libraryvia mp0 (books read locally) - Update Caddy:
books.hubris.network→ new IP - Update inventory.yaml
- Test: book browsing, calibre-web access
No NFS export needed
With the data split by origin, hubris guests that only need user-generated content (documents, images, cloud, repos, sophia) still access them from the original library SSD — no cross-node NFS required. The two sides are independent.
Result after Phase 2: hubris frees 26 GiB RAM (4 migrated guests) + 344G of library I/O burden. Strong becomes the media/books powerhouse.
Phase 3 — Migrate mule-images (120) to strong
Move photo management (12 GiB RAM, 6 cores, iGPU) last because it needs:
/mnt/libraryaccess (now NFS from strong — already set up in Phase 2d)/dev/dri/renderD128(Radeon 680M — confirm VAAPI compatibility first)
Steps:
- Stop mule-images on hubris, rsync the 100G rootfs to strong (faster than vzdump)
- Restore on strong with IP on vmbr1
- Pass Radeon 680M iGPU
- Reconfigure library paths →
/mnt/media_local(or keep NFS mount) - Update Caddy:
photos.hubris.network→ new IP - Test photo import + processing pipeline
Phase 4 — Tier 3 moves (optional)
Migrate grimmory (130), teddycloud (131), rclone (132), trmnl (128), sophia (119) as needed — each frees 1–2 GiB. Not urgent; do when convenient.
Phase 5 — Follow-up
- QDevice: add a tiebreaker for 2-node quorum
- Gaming VM: strong's 6850U has enough cores alongside migrated LXCs
- Hubris library cleanup: after all guests are confirmed working, decide whether to keep the original library SSD as backup or repurpose it
Resource math after Phase 3 (all Tier 1 + 2 moved)
| hubris | strong | |
|---|---|---|
| Guests | 11 LXC + 2 VM | 5 LXC |
| RAM allocated | ~25 GiB | ~45 GiB |
| RAM capacity | 28 GiB | 28 GiB |
| Overcommit | 0.9× (under-committed) | 1.6× (manageable) |
| Library disk | Local ext4 (3.7T) → NFS client | Local ext4 on ludo-lvm (1.8T) |
| GPU | Radeon 760M (idle) | Radeon 680M (jellyfin + mule-images) |
strong becomes the media/library powerhouse. hubris becomes a lean core-infra node (DNS, auth, git, docs, caddy, HA).
Risk register
| Risk | Impact | Mitigation |
|---|---|---|
| NFS latency for library reads (jellyfin, arriman) | Media playback stutter, slow downloads | Test iperf between strong↔hubris first. If 2.5G link, NFS throughput is fine (~1 Gbit/s). |
| GPU passthrough on strong (680M vs 760M) | Transcode quality/compat differences | Both are AMD VAAPI — same driver stack. Test vainfo inside LXC before going live. |
| Caddy backend IP churn | Service outage if IP wrong | Update Caddyfile in git repo (caddy-conf), test each route before destroying old LXC. |
| vzdump/restore downtime | Service unavailable during migration | Schedule off-hours. Use rsync for large rootfs (120's 100G) to minimize freeze window. |
| 2-node quorum still fragile | If hubris goes down, strong /etc/pve goes read-only | Guests keep running. Add QDevice as follow-up. |
| Library data integrity during NFS transition | Permission drift | NFS all_squash,anonuid=33,anongid=10000 matches existing LXC 102 config. Verify with ls -la /mnt/library after mount. |
Open questions for operator
- Internal bridge on strong: proceed with
vmbr1on192.168.8.3/24(Option A), or use192.168.178.xguest IPs (Option B)? - Migration method:
vzdump/restore (clean, downtime) vsrsyncrootfs (faster for large disks, needs manual config copy)? - Phase 1 priority: move elementsynapse + house first (quick wins), or go straight to Phase 2 (mule-images/jellyfin/arriman) for maximum relief?
- Should we add a QDevice now before moving anything, to protect management plane during the migration?
Changelog
2026-07-05 — Phase 2d complete (grimmory migrated; media NFS to zimaos)
grimmory (130) → 192.168.8.247 on strong. Rsync'd /books (2.6G) to ludo-lvm. LXC 102 (nfs-export) now mounts strong's NFS at /mnt/media and exports it as a second share alongside /mnt/library. Zimaos mounts both: /media/library (hubris user-generated) and /media/media (strong media+books). See hosts/strong.md changelog.
2026-07-05 — Phase 2 complete (arriman + jellyfin migrated; library on ludo-lvm)
arriman (122) → 192.168.8.245, jellyfin (101) → 192.168.8.246. Created 1.5T thin volume on ludo-lvm, rsync'd 363G of media data. Both containers use local ext4 mount — no NFS. Jellyfin has 680M + RX 7600 GPU passthrough. Caddy backends updated. See hosts/strong.md changelog.
2026-07-05 — Phase 1 complete (elementsynapse + house migrated to strong)
Both Tier 1 guests moved: elementsynapse (118) → 192.168.8.242, house (129) → 192.168.8.244. Strong now has vmbr1 at 192.168.8.241/28. Hubris has proxy ARP + /32 routes for strong guest range. DHCP scope narrowed to 192.168.8.100-239 to avoid conflicts. Teddycloud (LXC 131) given static IP 192.168.8.150 due to IP conflict with previous DHCP allocation at 192.168.8.243. See hosts/strong.md changelog for full steps.
2026-07-05 — assessment created
Built from live pct config + pvesm status + free -h data pulled from
both nodes. Supersedes the storage-migration framing of the original
library-SSD plan — this assessment treats the SSD move as optional and
focuses on guest relocation via NFS.