strong migration Phase 1+2: move 5 LXCs + library split to ludo-lvm
This commit is contained in:
358
.hermes/plans/2026-07-05_strong-migration-assessment.md
Normal file
358
.hermes/plans/2026-07-05_strong-migration-assessment.md
Normal file
@@ -0,0 +1,358 @@
|
||||
# Assessment: Which nodes can move to `strong`
|
||||
|
||||
## Executive summary
|
||||
|
||||
hubris is **memory-starved**: 28 GiB RAM, 71.9 GiB allocated across 18 LXC + 2 VM
|
||||
(2.5× overcommit), 10 GiB swap in active use. strong sits **completely empty** —
|
||||
28 GiB RAM, 25 GiB free, 0 guests, 2.7 TiB unused storage. The single most
|
||||
effective decongestion move is to shift guests off hubris onto strong.
|
||||
|
||||
This document assesses every guest for move-readiness, grouped by constraints
|
||||
(library dependency, GPU, core-infra status), and proposes a phased migration
|
||||
that does **not** require the physical library-SSD move (the blocker of the
|
||||
original plan) — library access from strong is provided via NFS from hubris.
|
||||
|
||||
---
|
||||
|
||||
## Current resource state (live, 2026-07-05)
|
||||
|
||||
### hubris — overloaded
|
||||
|
||||
| Resource | Capacity | Allocated (all guests) | Actual use | Status |
|
||||
|----------|----------|----------------------|------------|--------|
|
||||
| RAM | 28 GiB | 70.2 GiB (2.5× overcommit) | 18 GiB used + 10 GiB swap | ⚠️ heavy swap pressure |
|
||||
| CPU | 12 vCPU (6c/12t) | 43 vCPU (3.6× overcommit) | ~43% scaling MHz | OK (shares) |
|
||||
| local-lvm | 856 GiB | 582 GiB allocated (29% thin) | — | OK |
|
||||
| library (lvmthin) | 3.7 TiB | — | 1.2 TiB used (34%) | OK, 2.3 TiB free |
|
||||
|
||||
### strong — empty, ready
|
||||
|
||||
| Resource | Capacity | Used | Status |
|
||||
|----------|----------|------|--------|
|
||||
| RAM | 28 GiB | 2.3 GiB (host only) | 25 GiB free |
|
||||
| CPU | 16 vCPU (8c/16t) | idle (0.08 load) | 100% free |
|
||||
| local-lvm | 856 GiB | 0 | empty |
|
||||
| ludo-lvm | 1.8 TiB | 0 | empty |
|
||||
| Guests | — | 0 LXC, 0 VM | nothing running |
|
||||
|
||||
### Network topology constraint
|
||||
|
||||
```
|
||||
Fritz!Box (192.168.178.1)
|
||||
└── SODOLA 2.5G switch
|
||||
├── hubris eno1 → vmbr1 (192.168.178.10) → vmbr0 (192.168.8.0/24)
|
||||
│ └── all 20 guests on 192.168.8.x
|
||||
└── strong vmbr0 (192.168.178.181)
|
||||
└── no internal bridge yet, guests would be on 192.168.178.x
|
||||
```
|
||||
|
||||
strong reaches `192.168.8.0/24` via the Fritz static route through hubris.
|
||||
Guests on strong get `192.168.178.x` IPs unless we add an internal bridge
|
||||
on strong (Phase 0 prerequisite — see below).
|
||||
|
||||
---
|
||||
|
||||
## Per-guest assessment
|
||||
|
||||
### Tier 1 — Move immediately (no library dependency, no core-infra)
|
||||
|
||||
These guests mount **no** `/mnt/library` and are not part of the core
|
||||
infrastructure spine (caddy/dns/auth/mcp). They are the easiest wins.
|
||||
|
||||
| ID | Name | Cores | RAM | Library? | GPU? | Notes |
|
||||
|----|------|-------|-----|----------|------|-------|
|
||||
| 118 | elementsynapse | 2 | 4 GiB | ❌ | ❌ | Matrix homeserver. Public via Caddy (`matrix.hubris.network`). Only change: Caddy backend IP. **Easiest move in the fleet.** |
|
||||
| 129 | house | 2 | 3 GiB | ❌ | ❌ | Yuvomi family planner (Docker). No public Caddy route yet (uses VPS traefik directly). Self-contained. |
|
||||
|
||||
**Combined RAM freed from hubris: 7 GiB.** No NFS, no library, no GPU.
|
||||
|
||||
### Tier 2 — Move with library NFS (high resource consumers)
|
||||
|
||||
These are the heaviest guests and the original migration plan's primary
|
||||
targets. They mount `/mnt/library` and two use the iGPU. Moving them
|
||||
requires an NFS export from hubris → strong (reverse of the original
|
||||
plan's direction, since the physical SSD hasn't moved).
|
||||
|
||||
| ID | Name | Cores | RAM | Library? | GPU? | I/O profile | Notes |
|
||||
|----|------|-------|-----|----------|------|-------------|-------|
|
||||
| 120 | mule-images | 6 | 12 GiB | ✅ mp0 | ✅ iGPU | Write-heavy (photo processing) | **#1 RAM consumer.** Strong has Radeon 680M iGPU (VAAPI works). |
|
||||
| 122 | arriman | 4 | 8 GiB | ✅ mp0 | ❌ | Write-heavy (downloads) | *arr stack + qbit + sab. Mounts library for download writes. |
|
||||
| 101 | jellyfin | 4 | 8 GiB | ✅ mp0 | ✅ iGPU | Read-heavy sequential | Media streaming + transcode. Strong 680M handles VAAPI. |
|
||||
|
||||
**Combined RAM freed: 28 GiB.** This alone would eliminate hubris's swap
|
||||
pressure entirely.
|
||||
|
||||
### Tier 3 — Could move, low urgency
|
||||
|
||||
| ID | Name | Cores | RAM | Library? | Notes |
|
||||
|----|------|-------|-----|----------|-------|
|
||||
| 130 | grimmory | 1 | 2 GiB | ✅ mp0 | Book library (Docker). Migrated from apps LXC recently. |
|
||||
| 131 | teddycloud | 1 | 1 GiB | ✅ mp0 | New (not in inventory.yaml yet). |
|
||||
| 132 | rclone | 1 | 2 GiB | ✅ mp0 (ro) | Backup container. Read-only library mount. |
|
||||
| 128 | trmnl | 1 | 768 MiB | ❌ | TRMNL middleware. No library. Could move but tiny. |
|
||||
| 119 | sophia | 2 | 1 GiB | ✅ mp0 | Workshop. Light use. |
|
||||
|
||||
### Stay on hubris (core infrastructure)
|
||||
|
||||
| ID | Name | Cores | RAM | Why it stays |
|
||||
|----|------|-------|-----|--------------|
|
||||
| 121 | caddy | 1 | 512 MiB | Reverse proxy — terminates all `*.hubris.network`. Must stay on hubris for LAN-side reachability. **Needs backend IP updates** when guests move. |
|
||||
| 107 | dns | 1 | 1 GiB | Technitium DNS, split-horizon. Core. |
|
||||
| 106 | auth-outpost | 1 | 512 MiB | Authentik SSO enforcement. Core. |
|
||||
| 105 | apps | 2 | 4 GiB | homelab MCP + secrets-issuance + artifacto. Core infra. Mounts library. |
|
||||
| 104 | gitea | 1 | 1 GiB | Git server. NFS would hurt git lock/stat perf. Mounts library (bare repos). |
|
||||
| 103 | paperless | 2 | 3 GiB | Document archive. Moderate I/O, OCR writes. Mounts library. |
|
||||
| 102 | nfs-export | 1 | 512 MiB | Exports library to zimaos via NFS. Must stay with the physical library. |
|
||||
| 100 | zimaos | 4 | 8 GiB (VM) | NAS frontend eval. Already NFS-mounts library from 102. |
|
||||
| 108 | haos | 2 | 4 GiB (VM) | Home Assistant OS. Hardware access, low latency. |
|
||||
|
||||
---
|
||||
|
||||
## Constraints & prerequisites
|
||||
|
||||
### 1. Network — strong needs an internal bridge (Phase 0)
|
||||
|
||||
strong currently has only `vmbr0` on `192.168.178.0/24`. Guests created there
|
||||
get household-LAN IPs, not homelab-subnet IPs. Two options:
|
||||
|
||||
- **Option A (recommended):** Add `vmbr1` on strong as a portless internal
|
||||
bridge with a `192.168.8.x/24` address (e.g. `192.168.8.3`). Route between
|
||||
strong's `vmbr0` and `vmbr1` the same way hubris does. Guests go on `vmbr1`
|
||||
and get `192.168.8.x` IPs — transparent to Caddy, DNS, and inter-LXC refs.
|
||||
Requires adding a static route on Fritz (or relying on hubris's existing
|
||||
route — strong would need IP forwarding + a route to 192.168.8.0/24 via vmbr1).
|
||||
|
||||
- **Option B (simpler, messier):** Put guests on `192.168.178.x` directly.
|
||||
Caddy can still reach them (hubris routes to 192.168.178.0/24). But DNS
|
||||
records, inter-LXC references, and firewall rules all assume `192.168.8.x`.
|
||||
More config churn per guest.
|
||||
|
||||
### 2. Storage — rootfs migration (no shared storage)
|
||||
|
||||
`local-lvm` is per-node (not shared). Moving an LXC requires either:
|
||||
- `vzdump` → restore on strong (clean, but needs temp disk space + downtime)
|
||||
- `rsync` the rootfs to a new LXC on strong (faster for large rootfs like 120's 100G)
|
||||
- `pct migrate` only works with shared storage — **not applicable here**
|
||||
|
||||
For VMs (100, 108): `qm migrate` also needs shared storage. Not moving VMs.
|
||||
|
||||
### 3. Library access — NFS from hubris to strong
|
||||
|
||||
Since the physical library SSD is still on hubris, strong's guests that need
|
||||
`/mnt/library` must NFS-mount it from hubris. Options:
|
||||
|
||||
- **Export from hubris host directly** (simplest): add `/mnt/library` to
|
||||
`/etc/exports` on hubris with the same squash params as LXC 102
|
||||
(`rw,all_squash,anonuid=33,anongid=10000,no_subtree_check`). Mount on strong
|
||||
at `/mnt/library`. Strong's guests bind-mount it just like hubris's guests do.
|
||||
|
||||
- **Use existing nfs-export LXC 102**: strong NFS-mounts from `192.168.8.200`
|
||||
(LXC 102). This already has the right squash config. Less host-level change.
|
||||
**This is the path of least resistance.**
|
||||
|
||||
### 4. GPU — iGPU passthrough on strong
|
||||
|
||||
strong has a Ryzen 7 PRO 6850U with Radeon 680M iGPU. For jellyfin (VAAPI
|
||||
transcoding) and mule-images (photo processing), we need:
|
||||
- `/dev/dri/renderD128` passed to the LXC (`lxc.cgroup2.devices.allow` +
|
||||
`lxc.mount.entry` or Proxmox's `dev0:` passthrough)
|
||||
- `video` / `render` group membership inside the container
|
||||
- Confirm `amdgpu` driver loads on strong's host kernel (it should — same APU family)
|
||||
|
||||
### 5. Quorum — 2-node cluster, no QDevice
|
||||
|
||||
Moving guests to strong does NOT fix the quorum issue but **reduces blast
|
||||
radius**: if hubris reboots (its known thermal instability), the guests on
|
||||
strong keep running independently. Consider adding a QDevice as a separate
|
||||
follow-up — it's orthogonal to this migration.
|
||||
|
||||
---
|
||||
|
||||
## Revised migration phases
|
||||
|
||||
The original plan's NFS-over-LAN approach has been superseded. Instead,
|
||||
**media library data moves to ludo-lvm** on strong so migrated guests access
|
||||
it as a local ext4 mount. Data is split by origin:
|
||||
|
||||
```
|
||||
hubris (stays): library SSD (3.7T, 1.2T used)
|
||||
└── /mnt/library/{documents,images,cloud,homecloud,notes,repos,sophia}
|
||||
↑ user-generated content (docs, photos, cloud sync, notes, repos, workshop)
|
||||
|
||||
strong (moves): ludo-lvm (1.8T, 0 used at start)
|
||||
└── /mnt/media_local ← 1.5T thin volume
|
||||
└── {downloads,movies,music,tv,anime,books}
|
||||
↑ non-user-generated content (media arr stack, book library)
|
||||
```
|
||||
|
||||
| Category | Stays on hubris | Moves to strong |
|
||||
|----------|----------------|-----------------|
|
||||
| Media | — | downloads (25G), movies (51G), music (29G), tv (30G), anime (206G) |
|
||||
| Books | — | books (2.6G) |
|
||||
| Docs/Photos | documents (249M), images (4K) | — |
|
||||
| Cloud sync | cloud (287G), homecloud (367G) | — |
|
||||
| Personal | notes (6.7M), repos (84M), sophia (151G) | — |
|
||||
| **Total** | **~805G** | **~344G** |
|
||||
|
||||
ludo-lvm (1.8T) fits all media + books with ~1.15T headroom for growth.
|
||||
hubris library SSD (3.7T, 1.2T used) retains the user-generated content.
|
||||
Both sides keep their data local — no cross-node NFS needed for daily I/O.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2a — Prepare ludo-lvm on strong
|
||||
|
||||
1. Create a ext4 filesystem on ludo-lvm for media:
|
||||
```bash
|
||||
lvcreate -n media -L 1.5T ludo-lvm
|
||||
mkfs.ext4 /dev/ludo-lvm/media
|
||||
```
|
||||
2. Mount at `/mnt/media_local` on strong, add to `/etc/fstab`
|
||||
3. rsync media directories from hubris → strong:
|
||||
```bash
|
||||
rsync -av --progress /mnt/library/{movies,tv,anime,downloads,music,books} strong:/mnt/media_local/
|
||||
```
|
||||
|
||||
### Phase 2b — Migrate arriman (122) to strong
|
||||
|
||||
1. Stop arriman on hubris, dump rootfs (24G)
|
||||
2. Restore on strong with IP `192.168.8.245/28` on vmbr1
|
||||
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (downloads land locally)
|
||||
4. Update Caddy: jellyseerr/qbit/sab backends → new IP
|
||||
5. Update inventory.yaml
|
||||
|
||||
### Phase 2c — Migrate jellyfin (101) to strong
|
||||
|
||||
1. Stop jellyfin on hubris, dump rootfs (16G)
|
||||
2. Restore on strong with IP `192.168.8.246/28` on vmbr1
|
||||
3. Pass `/dev/dri/renderD128` + `/dev/dri/card0` (Radeon 680M + RX 7600)
|
||||
4. Mount `/mnt/media_local` → `/mnt/library` via mp0 (media reads locally)
|
||||
5. Update Caddy: `media.hubris.network` → new IP
|
||||
6. Reinstall `sso-inject.js` in web dir (lost on every apt upgrade)
|
||||
7. Test VAAPI transcoding, SSO login, media playback
|
||||
|
||||
### Phase 2d — Migrate grimmory (130) to strong
|
||||
|
||||
1. Stop grimmory on hubris, dump rootfs (16G)
|
||||
2. Restore on strong with IP `192.168.8.247/28` on vmbr1
|
||||
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (books read locally)
|
||||
4. Update Caddy: `books.hubris.network` → new IP
|
||||
5. Update inventory.yaml
|
||||
6. Test: book browsing, calibre-web access
|
||||
|
||||
### No NFS export needed
|
||||
|
||||
With the data split by origin, hubris guests that only need user-generated
|
||||
content (documents, images, cloud, repos, sophia) still access them from the
|
||||
original library SSD — no cross-node NFS required. The two sides are
|
||||
independent.
|
||||
|
||||
**Result after Phase 2: hubris frees 26 GiB RAM (4 migrated guests) + 344G of
|
||||
library I/O burden. Strong becomes the media/books powerhouse.**
|
||||
|
||||
---
|
||||
|
||||
### Phase 3 — Migrate mule-images (120) to strong
|
||||
|
||||
Move photo management (12 GiB RAM, 6 cores, iGPU) last because it needs:
|
||||
- `/mnt/library` access (now NFS from strong — already set up in Phase 2d)
|
||||
- `/dev/dri/renderD128` (Radeon 680M — confirm VAAPI compatibility first)
|
||||
|
||||
Steps:
|
||||
1. Stop mule-images on hubris, rsync the 100G rootfs to strong (faster than vzdump)
|
||||
2. Restore on strong with IP on vmbr1
|
||||
3. Pass Radeon 680M iGPU
|
||||
4. Reconfigure library paths → `/mnt/media_local` (or keep NFS mount)
|
||||
5. Update Caddy: `photos.hubris.network` → new IP
|
||||
6. Test photo import + processing pipeline
|
||||
|
||||
---
|
||||
|
||||
### Phase 4 — Tier 3 moves (optional)
|
||||
|
||||
Migrate grimmory (130), teddycloud (131), rclone (132), trmnl (128), sophia (119)
|
||||
as needed — each frees 1–2 GiB. Not urgent; do when convenient.
|
||||
|
||||
---
|
||||
|
||||
### Phase 5 — Follow-up
|
||||
|
||||
- **QDevice**: add a tiebreaker for 2-node quorum
|
||||
- **Gaming VM**: strong's 6850U has enough cores alongside migrated LXCs
|
||||
- **Hubris library cleanup**: after all guests are confirmed working, decide
|
||||
whether to keep the original library SSD as backup or repurpose it
|
||||
|
||||
---
|
||||
|
||||
## Resource math after Phase 3 (all Tier 1 + 2 moved)
|
||||
|
||||
| | hubris | strong |
|
||||
|---|--------|--------|
|
||||
| Guests | 11 LXC + 2 VM | 5 LXC |
|
||||
| RAM allocated | ~25 GiB | ~45 GiB |
|
||||
| RAM capacity | 28 GiB | 28 GiB |
|
||||
| Overcommit | 0.9× (under-committed) | 1.6× (manageable) |
|
||||
| Library disk | Local ext4 (3.7T) → NFS client | Local ext4 on ludo-lvm (1.8T) |
|
||||
| GPU | Radeon 760M (idle) | Radeon 680M (jellyfin + mule-images) |
|
||||
|
||||
strong becomes the media/library powerhouse. hubris becomes a lean core-infra
|
||||
node (DNS, auth, git, docs, caddy, HA).
|
||||
|
||||
---
|
||||
|
||||
---
|
||||
|
||||
## Risk register
|
||||
|
||||
| Risk | Impact | Mitigation |
|
||||
|------|--------|------------|
|
||||
| NFS latency for library reads (jellyfin, arriman) | Media playback stutter, slow downloads | Test iperf between strong↔hubris first. If 2.5G link, NFS throughput is fine (~1 Gbit/s). |
|
||||
| GPU passthrough on strong (680M vs 760M) | Transcode quality/compat differences | Both are AMD VAAPI — same driver stack. Test `vainfo` inside LXC before going live. |
|
||||
| Caddy backend IP churn | Service outage if IP wrong | Update Caddyfile in git repo (caddy-conf), test each route before destroying old LXC. |
|
||||
| vzdump/restore downtime | Service unavailable during migration | Schedule off-hours. Use rsync for large rootfs (120's 100G) to minimize freeze window. |
|
||||
| 2-node quorum still fragile | If hubris goes down, strong /etc/pve goes read-only | Guests keep running. Add QDevice as follow-up. |
|
||||
| Library data integrity during NFS transition | Permission drift | NFS `all_squash,anonuid=33,anongid=10000` matches existing LXC 102 config. Verify with `ls -la /mnt/library` after mount. |
|
||||
|
||||
---
|
||||
|
||||
## Open questions for operator
|
||||
|
||||
1. **Internal bridge on strong**: proceed with `vmbr1` on `192.168.8.3/24`
|
||||
(Option A), or use `192.168.178.x` guest IPs (Option B)?
|
||||
2. **Migration method**: `vzdump`/restore (clean, downtime) vs `rsync` rootfs
|
||||
(faster for large disks, needs manual config copy)?
|
||||
3. **Phase 1 priority**: move elementsynapse + house first (quick wins), or
|
||||
go straight to Phase 2 (mule-images/jellyfin/arriman) for maximum relief?
|
||||
4. **Should we add a QDevice now** before moving anything, to protect
|
||||
management plane during the migration?
|
||||
|
||||
---
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2026-07-05 — Phase 2d complete (grimmory migrated; media NFS to zimaos)
|
||||
grimmory (130) → 192.168.8.247 on strong. Rsync'd /books (2.6G) to ludo-lvm.
|
||||
LXC 102 (nfs-export) now mounts strong's NFS at /mnt/media and exports it as
|
||||
a second share alongside /mnt/library. Zimaos mounts both: /media/library
|
||||
(hubris user-generated) and /media/media (strong media+books).
|
||||
See hosts/strong.md changelog.
|
||||
|
||||
### 2026-07-05 — Phase 2 complete (arriman + jellyfin migrated; library on ludo-lvm)
|
||||
arriman (122) → 192.168.8.245, jellyfin (101) → 192.168.8.246. Created 1.5T
|
||||
thin volume on ludo-lvm, rsync'd 363G of media data. Both containers use local
|
||||
ext4 mount — no NFS. Jellyfin has 680M + RX 7600 GPU passthrough.
|
||||
Caddy backends updated. See hosts/strong.md changelog.
|
||||
|
||||
### 2026-07-05 — Phase 1 complete (elementsynapse + house migrated to strong)
|
||||
Both Tier 1 guests moved: elementsynapse (118) → 192.168.8.242, house (129) → 192.168.8.244.
|
||||
Strong now has vmbr1 at 192.168.8.241/28. Hubris has proxy ARP + /32 routes for strong
|
||||
guest range. DHCP scope narrowed to 192.168.8.100-239 to avoid conflicts.
|
||||
Teddycloud (LXC 131) given static IP 192.168.8.150 due to IP conflict with
|
||||
previous DHCP allocation at 192.168.8.243.
|
||||
See hosts/strong.md changelog for full steps.
|
||||
|
||||
### 2026-07-05 — assessment created
|
||||
Built from live `pct config` + `pvesm status` + `free -h` data pulled from
|
||||
both nodes. Supersedes the storage-migration framing of the original
|
||||
library-SSD plan — this assessment treats the SSD move as optional and
|
||||
focuses on guest relocation via NFS.
|
||||
Reference in New Issue
Block a user