- Migrations 010 (content_hash) + 011 (search tsvector column) - new: internal/knowledge/seed.go — knowledge seed ingest engine - new: internal/httpapi/knowledge.go — SearchKnowledge + GetEntityKnowledge - wire knowledge ingest into oikos seed pipeline - convert all 36 wiki docs + 6 investigations + 12 runbooks → seeds/knowledge.yaml - archive: knowledge/wiki/→archive/, oikos/cards/→archive/, .hermes/plans/→archive/ - delete: 9 superseded Python kernel files, ledger/, mcp/build_host_files.py - remove empty knowledge/ directory tree
359 lines
16 KiB
Markdown
359 lines
16 KiB
Markdown
# Assessment: Which nodes can move to `strong`
|
||
|
||
## Executive summary
|
||
|
||
hubris is **memory-starved**: 28 GiB RAM, 71.9 GiB allocated across 18 LXC + 2 VM
|
||
(2.5× overcommit), 10 GiB swap in active use. strong sits **completely empty** —
|
||
28 GiB RAM, 25 GiB free, 0 guests, 2.7 TiB unused storage. The single most
|
||
effective decongestion move is to shift guests off hubris onto strong.
|
||
|
||
This document assesses every guest for move-readiness, grouped by constraints
|
||
(library dependency, GPU, core-infra status), and proposes a phased migration
|
||
that does **not** require the physical library-SSD move (the blocker of the
|
||
original plan) — library access from strong is provided via NFS from hubris.
|
||
|
||
---
|
||
|
||
## Current resource state (live, 2026-07-05)
|
||
|
||
### hubris — overloaded
|
||
|
||
| Resource | Capacity | Allocated (all guests) | Actual use | Status |
|
||
|----------|----------|----------------------|------------|--------|
|
||
| RAM | 28 GiB | 70.2 GiB (2.5× overcommit) | 18 GiB used + 10 GiB swap | ⚠️ heavy swap pressure |
|
||
| CPU | 12 vCPU (6c/12t) | 43 vCPU (3.6× overcommit) | ~43% scaling MHz | OK (shares) |
|
||
| local-lvm | 856 GiB | 582 GiB allocated (29% thin) | — | OK |
|
||
| library (lvmthin) | 3.7 TiB | — | 1.2 TiB used (34%) | OK, 2.3 TiB free |
|
||
|
||
### strong — empty, ready
|
||
|
||
| Resource | Capacity | Used | Status |
|
||
|----------|----------|------|--------|
|
||
| RAM | 28 GiB | 2.3 GiB (host only) | 25 GiB free |
|
||
| CPU | 16 vCPU (8c/16t) | idle (0.08 load) | 100% free |
|
||
| local-lvm | 856 GiB | 0 | empty |
|
||
| ludo-lvm | 1.8 TiB | 0 | empty |
|
||
| Guests | — | 0 LXC, 0 VM | nothing running |
|
||
|
||
### Network topology constraint
|
||
|
||
```
|
||
Fritz!Box (192.168.178.1)
|
||
└── SODOLA 2.5G switch
|
||
├── hubris eno1 → vmbr1 (192.168.178.10) → vmbr0 (192.168.8.0/24)
|
||
│ └── all 20 guests on 192.168.8.x
|
||
└── strong vmbr0 (192.168.178.181)
|
||
└── no internal bridge yet, guests would be on 192.168.178.x
|
||
```
|
||
|
||
strong reaches `192.168.8.0/24` via the Fritz static route through hubris.
|
||
Guests on strong get `192.168.178.x` IPs unless we add an internal bridge
|
||
on strong (Phase 0 prerequisite — see below).
|
||
|
||
---
|
||
|
||
## Per-guest assessment
|
||
|
||
### Tier 1 — Move immediately (no library dependency, no core-infra)
|
||
|
||
These guests mount **no** `/mnt/library` and are not part of the core
|
||
infrastructure spine (caddy/dns/auth/mcp). They are the easiest wins.
|
||
|
||
| ID | Name | Cores | RAM | Library? | GPU? | Notes |
|
||
|----|------|-------|-----|----------|------|-------|
|
||
| 118 | elementsynapse | 2 | 4 GiB | ❌ | ❌ | Matrix homeserver. Public via Caddy (`matrix.hubris.network`). Only change: Caddy backend IP. **Easiest move in the fleet.** |
|
||
| 129 | house | 2 | 3 GiB | ❌ | ❌ | Yuvomi family planner (Docker). No public Caddy route yet (uses VPS traefik directly). Self-contained. |
|
||
|
||
**Combined RAM freed from hubris: 7 GiB.** No NFS, no library, no GPU.
|
||
|
||
### Tier 2 — Move with library NFS (high resource consumers)
|
||
|
||
These are the heaviest guests and the original migration plan's primary
|
||
targets. They mount `/mnt/library` and two use the iGPU. Moving them
|
||
requires an NFS export from hubris → strong (reverse of the original
|
||
plan's direction, since the physical SSD hasn't moved).
|
||
|
||
| ID | Name | Cores | RAM | Library? | GPU? | I/O profile | Notes |
|
||
|----|------|-------|-----|----------|------|-------------|-------|
|
||
| 120 | mule-images | 6 | 12 GiB | ✅ mp0 | ✅ iGPU | Write-heavy (photo processing) | **#1 RAM consumer.** Strong has Radeon 680M iGPU (VAAPI works). |
|
||
| 122 | arriman | 4 | 8 GiB | ✅ mp0 | ❌ | Write-heavy (downloads) | *arr stack + qbit + sab. Mounts library for download writes. |
|
||
| 101 | jellyfin | 4 | 8 GiB | ✅ mp0 | ✅ iGPU | Read-heavy sequential | Media streaming + transcode. Strong 680M handles VAAPI. |
|
||
|
||
**Combined RAM freed: 28 GiB.** This alone would eliminate hubris's swap
|
||
pressure entirely.
|
||
|
||
### Tier 3 — Could move, low urgency
|
||
|
||
| ID | Name | Cores | RAM | Library? | Notes |
|
||
|----|------|-------|-----|----------|-------|
|
||
| 130 | grimmory | 1 | 2 GiB | ✅ mp0 | Book library (Docker). Migrated from apps LXC recently. |
|
||
| 131 | teddycloud | 1 | 1 GiB | ✅ mp0 | New (not in inventory.yaml yet). |
|
||
| 132 | rclone | 1 | 2 GiB | ✅ mp0 (ro) | Backup container. Read-only library mount. |
|
||
| 128 | trmnl | 1 | 768 MiB | ❌ | TRMNL middleware. No library. Could move but tiny. |
|
||
| 119 | sophia | 2 | 1 GiB | ✅ mp0 | Workshop. Light use. |
|
||
|
||
### Stay on hubris (core infrastructure)
|
||
|
||
| ID | Name | Cores | RAM | Why it stays |
|
||
|----|------|-------|-----|--------------|
|
||
| 121 | caddy | 1 | 512 MiB | Reverse proxy — terminates all `*.hubris.network`. Must stay on hubris for LAN-side reachability. **Needs backend IP updates** when guests move. |
|
||
| 107 | dns | 1 | 1 GiB | Technitium DNS, split-horizon. Core. |
|
||
| 106 | auth-outpost | 1 | 512 MiB | Authentik SSO enforcement. Core. |
|
||
| 105 | apps | 2 | 4 GiB | homelab MCP + secrets-issuance + artifacto. Core infra. Mounts library. |
|
||
| 104 | gitea | 1 | 1 GiB | Git server. NFS would hurt git lock/stat perf. Mounts library (bare repos). |
|
||
| 103 | paperless | 2 | 3 GiB | Document archive. Moderate I/O, OCR writes. Mounts library. |
|
||
| 102 | nfs-export | 1 | 512 MiB | Exports library to zimaos via NFS. Must stay with the physical library. |
|
||
| 100 | zimaos | 4 | 8 GiB (VM) | NAS frontend eval. Already NFS-mounts library from 102. |
|
||
| 108 | haos | 2 | 4 GiB (VM) | Home Assistant OS. Hardware access, low latency. |
|
||
|
||
---
|
||
|
||
## Constraints & prerequisites
|
||
|
||
### 1. Network — strong needs an internal bridge (Phase 0)
|
||
|
||
strong currently has only `vmbr0` on `192.168.178.0/24`. Guests created there
|
||
get household-LAN IPs, not homelab-subnet IPs. Two options:
|
||
|
||
- **Option A (recommended):** Add `vmbr1` on strong as a portless internal
|
||
bridge with a `192.168.8.x/24` address (e.g. `192.168.8.3`). Route between
|
||
strong's `vmbr0` and `vmbr1` the same way hubris does. Guests go on `vmbr1`
|
||
and get `192.168.8.x` IPs — transparent to Caddy, DNS, and inter-LXC refs.
|
||
Requires adding a static route on Fritz (or relying on hubris's existing
|
||
route — strong would need IP forwarding + a route to 192.168.8.0/24 via vmbr1).
|
||
|
||
- **Option B (simpler, messier):** Put guests on `192.168.178.x` directly.
|
||
Caddy can still reach them (hubris routes to 192.168.178.0/24). But DNS
|
||
records, inter-LXC references, and firewall rules all assume `192.168.8.x`.
|
||
More config churn per guest.
|
||
|
||
### 2. Storage — rootfs migration (no shared storage)
|
||
|
||
`local-lvm` is per-node (not shared). Moving an LXC requires either:
|
||
- `vzdump` → restore on strong (clean, but needs temp disk space + downtime)
|
||
- `rsync` the rootfs to a new LXC on strong (faster for large rootfs like 120's 100G)
|
||
- `pct migrate` only works with shared storage — **not applicable here**
|
||
|
||
For VMs (100, 108): `qm migrate` also needs shared storage. Not moving VMs.
|
||
|
||
### 3. Library access — NFS from hubris to strong
|
||
|
||
Since the physical library SSD is still on hubris, strong's guests that need
|
||
`/mnt/library` must NFS-mount it from hubris. Options:
|
||
|
||
- **Export from hubris host directly** (simplest): add `/mnt/library` to
|
||
`/etc/exports` on hubris with the same squash params as LXC 102
|
||
(`rw,all_squash,anonuid=33,anongid=10000,no_subtree_check`). Mount on strong
|
||
at `/mnt/library`. Strong's guests bind-mount it just like hubris's guests do.
|
||
|
||
- **Use existing nfs-export LXC 102**: strong NFS-mounts from `192.168.8.200`
|
||
(LXC 102). This already has the right squash config. Less host-level change.
|
||
**This is the path of least resistance.**
|
||
|
||
### 4. GPU — iGPU passthrough on strong
|
||
|
||
strong has a Ryzen 7 PRO 6850U with Radeon 680M iGPU. For jellyfin (VAAPI
|
||
transcoding) and mule-images (photo processing), we need:
|
||
- `/dev/dri/renderD128` passed to the LXC (`lxc.cgroup2.devices.allow` +
|
||
`lxc.mount.entry` or Proxmox's `dev0:` passthrough)
|
||
- `video` / `render` group membership inside the container
|
||
- Confirm `amdgpu` driver loads on strong's host kernel (it should — same APU family)
|
||
|
||
### 5. Quorum — 2-node cluster, no QDevice
|
||
|
||
Moving guests to strong does NOT fix the quorum issue but **reduces blast
|
||
radius**: if hubris reboots (its known thermal instability), the guests on
|
||
strong keep running independently. Consider adding a QDevice as a separate
|
||
follow-up — it's orthogonal to this migration.
|
||
|
||
---
|
||
|
||
## Revised migration phases
|
||
|
||
The original plan's NFS-over-LAN approach has been superseded. Instead,
|
||
**media library data moves to ludo-lvm** on strong so migrated guests access
|
||
it as a local ext4 mount. Data is split by origin:
|
||
|
||
```
|
||
hubris (stays): library SSD (3.7T, 1.2T used)
|
||
└── /mnt/library/{documents,images,cloud,homecloud,notes,repos,sophia}
|
||
↑ user-generated content (docs, photos, cloud sync, notes, repos, workshop)
|
||
|
||
strong (moves): ludo-lvm (1.8T, 0 used at start)
|
||
└── /mnt/media_local ← 1.5T thin volume
|
||
└── {downloads,movies,music,tv,anime,books}
|
||
↑ non-user-generated content (media arr stack, book library)
|
||
```
|
||
|
||
| Category | Stays on hubris | Moves to strong |
|
||
|----------|----------------|-----------------|
|
||
| Media | — | downloads (25G), movies (51G), music (29G), tv (30G), anime (206G) |
|
||
| Books | — | books (2.6G) |
|
||
| Docs/Photos | documents (249M), images (4K) | — |
|
||
| Cloud sync | cloud (287G), homecloud (367G) | — |
|
||
| Personal | notes (6.7M), repos (84M), sophia (151G) | — |
|
||
| **Total** | **~805G** | **~344G** |
|
||
|
||
ludo-lvm (1.8T) fits all media + books with ~1.15T headroom for growth.
|
||
hubris library SSD (3.7T, 1.2T used) retains the user-generated content.
|
||
Both sides keep their data local — no cross-node NFS needed for daily I/O.
|
||
|
||
---
|
||
|
||
### Phase 2a — Prepare ludo-lvm on strong
|
||
|
||
1. Create a ext4 filesystem on ludo-lvm for media:
|
||
```bash
|
||
lvcreate -n media -L 1.5T ludo-lvm
|
||
mkfs.ext4 /dev/ludo-lvm/media
|
||
```
|
||
2. Mount at `/mnt/media_local` on strong, add to `/etc/fstab`
|
||
3. rsync media directories from hubris → strong:
|
||
```bash
|
||
rsync -av --progress /mnt/library/{movies,tv,anime,downloads,music,books} strong:/mnt/media_local/
|
||
```
|
||
|
||
### Phase 2b — Migrate arriman (122) to strong
|
||
|
||
1. Stop arriman on hubris, dump rootfs (24G)
|
||
2. Restore on strong with IP `192.168.8.245/28` on vmbr1
|
||
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (downloads land locally)
|
||
4. Update Caddy: jellyseerr/qbit/sab backends → new IP
|
||
5. Update inventory.yaml
|
||
|
||
### Phase 2c — Migrate jellyfin (101) to strong
|
||
|
||
1. Stop jellyfin on hubris, dump rootfs (16G)
|
||
2. Restore on strong with IP `192.168.8.246/28` on vmbr1
|
||
3. Pass `/dev/dri/renderD128` + `/dev/dri/card0` (Radeon 680M + RX 7600)
|
||
4. Mount `/mnt/media_local` → `/mnt/library` via mp0 (media reads locally)
|
||
5. Update Caddy: `media.hubris.network` → new IP
|
||
6. Reinstall `sso-inject.js` in web dir (lost on every apt upgrade)
|
||
7. Test VAAPI transcoding, SSO login, media playback
|
||
|
||
### Phase 2d — Migrate grimmory (130) to strong
|
||
|
||
1. Stop grimmory on hubris, dump rootfs (16G)
|
||
2. Restore on strong with IP `192.168.8.247/28` on vmbr1
|
||
3. Mount `/mnt/media_local` → `/mnt/library` via mp0 (books read locally)
|
||
4. Update Caddy: `books.hubris.network` → new IP
|
||
5. Update inventory.yaml
|
||
6. Test: book browsing, calibre-web access
|
||
|
||
### No NFS export needed
|
||
|
||
With the data split by origin, hubris guests that only need user-generated
|
||
content (documents, images, cloud, repos, sophia) still access them from the
|
||
original library SSD — no cross-node NFS required. The two sides are
|
||
independent.
|
||
|
||
**Result after Phase 2: hubris frees 26 GiB RAM (4 migrated guests) + 344G of
|
||
library I/O burden. Strong becomes the media/books powerhouse.**
|
||
|
||
---
|
||
|
||
### Phase 3 — Migrate mule-images (120) to strong
|
||
|
||
Move photo management (12 GiB RAM, 6 cores, iGPU) last because it needs:
|
||
- `/mnt/library` access (now NFS from strong — already set up in Phase 2d)
|
||
- `/dev/dri/renderD128` (Radeon 680M — confirm VAAPI compatibility first)
|
||
|
||
Steps:
|
||
1. Stop mule-images on hubris, rsync the 100G rootfs to strong (faster than vzdump)
|
||
2. Restore on strong with IP on vmbr1
|
||
3. Pass Radeon 680M iGPU
|
||
4. Reconfigure library paths → `/mnt/media_local` (or keep NFS mount)
|
||
5. Update Caddy: `photos.hubris.network` → new IP
|
||
6. Test photo import + processing pipeline
|
||
|
||
---
|
||
|
||
### Phase 4 — Tier 3 moves (optional)
|
||
|
||
Migrate grimmory (130), teddycloud (131), rclone (132), trmnl (128), sophia (119)
|
||
as needed — each frees 1–2 GiB. Not urgent; do when convenient.
|
||
|
||
---
|
||
|
||
### Phase 5 — Follow-up
|
||
|
||
- **QDevice**: add a tiebreaker for 2-node quorum
|
||
- **Gaming VM**: strong's 6850U has enough cores alongside migrated LXCs
|
||
- **Hubris library cleanup**: after all guests are confirmed working, decide
|
||
whether to keep the original library SSD as backup or repurpose it
|
||
|
||
---
|
||
|
||
## Resource math after Phase 3 (all Tier 1 + 2 moved)
|
||
|
||
| | hubris | strong |
|
||
|---|--------|--------|
|
||
| Guests | 11 LXC + 2 VM | 5 LXC |
|
||
| RAM allocated | ~25 GiB | ~45 GiB |
|
||
| RAM capacity | 28 GiB | 28 GiB |
|
||
| Overcommit | 0.9× (under-committed) | 1.6× (manageable) |
|
||
| Library disk | Local ext4 (3.7T) → NFS client | Local ext4 on ludo-lvm (1.8T) |
|
||
| GPU | Radeon 760M (idle) | Radeon 680M (jellyfin + mule-images) |
|
||
|
||
strong becomes the media/library powerhouse. hubris becomes a lean core-infra
|
||
node (DNS, auth, git, docs, caddy, HA).
|
||
|
||
---
|
||
|
||
---
|
||
|
||
## Risk register
|
||
|
||
| Risk | Impact | Mitigation |
|
||
|------|--------|------------|
|
||
| NFS latency for library reads (jellyfin, arriman) | Media playback stutter, slow downloads | Test iperf between strong↔hubris first. If 2.5G link, NFS throughput is fine (~1 Gbit/s). |
|
||
| GPU passthrough on strong (680M vs 760M) | Transcode quality/compat differences | Both are AMD VAAPI — same driver stack. Test `vainfo` inside LXC before going live. |
|
||
| Caddy backend IP churn | Service outage if IP wrong | Update Caddyfile in git repo (caddy-conf), test each route before destroying old LXC. |
|
||
| vzdump/restore downtime | Service unavailable during migration | Schedule off-hours. Use rsync for large rootfs (120's 100G) to minimize freeze window. |
|
||
| 2-node quorum still fragile | If hubris goes down, strong /etc/pve goes read-only | Guests keep running. Add QDevice as follow-up. |
|
||
| Library data integrity during NFS transition | Permission drift | NFS `all_squash,anonuid=33,anongid=10000` matches existing LXC 102 config. Verify with `ls -la /mnt/library` after mount. |
|
||
|
||
---
|
||
|
||
## Open questions for operator
|
||
|
||
1. **Internal bridge on strong**: proceed with `vmbr1` on `192.168.8.3/24`
|
||
(Option A), or use `192.168.178.x` guest IPs (Option B)?
|
||
2. **Migration method**: `vzdump`/restore (clean, downtime) vs `rsync` rootfs
|
||
(faster for large disks, needs manual config copy)?
|
||
3. **Phase 1 priority**: move elementsynapse + house first (quick wins), or
|
||
go straight to Phase 2 (mule-images/jellyfin/arriman) for maximum relief?
|
||
4. **Should we add a QDevice now** before moving anything, to protect
|
||
management plane during the migration?
|
||
|
||
---
|
||
|
||
## Changelog
|
||
|
||
### 2026-07-05 — Phase 2d complete (grimmory migrated; media NFS to zimaos)
|
||
grimmory (130) → 192.168.8.247 on strong. Rsync'd /books (2.6G) to ludo-lvm.
|
||
LXC 102 (nfs-export) now mounts strong's NFS at /mnt/media and exports it as
|
||
a second share alongside /mnt/library. Zimaos mounts both: /media/library
|
||
(hubris user-generated) and /media/media (strong media+books).
|
||
See hosts/strong.md changelog.
|
||
|
||
### 2026-07-05 — Phase 2 complete (arriman + jellyfin migrated; library on ludo-lvm)
|
||
arriman (122) → 192.168.8.245, jellyfin (101) → 192.168.8.246. Created 1.5T
|
||
thin volume on ludo-lvm, rsync'd 363G of media data. Both containers use local
|
||
ext4 mount — no NFS. Jellyfin has 680M + RX 7600 GPU passthrough.
|
||
Caddy backends updated. See hosts/strong.md changelog.
|
||
|
||
### 2026-07-05 — Phase 1 complete (elementsynapse + house migrated to strong)
|
||
Both Tier 1 guests moved: elementsynapse (118) → 192.168.8.242, house (129) → 192.168.8.244.
|
||
Strong now has vmbr1 at 192.168.8.241/28. Hubris has proxy ARP + /32 routes for strong
|
||
guest range. DHCP scope narrowed to 192.168.8.100-239 to avoid conflicts.
|
||
Teddycloud (LXC 131) given static IP 192.168.8.150 due to IP conflict with
|
||
previous DHCP allocation at 192.168.8.243.
|
||
See hosts/strong.md changelog for full steps.
|
||
|
||
### 2026-07-05 — assessment created
|
||
Built from live `pct config` + `pvesm status` + `free -h` data pulled from
|
||
both nodes. Supersedes the storage-migration framing of the original
|
||
library-SSD plan — this assessment treats the SSD move as optional and
|
||
focuses on guest relocation via NFS.
|