After preflight detects MESH_CONNECTED=netbird and the host yaml says
kind != lxc (i.e. workstation, vm, proxmox-host), drop a Host block into
the enrolling user's ~/.ssh/config:
Host *.netbird.selfhosted
ControlMaster auto
ControlPath ~/.ssh/cm/%C
ControlPersist 2h
This is the real workaround for netbird's SSH JWT cache being flaky in
0.71.2 — we discovered that --ssh-jwt-cache-ttl can leave the daemon in
a state where stale cached tokens get sent and rejected with no fallback
to fresh SSO. ssh ControlMaster bypasses netbird-ssh-proxy entirely for
subsequent sessions: one SSO at the start of a working window covers all
back-to-back ssh / scp / `pct exec` ops until ControlPersist expires.
Validated 2026-05-21 on republic-laptop: ssh #1 prompted SSO once,
sshes #2 and #3 ran in ~1.1s each with no prompt.
Idempotent (sentinel comment check); writes ~/.ssh/cm/ with 700; uses
SUDO_USER's home when invoked via sudo.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After preflight detects MESH_CONNECTED=netbird, run `netbird down && netbird up
--ssh-jwt-cache-ttl=86400` so ssh into mesh peers (e.g. `ssh -p 22022
root@proxmox-server.netbird.selfhosted ...`) stops triggering device-code SSO
on every connection.
Validated 2026-05-21 on republic-laptop: after one SSO, subsequent ssh
sessions within 24h skip the device-code flow and run instantly. Fleet
operations (e.g. pct exec through hubris into LXCs) reuse the cached JWT.
Notes:
- Flag is supported in netbird 0.71.x+ (netbirdio/netbird#4015). A version
probe (`netbird up --help | grep ssh-jwt-cache-ttl`) skips the section on
older clients.
- Flag belongs on `netbird up` (client config), NOT on the daemon's
ExecStart — putting it there crashes the daemon with "unknown flag".
- Runs LAST in bootstrap, after secrets issuance + MCP wiring, so the brief
mesh down/up doesn't disrupt earlier steps.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two additions:
1. secrets-issuance backup: daily timer snapshots /var/lib/secrets-issuance
to /mnt/library/.secrets-issuance-backup/ as a date-stamped tar.gz,
keeping the last 14 days. Closes the catastrophic-fail-mode where an
LXC 105 loss wipes every client's age key with no recovery path.
Caveat: privileged LXCs that mount /mnt/library can read the backup
(root-uid maps to host root); encrypted-tarball variant is a future
refinement.
2. homelab doctor: 10 invariant checks for an enrolled client — clone
present, sync timer/launchd job active, age key perms, CLI symlinked,
AGENTS.md linked, inventory entry exists, MCP reachable, secrets
/health responds, sops canary decrypts, git creds present. Returns
nonzero on any 'fail'. Useful after enrollment or whenever something
smells off.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Discovered via ARP (MAC 02:E1:73:18:EA:49 from qm config). The Tailscale
FQDN 'homeassistant' is fine for tailscale peers but unreachable from
Netbird peers like republic. lan_ip works from both — Netbird routes
192.168.8.0/24 through hubris.
zimaos (VM 100) remains without lan_ip because its IP wasn't in the
hubris ARP table at audit time and we don't want a network scan. The
zimaos service is still reachable via caddy at zimaos.hubris.network
(verified in 'homelab status' SERVICE column).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Audit against actual netbird+tailscale peer lists:
- hubris is the only LXC-host on Netbird; only workstations + hubris
have netbird entries
- 10 LXCs+VMs have real Tailscale FQDNs: apps, jellyfin, paperless,
gitea, nextcloud, elementsynapse, sophia, mule-images→muleimage,
arriman→arr, haos→homeassistant
- 7 hosts are LAN-only (no mesh block): nfs-export, caddy, claudio-bot,
authentik, plato, mule-photos-new, zimaos
- mac-mini's netbird FQDN corrected to the actual peer name
(mac-mini-234-17.netbird.selfhosted)
Also: bin/homelab host_address() now prefers lan_ip first — universally
reachable from any LAN client and from any Netbird peer via the
192.168.8.0/24 network resource routed through hubris. Mesh FQDNs are
fallbacks for roaming workstations without a fixed lan_ip.
This makes 'homelab status' from republic show all backends 'ok' instead
of falsely reporting them 'down' against unresolvable netbird FQDNs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
End of enrollment, if secrets/gitea-pat.yaml is decryptable with the
just-issued age key (i.e. the operator has already run --finalize-pubkey
from another client), upgrade /etc/homelab-context/git-credentials from
the read-only bootstrap PAT to the write-scoped one. Best-effort: fails
silently if not yet a recipient, with a clear hint about what to do next.
Removes the post-bootstrap manual 'homelab refresh-creds' step from the
common flow; falls back to the documented one-liner for first-bootstrap
clients.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Service keys in inventory use _ for python-attribute friendliness but the
actual systemd units use dashes. Explicit systemd_unit field disambiguates
for MCP tail_log / get_service_status.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Unit was hard-coding HOMELAB_MCP_SSH_USER=mcp-reader, overriding the new
code's 'root' default. No mcp-reader user exists on hubris — the key is
authorized for root, with a strict command= wrapper. Also pin
HOMELAB_MCP_HUBRIS_HOST + KNOWN_HOSTS env so the service doesn't depend
on the python defaults staying in sync with the unit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The systemd unit's ProtectHome=true blocks ~/.ssh access. SSH then had
no place to write known_hosts (StrictHostKeyChecking=accept-new) and
silently produced empty results. Use /etc/homelab-mcp/known_hosts (which
ProtectSystem=strict still allows reading) and StrictHostKeyChecking=yes.
Operator pre-populates the file via:
ssh-keyscan -t ed25519 192.168.8.77 > /etc/homelab-mcp/known_hosts
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The 5 management tools (get_service_status, tail_log, list_lxcs,
get_lxc_state, ping_service) were registered with sse but all SSH calls
went to per-host targets via a 'mcp-reader' user that didn't exist
anywhere. New design routes every management call through ONE channel:
LXC 105 -> hubris (SSH key + restricted authorized_keys command), then
hubris pct-execs into the right LXC where needed.
Adds mcp/mcp-reader-shell — a strict allowlist wrapper read from
$SSH_ORIGINAL_COMMAND. Rejects shell metacharacters up front and then
matches against a fixed set of read-only patterns (systemctl is-active/
is-enabled, journalctl -u, pct list/status/config, pct exec for the
same subset). Logged to syslog tag mcp-reader.
Authorized_keys line on hubris:
command="/usr/local/bin/mcp-reader-shell",restrict ssh-ed25519 ... mcp-reader@homelab-mcp
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
yaml.safe_load + safe_dump stripped every comment from inventory.yaml
on each enrollment, eroding the file's documentation value. New
_inventory_set_age_pubkey / _inventory_remove_host / _inventory_append_host
helpers do line-based edits so comments outside the modified region
survive. inventory.yaml gets its top-of-file conventions block back.
build_host_files.py round-trips cleanly (--check returns 0).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Critical bug — the loop reset age_block_last_idx unconditionally on
every '- path_regex:' line. So after finding the target rule's age
block, encountering the NEXT rule wiped the result and the function
returned False. _grant_shared_secrets / _revoke_shared_secrets both
silently no-op'd because of this.
Fix: break out of the scan once we've collected the target rule's
data. Refactored the remove helper to share a find_target_age_lines
inner so the dangling-trailing-comma fixup uses the same logic.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The remove flow ran 'sops updatekeys' but never edited .sops.yaml first,
so the removed client stayed a recipient on every shared secret —
exactly the opposite of what 'remove' should do. Adds the
_remove_recipient_from_sops_policy / _revoke_shared_secrets helpers
(inverse of the grant-side ones from the previous commit); cmd_client_remove
now resolves the pubkey from inventory before deletion and feeds it
through that pipeline.
Also adds the sudo re-exec pattern so 'homelab client remove' works
from a non-root user (matching cmd_secret / cmd_refresh_creds /
cmd_client_add).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
LXCs without a mesh CLI sit on 192.168.8.0/24 which is in the
issuance service's MESH_SUBNETS — they should be able to bootstrap
without netbird/tailscale installed. New third path probes the
issuance /health endpoint directly; if reachable, treat that as
satisfying the mesh precondition.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds infrastructure/homelab-context.md as the architecture reference for
the cross-client context + MCP + secrets-issuance system. Updates:
- 105-apps.md: two new ## Stacks sections (homelab-mcp, secrets-issuance)
with their deploy pipelines + a row each in the public-hostname table;
changelog entry.
- auto-deploy.md: both new pipelines added to the table (one repo, two
webhooks, same push); per-pipeline notes covering the clone-per-service
pattern and the deploy.sh self-restart caveat; changelog entry.
- README.md: link to the new infrastructure page.
Operational walkthrough already lives at operations/agent-enrollment.md;
this commit is the architecture side of the same story.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Setting the age_pubkey is half the enrollment; the new client also needs
to be a recipient on shared secrets (hello.yaml, gitea-pat.yaml) to
actually use them. Now --finalize-pubkey:
1. writes hosts.<name>.age_pubkey
2. appends the pubkey to each shared-secret rule in .sops.yaml
(preserving comments via line-by-line edit, not yaml round-trip)
3. runs sops updatekeys -y on each shared file
4. commits inventory + hosts/ + .sops.yaml + secrets/ as one commit
Also: cmd_client_add now re-execs via sudo when invoked as a regular
user (matches the pattern in cmd_secret + cmd_refresh_creds).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds secrets/gitea-pat.yaml (SOPS-encrypted, dtoro PAT with read+write
scopes) so any enrolled client can push to dtoro/Homelab-Docs — not just
where I have SSH. Recipient set = hello.yaml's (hubris, apps, republic);
expand alongside hello.yaml when enrolling new clients.
bin/homelab gains 'refresh-creds': decrypts gitea-pat.yaml, rewrites
/etc/homelab-context/git-credentials with the write token, repoints
git's --system credential helper. Re-execs via sudo for non-root callers
(same pattern as 'homelab secret').
After this lands, 'homelab client add/remove' and wiki edits can run
from any client. The initial bootstrap still needs an operator-supplied
read-only PAT (chicken-and-egg); 'refresh-creds' upgrades the client
to write afterwards.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The systemd sync timer runs git without HOME set, so git config --global
(which writes /root/.gitconfig) is invisible to the timer's process — the
timer fails with 'could not read Username' silently. Switching to
--system writes to /etc/gitconfig which is HOME-agnostic.
Migration for already-bootstrapped hosts captured in agent-enrollment.md
troubleshooting.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The 5-min sync pulls /opt/homelab-context but does not re-install the
CLI. A copy at /usr/local/bin/homelab therefore goes stale after every
CLI fix until someone re-runs bootstrap. Symlinking points
/usr/local/bin/homelab directly at the synced source, so updates land
on the next pull. Doc updated with the one-line migration for hosts
bootstrapped before this commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
/etc/age/key.txt is 0600 root and /etc/age is 0700 root, so the CLI's
Path.exists() check was returning False under regular users — making the
subcommand look broken when bootstrap had actually written the key fine.
Re-exec via sudo preserves the existing UX (one password prompt, then
plaintext) without loosening the key's perms.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 2 first workstation enrolled. age1vf8... is republic-laptop's
issued pubkey; added as a recipient on hello.yaml so the post-bootstrap
decrypt test works there.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Captures the full enrollment flow validated during Phase 2 rollout: per-OS
dep install (dnf/apt/brew), Gitea PAT prerequisite, DNS gotchas, the
bootstrap command, post-bootstrap verification, the homelab client add
ceremony for new inventory entries, secret grant/revoke, and a
troubleshooting table mapping every failure mode we hit during validation
to the commit that fixed it.
Linked from README under Operations.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Caddy + split-horizon DNS now resolve these to LXC 105 (via 121).
Workstations off-LAN reach them via Netbird (192.168.8.0/24 is a
network resource routed through hubris).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Was apt-only on Linux; republic-laptop is Nobara so installs use dnf and
yaml is python3-pyyaml (not python3-yaml).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
We're nftables-gated to mesh+LAN; the browser-attack threat doesn't
apply, and the default whitelist (127.0.0.1/localhost/[::1] only) blocks
every LAN/mesh client.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The MCP service unit expects port 9810 per the inventory; FastMCP only
binds correctly when we set mcp.settings.host/port before run().
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Encrypted to hubris + apps; expand recipients as new clients enrol via
'sops updatekeys -y secrets/hello.yaml'. Tests the full sops + age path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Hubris is Netbird-only and LXC 105 is Tailscale-only; they share LAN
but not a mesh, so on-host \`curl http://192.168.8.205:9820/issue\`
arrives with source IP 192.168.8.77. In a homelab LAN with no
untrusted devices the trust boundary is reasonable; if that changes
later, narrow this.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The homelab CLI imports yaml; missing on a fresh LXC. Preflight now
checks and emits the right install hint per OS.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
git credential helper does exact prefix match including scheme. Hardcoding
https:// breaks for in-LAN clones using http://192.168.8.121:3000 (which
LXC 105 needs because its DNS doesn't have the split-horizon override).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Each non-hubris client needs HTTPS auth against gitea for the initial
context clone (chicken-and-egg: a PAT stored in SOPS can't be fetched
until after the clone exists). Adds a --gitea-token flag that writes
credentials to /etc/homelab-context/git-credentials and points git's
credential.helper at it, so the clone and all future pulls succeed
without prompting.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add the foundation for distributing homelab context to every client
(LXCs, VMs, workstations including republic-laptop, mac-mini, ludo-mini)
with a single source of truth, structured query layer (MCP), and per-client
age-key issuance for secrets:
- inventory.yaml — canonical topology (hosts, services, mesh addresses)
- hosts/*.yaml — per-host identity files generated from inventory by
mcp/build_host_files.py; do not edit by hand
- AGENTS.md — orientation doc symlinked to /root/AGENTS.md on every client
- bootstrap.sh — one-shot enroll (Linux + macOS), clones repo, fetches age
key from issuance, installs sync timer/launchd job, drops the homelab CLI
- bin/homelab — single-binary Python CLI: whoami, list, ssh, pct, logs,
restart, open, status, secret, sync, mcp, client add/remove, nuke
- mcp/server.py — FastMCP server: context tools + read-only management
tools (no mutations exposed); shell-outs use mcp-reader restricted ssh key
- mcp/deploy/ — claudio-monitor-style gitea webhook deploy scaffold for the
MCP service on LXC 105 (ports 9810 mcp, 9811 webhook)
- secrets-issuance/ — per-client age key auto-provisioning over the mesh;
source-IP gated against inventory, with denylist for revoked clients
(ports 9820 issue, 9821 webhook)
- secrets/, .sops.yaml — SOPS recipient scaffolding; the operator fills in
age public keys after Phase 3a generates them
- scripts/sync/ — systemd timer (Linux) + launchd plist (macOS) pulling
/opt/homelab-context every 5 min
Mesh: both Netbird (preferred, 100.122.0.0/16) and Tailscale accepted
during the in-flight migration; no client is gated on completing the move.
Plan reference: /root/.claude/plans/lets-make-a-plan-fluttering-trinket.md
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PhotoPrism rotates its session HMAC key on every container start, so any
auto-deploy that recreated pp-app invalidated in-flight OIDC logins. The
deploy script was force-recreating + image-pulling on each push; pinned
both so pp-app survives a routine code deploy.
Measured cold-cache behaviour: thumbnails ~2ms, HEVC video playback
12-21s TTFB because libx264 transcodes inline and serialised one
ffmpeg at a time. Sibling thumbs unaffected (~2ms during transcode).
With 11/45 .mov files already pre-baked, ~75% of clicks were cold.
Ran `photoprism convert` to bring sidecar coverage to 45/45; previously
cold videos now serve in ~2ms TTFB.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
NFS was mounted at /DATA/library, so the icewhale-files trash regex
(^/media/([^/]+)) derived drive=ZimaOS-HD and tried to rename into
/media/ZimaOS-HD/.trash on local ext4 -- cross-device EXDEV from NFS.
Remounting at /media/library makes 'library' its own /media/<name>
segment, so trash resolves to /media/library/.trash on the NFS itself.
Synapse runs on SQLite (not Postgres — Postgres only hosts the
mautrix bridge dbs). Documented the post-disk-full
event_push_actions cleanup that took @admin'\''s phantom count
from 125 to 4.