netbird mgmt might reject unknown top-level JSON keys depending on
parser flags. The rendering note stays in turnserver.conf (# comments
are valid INI syntax). After this, the rendered management.json will
be byte-identical to what's currently live on the VPS, so the dry-run
will show /opt/management.json as unchanged.
Lets the operator see exactly what would change on the VPS before
applying. Secret values from the decrypted sops files are masked as
<REDACTED> so the diff is safe to paste into chat/PRs.
The IONOS netbird VPS held two credentials in plaintext that were the last
holdouts from the homelab's sops+age secrets pattern:
- /root/turn-pass.txt (coturn long-term-credential password)
- PKCEAuthorizationFlow.ProviderConfig.ClientSecret inline in
/opt/management.json (Authentik OIDC client secret)
This commit moves both into sops-encrypted YAML in the repo and adds a render
command that recreates the VPS config files from templates + decrypted secrets:
* secrets/turn-shared-secret.yaml — encrypted `password: <coturn pwd>`
* secrets/netbird-authentik-oidc.yaml — encrypted `client_secret: <...>`
Both recipients = hubris + apps + republic-laptop (same 3 as hello.yaml).
* vps/turnserver.conf.tmpl + vps/management.json.tmpl — templates with
{{TURN_PASSWORD}} + {{AUTHENTIK_CLIENT_SECRET}} placeholders.
* bin/homelab new subcommand `render-vps-configs`:
- Decrypts both secrets locally (works on any recipient).
- Substitutes placeholders into templates.
- Diffs against current VPS state via ssh, prompts, applies atomically
(write `.new` then mv), restarts coturn + netbird-mgmt.
- --dry-run + -y flags. Hops through hubris when not running on hubris
itself, since VPS sshd is locked to hubris's pubkey.
* inventory.yaml adds the VPS as `kind: external` (new kind; reserved for
ssh-managed hosts that aren't homelab clients themselves — no age key,
no /opt/homelab-context). hosts/netbird-vps.yaml regenerated.
* SHARED_SECRETS list includes both new secrets so re-keys on enrollment
changes pick them up automatically.
After this lands + the 5-min sync propagates to hubris, run from hubris PVE
shell (or any client; hubris just skips the extra ssh hop):
homelab render-vps-configs --dry-run # see plan, no changes
homelab render-vps-configs -y # apply + restart services
Once verified working, the plaintext `/root/turn-pass.txt` should be deleted
on the VPS (the rendered /etc/turnserver.conf no longer needs it as a
reference).
Three coordinated changes addressing follow-ups from the 2026-05-21 netbird vanilla
migration, plus a related off-LAN ergonomics fix:
bin/homelab:
- New ssh_target(name, force_mesh=False) helper resolves (addr, port, user)
from inventory, honoring ssh.netbird_port (forces mesh path) and ssh.user
(default "root"). Falls back to the netbird FQDN when LAN IP fails a
cached 1.5s TCP probe — helps off-LAN operators on VPN/symmetric-NAT
paths where the netbird subnet route doesn't reach 192.168.8.0/24.
- New ssh_base() builds the full `ssh ... user@addr` invocation; hubris_ssh()
is now a back-compat shim. cmd_ssh, cmd_logs, cmd_restart, cmd_nuke
refactored to use it — no more hardcoded "root@" anywhere.
inventory.yaml:
- New ssh.user convention (root by default, explicit per workstation).
- hubris.ssh.user=root (explicit, documents convention).
- republic-laptop, mac-mini: ssh.user=dtoro. ludo-mini left default (TODO).
- Comment block in the header explains the field + why it exists (netbird-ssh
defaults to LOCAL username; "user not found" on LXCs is the gotcha).
- hosts/*.yaml regenerated from build_host_files.py.
infrastructure/mesh.md:
- Migration changelog entry updated: Device Code Stage is now configured
(was "NOT yet" — landed in d41d73f); --setup-key no longer required.
- New subsection documenting the post-migration JWT-issuer cache bug:
client/internal/engine_ssh.go's updateSSH() bails out when sshServer is
already running, so systemctl restart and netbird down/up don't refresh
the SSH JWT validator. Full daemon stop/start is the fix.
- Companion username gotcha (`netbird ssh` defaulting to local username).
operations/agent-enrollment.md:
- Three new troubleshooting rows: JWT-issuer cache, user-not-found, and
LAN-unreachable-from-mesh-peer (the new homelab CLI behavior).
Verification: ssh_target resolution against the live inventory yields
- hubris → ssh -p 22022 root@proxmox-server.netbird.selfhosted (mesh-forced)
- jellyfin/gitea → ssh root@192.168.8.x (LAN reachable, probe passed)
- republic-laptop/mac-mini → ssh dtoro@<fqdn> (per ssh.user)
- ludo-mini → ssh root@<fqdn> (default)
Two changes that together collapse new-workstation onboarding from ~7 steps
to ~2 commands:
* bootstrap.sh:
- Dep-check now AUTO-INSTALLS missing tools (apt/dnf/brew) instead of
printing instructions and exiting. Re-verifies after install.
- New pre-mesh-check block: if netbird isn't installed, installs it
from the netbird apt/dnf repo (or `brew install --cask netbird` on
Darwin), then if mgmt isn't connected, runs `netbird up
--management-url=https://netbird.hubris.network --ssh-jwt-cache-ttl 86400`.
Operator clicks the device-code URL inline. Waits up to ~30s for
Management: Connected before continuing. Skipped on --no-secrets +
--dry-run.
* containers/124-authentik.md: replaces the "KNOWN MISSING — Device Code
Stage" subsection with a working recipe — Authentik 2026.2 routes
/device via a BRAND-level "Device code flow" field, not a provider
field. Documented stage bindings for a `default-device-code-flow`
flow (identification → password → user-login → consent) and the
brand-level binding step.
* operations/agent-enrollment.md: Path B (interactive `netbird up`) is
now the default; Path A (setup-key) demoted to "headless/scripted"
alternative. "Install dependencies" section collapsed into a note
that bootstrap handles it, with the manual recipes kept in a
collapsible <details> block for air-gapped use.
The flow uniquely available to lab owners (single Authentik user today)
still relies on the first-time-owner sqlite promotion documented in
124-authentik.md.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Updates to four pages reflecting the combined → vanilla mgmt+signal+relay+coturn
cutover and the IONOS-3478-firewall-exception discovery:
* infrastructure/mesh.md — rewrites the ICE/STUN section to cover the new
TURN endpoint, the IONOS upstream TCP-3478 filtering (load-bearing,
undocumented before today), and the verification probe. New changelog
entry covering the migration outcome + Device Code Stage gap.
* infrastructure/vps-hardening.md — "At a glance" lists the new 6-service
docker stack + host coturn. Firewall section notes the new
`iifname ens6 tcp dport 3478 accept` rule plus the IONOS upstream
exception. New changelog entry.
* containers/124-authentik.md — replaces the "Netbird IdP integration —
DEFERRED" section with the LANDED state: Provider details (Public
client type — Confidential breaks PKCE on the dashboard SPA), the
first-time owner-promotion sqlite recipe, the missing Device Code
Stage gap + workaround (setup-keys), and a note that the old 2026-04-22
pre-work Provider/App is now obsolete and safe to delete. Updated
changelog (Phase 6 landed).
* operations/agent-enrollment.md — new "Getting onto Netbird" subsection
explaining the setup-key path (currently the only working flow until
Device Code Stage lands) and why direct OIDC from the public internet
fails (auth.hubris.network is mesh-only-reachable). Prerequisites table
row updated to point at the new section.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Bundles the documentation slice of the apt-sweep backlog:
* operations/runbook-dpkg-interrupted.md (NEW) — Path A (ssh-reachable
recovery) + Path B (PVE web Shell when the netbird mesh broke
alongside the dpkg state, as happened during Wave 6 on hubris).
Closes B2.
* operations/commands.md — new "Fleet apt operations" section
documenting `homelab apt-audit` and `homelab apt-upgrade`
(--status / --safe / --force). Adds the dpkg-interrupted runbook to
Related.
* operations/agent-enrollment.md —
- new "Claude Code permissions for fleet ops" section with the
`permissions.allow` snippet (`Bash(ssh -p 22022 *)`,
`Bash(homelab *)`) for `~/.claude/settings.json`. Closes A3.
- two new Troubleshooting rows: chat-mode `!` sudo no-tty gotcha
(G2) and the cosmetic netbird DNS-probe warning.
* infrastructure/auto-deploy.md — new "Custom-built binaries that
overlap apt-managed paths" section describing the two acceptable
patterns (epoch-versioned .deb à la caddy 1:2.11.3-hubris1; or
apt-mark hold) and the discovery path via `homelab apt-audit`'s
NONAPT column. Closes D3.
Remaining backlog after this commit: A4 (upstream OpenSSH/netbird mux
bug), D1 (apt-mark hold caddy in caddy-conf bootstrap — superseded
in practice by the epoch .deb), E1/E2 (LXC DNS fallback for
tailscale-managed resolv.conf), F1/F2 (vzdump fallback doc; F3 already
shipped).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
G1 from the apt-sweep backlog. After ssh ControlMaster setup, install
the `mcp[cli]` python package via pipx so `homelab mcp <tool>` works
out-of-the-box on new workstations. Skipped on LXCs / VMs.
Idempotent (`command -v mcp` guard), respects --dry-run, falls back
through brew (Darwin) → dnf → apt for pipx itself if not already
installed. Runs the install as $SUDO_USER (not root) so the binary
lands in the user's pipx env.
This closes one of the "discovered missing dep" gaps from the 2026-05-21
fleet sweep: republic-laptop had pipx-installed homelab CLI but no `mcp`
binary, so `homelab mcp <tool>` died with an instructional message
instead of just working.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
F3 from the apt-sweep backlog. `homelab apt-upgrade --safe ...` takes a
pre-upgrade snapshot per LXC before launching the apt run, so rollback is
trivial if anything regresses.
Snapshot strategy per target:
- LXC: try `pct snapshot <id> preupgrade_<timestamp>` first (CoW, near
instant). If that refuses ("snapshot feature is not available" — the
failure mode for LXCs with host bind-mounts like `mp0: /mnt/library`),
fall back to `vzdump <id> --mode snapshot --storage local --compress
zstd`. Slower but works on bind-mounted LXCs.
- hubris (PVE host): skipped intentionally; no host-level snapshot in v1.
If any snapshot fails the entire run refuses unless --force is passed.
Snapshot rollback hints are printed after launch so the operator has the
recovery path one copy/paste away.
Validated 2026-05-21: py_compile clean; --status mode works on bind-
mounted LXCs (skips snapshot path). Live snapshot test deferred to next
real fleet sweep to avoid leaving stale artifacts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two new subcommands (B1+H1 + C1+C2 from the apt-sweep backlog):
* `homelab apt-audit [--target HOST]`
Per-host pre-flight table reporting dpkg-interrupted state, holds,
upgradable count, non-apt binaries in system paths, and DNS health.
Exits nonzero if any target has unconfigured packages (would catch the
latent dpkg-interrupted state we found on caddy + nextcloud before
touching them).
* `homelab apt-upgrade [--target HOST | --all] [--status] [--force]`
Launches `apt update && apt upgrade` inside a transient systemd unit
(`apt-upgrade-<hostname>`) on each target via `systemd-run --collect`,
so the run survives ssh teardown. Output tee'd to
`/var/log/homelab-apt-upgrade.log`. Apt configured with
`Acquire::Retries=3` + `ForceIPv4=true` to absorb transient mirror
blips (would have prevented the jellyfin ffmpeg7 fetch failure that
rolled back 131 packages mid-Wave-3).
Pre-flight audit gate refuses to run on hosts with dpkg-interrupted
state; override with --force.
`--status` mode prints each target's unit state and tails the upgrade
log; safe to run before, during, and after.
Also fixes a latent bug in `hubris_ssh()`: the function joined hubris's
`lan_ip` (192.168.8.77) with the netbird-only port 22022. Port 22022
listens on hubris's netbird interface ONLY — not on the LAN interface,
not even via the netbird subnet route. Result: every off-LAN client
hitting `homelab pct/logs/restart/status` got "Connection refused".
Now uses the netbird FQDN when netbird_port is in use; falls back to
host_address otherwise.
Targets exclude VMs and workstations (they have their own update flows).
Standard apt targets = hubris + every LXC (kind=lxc with a pve_id).
Validated 2026-05-21 on republic-laptop:
- apt-audit reported clean dpkg + 40 upgradable on hubris (the still-
deferred PVE/kernel packages) + 0 on every LXC, DNS ok across the
fleet
- apt-upgrade --target mule-photos-new launched the unit, ran apt to
completion (0 upgrades pending), log captured with =START= / =END
rc=0= markers, --status reports unit + log correctly
- --all --status fans out across all 16 targets without issue
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After preflight detects MESH_CONNECTED=netbird and the host yaml says
kind != lxc (i.e. workstation, vm, proxmox-host), drop a Host block into
the enrolling user's ~/.ssh/config:
Host *.netbird.selfhosted
ControlMaster auto
ControlPath ~/.ssh/cm/%C
ControlPersist 2h
This is the real workaround for netbird's SSH JWT cache being flaky in
0.71.2 — we discovered that --ssh-jwt-cache-ttl can leave the daemon in
a state where stale cached tokens get sent and rejected with no fallback
to fresh SSO. ssh ControlMaster bypasses netbird-ssh-proxy entirely for
subsequent sessions: one SSO at the start of a working window covers all
back-to-back ssh / scp / `pct exec` ops until ControlPersist expires.
Validated 2026-05-21 on republic-laptop: ssh #1 prompted SSO once,
sshes #2 and #3 ran in ~1.1s each with no prompt.
Idempotent (sentinel comment check); writes ~/.ssh/cm/ with 700; uses
SUDO_USER's home when invoked via sudo.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
After preflight detects MESH_CONNECTED=netbird, run `netbird down && netbird up
--ssh-jwt-cache-ttl=86400` so ssh into mesh peers (e.g. `ssh -p 22022
root@proxmox-server.netbird.selfhosted ...`) stops triggering device-code SSO
on every connection.
Validated 2026-05-21 on republic-laptop: after one SSO, subsequent ssh
sessions within 24h skip the device-code flow and run instantly. Fleet
operations (e.g. pct exec through hubris into LXCs) reuse the cached JWT.
Notes:
- Flag is supported in netbird 0.71.x+ (netbirdio/netbird#4015). A version
probe (`netbird up --help | grep ssh-jwt-cache-ttl`) skips the section on
older clients.
- Flag belongs on `netbird up` (client config), NOT on the daemon's
ExecStart — putting it there crashes the daemon with "unknown flag".
- Runs LAST in bootstrap, after secrets issuance + MCP wiring, so the brief
mesh down/up doesn't disrupt earlier steps.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two additions:
1. secrets-issuance backup: daily timer snapshots /var/lib/secrets-issuance
to /mnt/library/.secrets-issuance-backup/ as a date-stamped tar.gz,
keeping the last 14 days. Closes the catastrophic-fail-mode where an
LXC 105 loss wipes every client's age key with no recovery path.
Caveat: privileged LXCs that mount /mnt/library can read the backup
(root-uid maps to host root); encrypted-tarball variant is a future
refinement.
2. homelab doctor: 10 invariant checks for an enrolled client — clone
present, sync timer/launchd job active, age key perms, CLI symlinked,
AGENTS.md linked, inventory entry exists, MCP reachable, secrets
/health responds, sops canary decrypts, git creds present. Returns
nonzero on any 'fail'. Useful after enrollment or whenever something
smells off.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Discovered via ARP (MAC 02:E1:73:18:EA:49 from qm config). The Tailscale
FQDN 'homeassistant' is fine for tailscale peers but unreachable from
Netbird peers like republic. lan_ip works from both — Netbird routes
192.168.8.0/24 through hubris.
zimaos (VM 100) remains without lan_ip because its IP wasn't in the
hubris ARP table at audit time and we don't want a network scan. The
zimaos service is still reachable via caddy at zimaos.hubris.network
(verified in 'homelab status' SERVICE column).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Audit against actual netbird+tailscale peer lists:
- hubris is the only LXC-host on Netbird; only workstations + hubris
have netbird entries
- 10 LXCs+VMs have real Tailscale FQDNs: apps, jellyfin, paperless,
gitea, nextcloud, elementsynapse, sophia, mule-images→muleimage,
arriman→arr, haos→homeassistant
- 7 hosts are LAN-only (no mesh block): nfs-export, caddy, claudio-bot,
authentik, plato, mule-photos-new, zimaos
- mac-mini's netbird FQDN corrected to the actual peer name
(mac-mini-234-17.netbird.selfhosted)
Also: bin/homelab host_address() now prefers lan_ip first — universally
reachable from any LAN client and from any Netbird peer via the
192.168.8.0/24 network resource routed through hubris. Mesh FQDNs are
fallbacks for roaming workstations without a fixed lan_ip.
This makes 'homelab status' from republic show all backends 'ok' instead
of falsely reporting them 'down' against unresolvable netbird FQDNs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
End of enrollment, if secrets/gitea-pat.yaml is decryptable with the
just-issued age key (i.e. the operator has already run --finalize-pubkey
from another client), upgrade /etc/homelab-context/git-credentials from
the read-only bootstrap PAT to the write-scoped one. Best-effort: fails
silently if not yet a recipient, with a clear hint about what to do next.
Removes the post-bootstrap manual 'homelab refresh-creds' step from the
common flow; falls back to the documented one-liner for first-bootstrap
clients.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Service keys in inventory use _ for python-attribute friendliness but the
actual systemd units use dashes. Explicit systemd_unit field disambiguates
for MCP tail_log / get_service_status.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Unit was hard-coding HOMELAB_MCP_SSH_USER=mcp-reader, overriding the new
code's 'root' default. No mcp-reader user exists on hubris — the key is
authorized for root, with a strict command= wrapper. Also pin
HOMELAB_MCP_HUBRIS_HOST + KNOWN_HOSTS env so the service doesn't depend
on the python defaults staying in sync with the unit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The systemd unit's ProtectHome=true blocks ~/.ssh access. SSH then had
no place to write known_hosts (StrictHostKeyChecking=accept-new) and
silently produced empty results. Use /etc/homelab-mcp/known_hosts (which
ProtectSystem=strict still allows reading) and StrictHostKeyChecking=yes.
Operator pre-populates the file via:
ssh-keyscan -t ed25519 192.168.8.77 > /etc/homelab-mcp/known_hosts
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The 5 management tools (get_service_status, tail_log, list_lxcs,
get_lxc_state, ping_service) were registered with sse but all SSH calls
went to per-host targets via a 'mcp-reader' user that didn't exist
anywhere. New design routes every management call through ONE channel:
LXC 105 -> hubris (SSH key + restricted authorized_keys command), then
hubris pct-execs into the right LXC where needed.
Adds mcp/mcp-reader-shell — a strict allowlist wrapper read from
$SSH_ORIGINAL_COMMAND. Rejects shell metacharacters up front and then
matches against a fixed set of read-only patterns (systemctl is-active/
is-enabled, journalctl -u, pct list/status/config, pct exec for the
same subset). Logged to syslog tag mcp-reader.
Authorized_keys line on hubris:
command="/usr/local/bin/mcp-reader-shell",restrict ssh-ed25519 ... mcp-reader@homelab-mcp
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
yaml.safe_load + safe_dump stripped every comment from inventory.yaml
on each enrollment, eroding the file's documentation value. New
_inventory_set_age_pubkey / _inventory_remove_host / _inventory_append_host
helpers do line-based edits so comments outside the modified region
survive. inventory.yaml gets its top-of-file conventions block back.
build_host_files.py round-trips cleanly (--check returns 0).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Critical bug — the loop reset age_block_last_idx unconditionally on
every '- path_regex:' line. So after finding the target rule's age
block, encountering the NEXT rule wiped the result and the function
returned False. _grant_shared_secrets / _revoke_shared_secrets both
silently no-op'd because of this.
Fix: break out of the scan once we've collected the target rule's
data. Refactored the remove helper to share a find_target_age_lines
inner so the dangling-trailing-comma fixup uses the same logic.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The remove flow ran 'sops updatekeys' but never edited .sops.yaml first,
so the removed client stayed a recipient on every shared secret —
exactly the opposite of what 'remove' should do. Adds the
_remove_recipient_from_sops_policy / _revoke_shared_secrets helpers
(inverse of the grant-side ones from the previous commit); cmd_client_remove
now resolves the pubkey from inventory before deletion and feeds it
through that pipeline.
Also adds the sudo re-exec pattern so 'homelab client remove' works
from a non-root user (matching cmd_secret / cmd_refresh_creds /
cmd_client_add).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
LXCs without a mesh CLI sit on 192.168.8.0/24 which is in the
issuance service's MESH_SUBNETS — they should be able to bootstrap
without netbird/tailscale installed. New third path probes the
issuance /health endpoint directly; if reachable, treat that as
satisfying the mesh precondition.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds infrastructure/homelab-context.md as the architecture reference for
the cross-client context + MCP + secrets-issuance system. Updates:
- 105-apps.md: two new ## Stacks sections (homelab-mcp, secrets-issuance)
with their deploy pipelines + a row each in the public-hostname table;
changelog entry.
- auto-deploy.md: both new pipelines added to the table (one repo, two
webhooks, same push); per-pipeline notes covering the clone-per-service
pattern and the deploy.sh self-restart caveat; changelog entry.
- README.md: link to the new infrastructure page.
Operational walkthrough already lives at operations/agent-enrollment.md;
this commit is the architecture side of the same story.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Setting the age_pubkey is half the enrollment; the new client also needs
to be a recipient on shared secrets (hello.yaml, gitea-pat.yaml) to
actually use them. Now --finalize-pubkey:
1. writes hosts.<name>.age_pubkey
2. appends the pubkey to each shared-secret rule in .sops.yaml
(preserving comments via line-by-line edit, not yaml round-trip)
3. runs sops updatekeys -y on each shared file
4. commits inventory + hosts/ + .sops.yaml + secrets/ as one commit
Also: cmd_client_add now re-execs via sudo when invoked as a regular
user (matches the pattern in cmd_secret + cmd_refresh_creds).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds secrets/gitea-pat.yaml (SOPS-encrypted, dtoro PAT with read+write
scopes) so any enrolled client can push to dtoro/Homelab-Docs — not just
where I have SSH. Recipient set = hello.yaml's (hubris, apps, republic);
expand alongside hello.yaml when enrolling new clients.
bin/homelab gains 'refresh-creds': decrypts gitea-pat.yaml, rewrites
/etc/homelab-context/git-credentials with the write token, repoints
git's --system credential helper. Re-execs via sudo for non-root callers
(same pattern as 'homelab secret').
After this lands, 'homelab client add/remove' and wiki edits can run
from any client. The initial bootstrap still needs an operator-supplied
read-only PAT (chicken-and-egg); 'refresh-creds' upgrades the client
to write afterwards.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The systemd sync timer runs git without HOME set, so git config --global
(which writes /root/.gitconfig) is invisible to the timer's process — the
timer fails with 'could not read Username' silently. Switching to
--system writes to /etc/gitconfig which is HOME-agnostic.
Migration for already-bootstrapped hosts captured in agent-enrollment.md
troubleshooting.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The 5-min sync pulls /opt/homelab-context but does not re-install the
CLI. A copy at /usr/local/bin/homelab therefore goes stale after every
CLI fix until someone re-runs bootstrap. Symlinking points
/usr/local/bin/homelab directly at the synced source, so updates land
on the next pull. Doc updated with the one-line migration for hosts
bootstrapped before this commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
/etc/age/key.txt is 0600 root and /etc/age is 0700 root, so the CLI's
Path.exists() check was returning False under regular users — making the
subcommand look broken when bootstrap had actually written the key fine.
Re-exec via sudo preserves the existing UX (one password prompt, then
plaintext) without loosening the key's perms.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Phase 2 first workstation enrolled. age1vf8... is republic-laptop's
issued pubkey; added as a recipient on hello.yaml so the post-bootstrap
decrypt test works there.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Captures the full enrollment flow validated during Phase 2 rollout: per-OS
dep install (dnf/apt/brew), Gitea PAT prerequisite, DNS gotchas, the
bootstrap command, post-bootstrap verification, the homelab client add
ceremony for new inventory entries, secret grant/revoke, and a
troubleshooting table mapping every failure mode we hit during validation
to the commit that fixed it.
Linked from README under Operations.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Caddy + split-horizon DNS now resolve these to LXC 105 (via 121).
Workstations off-LAN reach them via Netbird (192.168.8.0/24 is a
network resource routed through hubris).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Was apt-only on Linux; republic-laptop is Nobara so installs use dnf and
yaml is python3-pyyaml (not python3-yaml).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
We're nftables-gated to mesh+LAN; the browser-attack threat doesn't
apply, and the default whitelist (127.0.0.1/localhost/[::1] only) blocks
every LAN/mesh client.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The MCP service unit expects port 9810 per the inventory; FastMCP only
binds correctly when we set mcp.settings.host/port before run().
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Encrypted to hubris + apps; expand recipients as new clients enrol via
'sops updatekeys -y secrets/hello.yaml'. Tests the full sops + age path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Hubris is Netbird-only and LXC 105 is Tailscale-only; they share LAN
but not a mesh, so on-host \`curl http://192.168.8.205:9820/issue\`
arrives with source IP 192.168.8.77. In a homelab LAN with no
untrusted devices the trust boundary is reasonable; if that changes
later, narrow this.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>