Files
oikos/containers/132-rclone.md
dtoro 6669feafdc docs(rclone): document protondrive silent-stall incident + watchdog
Two silent stalls hit in LXC 132's first 24h of real traffic: rclone's own
--timeout didn't catch a protondrive-specific hang (transfer at 100%, zero
bytes/errors/retries for hours). Added a 5-min watchdog timer that restarts
rclone-backup.service if transferred bytes are frozen for 15+ min. Also
found and fixed a monitoring bug in the runner (wrong stats-group key) that
made a healthy sync look falsely stalled for 22h in its own log.
2026-07-03 12:32:53 +02:00

12 KiB

132 — rclone

Off-host backup appliance. Mirrors selected /mnt/library folders to Proton Drive with a plain rclone sync (monthly), and serves rclone's Web GUI on the LAN for browsing and ad-hoc runs. Replaces the disabled restic-on-USB job — see backups.

Provisioned 2026-07-01. (LXC 131 was already taken by an undocumented teddycloud container, so this landed on 132.)

At a glance

  • Hostname: rclone
  • IP: 192.168.8.214 (static, set in PVE net0 config — same pattern as grimmory/authentik)
  • Privilege: privileged (root in-container = host root → reads every /mnt/library subtree, incl. homecloud/ and documents/, regardless of owner)
  • Resources: 1 core / 1 GiB RAM / 8 GiB rootfs (Debian 13)
  • Mounts: /mnt/library read-only (mp0: /mnt/library,mp=/mnt/library,ro=1) — a backup job must never be able to write into the library
  • Public hostname: none — the UI is LAN-only, no auth (by design)

Service / port map

Service Listen Notes
rclone Web GUI (rcd) 192.168.8.214:5572 rclone-rcd.service, --rc-no-auth, LAN-only. Browse /mnt/library + proton:, run ad-hoc syncs, live job status
monthly mirror rclone-backup.service + .timer (OnCalendar=*-*-01 03:00)

Backup design

  • Mode: plain mirror — rclone sync (Proton mirrors local; deletions propagate; no versioning).
  • Encryption: Proton Drive's built-in E2E only (no rclone crypt overlay → files stay browsable in Proton's web UI).
  • Selected set: /etc/rclone-backup/folders.list — one absolute source path per line (#/blank ignored). This file is the picked set the monthly timer mirrors. Extensible to other disks once bind-mounted into this LXC.
  • Path mapping: source Sproton:library-backup/<S without leading slash> (e.g. /mnt/library/notesproton:library-backup/mnt/library/notes).
  • Runner: /usr/local/sbin/rclone-backup.sh [folder ...] (Python, despite the .sh name — kept the path stable) — no arg = every enabled line. Submits each folder as an async job through the rclone rc API served by rclone-rcd.service (the same daemon backing the Web GUI on :5572), so scheduled/ad-hoc runs show up live in the GUI's Jobs panel, not just in logs. Gentle on Proton's rate limits (Transfers=4, TPSLimit=8, FastList=true via the rc _config payload). The rc API here requires POST for every call including job/status and core/stats — GET with query params 404s.
  • Logs / "past runs": per-run logs in /var/log/rclone-backup/<safe>-<ts>.log; one-line JSON summary per run appended to /var/log/rclone-backup/runs.jsonl.
  • Failure notify: OnFailure=rclone-backup-notify@%n.service → logs to journal today; TODO wire to Hermes send_message (Matrix) per backups.

rclone + Proton Drive

  • rclone installed from the official binary (not apt) so the protondrive backend is present (rclone v1.74.3).
  • Remote proton: (type protondrive). Config at /root/.config/rclone/rclone.conf, mode 600. This file is a secret (holds the obscured Proton password + TOTP secret + session) — never commit it. Escrow the Proton account creds in the password manager.
  • Config gotchas (from rclone docs/forum):
    • Log into Proton via a browser at least once first, or key generation fails.
    • For unattended runs, store the TOTP secret (not a 6-digit code) so rclone self-generates codes; obscure with rclone obscure.
    • Passwords with extended-ASCII characters are known to break auth.
    • Proton's API is rate-limited → keep --transfers/--tpslimit conservative (baked into the runner).
  • DR escrow (pending): store the Proton creds as sops secret secrets/protondrive.yaml, granted to this LXC's age key, so the remote can be rebuilt after a re-provision.

The UI (rclone Web GUI)

rclone rcd --rc-web-gui --rc-no-auth --rc-addr 0.0.0.0:5572 (assets auto-downloaded on first start). Reach it at http://192.168.8.214:5572 on the LAN.

Security note: --rc-no-auth exposes full rclone control — including deleting remote data — to anyone on the LAN (accepted per the design choice). The container has only a LAN NIC, so it is not publicly reachable. Harden later by adding --rc-user/--rc-pass or fronting it with Authentik.

Tracked config (deferred)

Not yet tracked. The runner, systemd units, and folders.list currently live as plain files directly on the LXC — fully functional, just not version-controlled or auto-deployed. A dtoro/rclone gitea repo package (runner, units, install.sh, webhook receiver) is pre-built and staged at /root/rclone-repo on the LXC for whenever this gets tracked (Shape A, like caddy). Gitea ALLOWED_HOST_LIST already includes 192.168.8.214 in anticipation. See auto-deploy.

Selected folders (live in /etc/rclone-backup/folders.list): /mnt/library/cloud (287G), /mnt/library/documents (249M), /mnt/library/repos (83M). /mnt/library/notes was synced once as a connectivity test (not in the recurring set). Proton quota checked: 2 TiB plan, ~1.65 TiB free after this set.

Enrollment gotcha: pct exec PATH

pct exec (lxc-attach) does not source /etc/environment or run a login shell, so /usr/local/bin (where bootstrap installs sops) isn't on $PATH by default — bootstrap's own command -v sops post-install check failed under pct exec even though the binary installed fine. Fixed by symlinking /usr/local/bin/{sops,homelab} into /usr/bin (always on the minimal PATH), rather than relying on /etc/environment. Same category as the documented pct exec no-initgroups gotcha — worth adding to agent-enrollment.md troubleshooting if it recurs on future LXC bootstraps.

Known issue: silent protondrive stalls + watchdog

The protondrive backend has been observed (twice in the first 24h) to hang a transfer at 100% (or mid-percentage) with zero bytes, zero errors, zero retries for hours — no timeout ever fires, including --timeout 5m/--contimeout 30s set via the rc _config (confirmed applied via options/get, made no difference). Signature: core/stats transferring list shows the exact same byte counts across repeated polls; journalctl -u rclone-rcd.service goes completely silent (no new lines at all) once it happens. The only reliable fix found is killing and restarting (systemctl restart rclone-backup.service) — rclone sync is idempotent on resume, so already- uploaded bytes aren't re-transferred.

rclone-backup-watchdog.timer (every 5 min) → rclone-backup-watchdog.sh: if rclone-backup.service is active but total transferred bytes (global core/stats on the rc API) haven't moved for 15 minutes, it restarts the service automatically. State kept in /var/lib/rclone-backup/watchdog-state.json, cleared whenever the service isn't running so a stale timestamp doesn't cause a false trigger next time it starts.

Runner monitoring bug (fixed 2026-07-03): the runner's per-job progress polling queried core/stats under group key job/<jobid>, but rclone actually tracks stats under whatever _group name the job was submitted with. This silently returned all-zero stats the entire time, making a perfectly healthy sync look stalled at 0 bytes for 22+ hours in the per-run log — the real progress was only visible via core/stats with no group filter (global) or a job/list+job/status cross- check. Fixed by using the same group variable throughout. Lesson: distrust the per-run log's "progress bytes=" line at a glance during this incident window; cross-check with unfiltered core/stats before concluding a stall is real.

Changelog

2026-07-03 — two silent protondrive stalls hit; watchdog added; monitoring bug fixed

The cloud sync stalled silently twice in its first ~24h (see "Known issue" above) — once pre-timeout-fix (~6.5h with zero progress before being caught), once post-fix (~22h, but that second one turned out to be partly a false alarm: a bug in the runner's stats-group key made a healthy, actively-transferring sync (real progress 48.6G → 62.4G confirmed via unfiltered core/stats) look completely frozen in its own log. Fixed the group-key bug, then caught and confirmed a second, genuine stall (zero rcd log activity for 15+ min, frozen byte counts) and restarted again. Added rclone-backup-watchdog.timer/.service (5-min interval, 15-min stall threshold) so future stalls auto-recover without manual intervention. Total transferred as of this entry: ~62.4 GB of the ~145 GB selected set.

2026-07-02 — runner rewritten to submit jobs via the rc API (GUI job visibility)

The original runner (rclone sync invoked as a standalone CLI subprocess) was invisible to the Web GUI's Jobs panel — the GUI only tracks work submitted through its own rcd process. Rewrote /usr/local/sbin/rclone-backup.sh in Python, submitting each folder via POST /sync/sync with _async: true against http://127.0.0.1:5572 (the running rclone-rcd.service), then polling POST /job/status + POST /core/stats (both must be POST — GET-with-querystring 404s on this rc API) until finished, logging periodic progress snapshots and the same runs.jsonl summary line as before. Verified live: submitted job visible in POST /job/list's runningIds while running, completed cleanly (success: true) once done. Deployed via atomic rename (write-then-mv) rather than truncating in place, specifically so it wouldn't risk corrupting the still-running original cloud+documents+repos sync mid-flight (verified after the fact: that sync's bash process was unaffected, kept running to completion under the old in-memory script content). The already-running scheduled sync from before this change is a standalone process and won't retroactively appear in the GUI; every run after this point will.

2026-07-02 — Proton Drive auth fixed; real folder set enabled; first live sync

Initial rclone config failed 2FA (422 ... auth/v4/2fa) because a live 6-digit TOTP code was entered instead of the TOTP secret — reconfigured with the secret, auth now works (rclone lsd proton: lists the Drive). Verified end-to-end with a real sync of /mnt/library/notes (219 objects, 5.964 MiB, exit 0) — confirmed files land as plain, browsable objects on Proton (not an opaque archive), matching the plain-mirror + Proton-E2E design. Checked Proton quota (2 TiB plan, 1.945 TiB free) before enabling a large folder. folders.list set to the real selection: cloud (287G), documents (249M), repos (83M); a full sync of that set was kicked off via the actual rclone-backup.service unit (not an ad-hoc call) to validate the real monthly path early rather than waiting for the Aug 1 timer. Tracked-repo step (dtoro/rclone on gitea) deferred by choice — runner/units/folders.list remain plain files on the LXC for now; the repo package stays staged at /root/rclone-repo for later.

2026-07-01 — provisioned; enrolled

LXC 132 created (Debian 13, privileged, 192.168.8.214, /mnt/library read-only). rclone v1.74.3 installed from the official binary (protondrive backend present). Runner + monthly timer + folders.list deployed; rclone Web GUI (rcd, LAN-only no-auth) live on :5572. Enrolled into homelab-context (--no-mesh, LAN-only issuance): age key issued, inventory finalized, shared secrets granted, homelab whoami + homelab secret hello verified. Gitea ALLOWED_HOST_LIST updated to include 192.168.8.214. Hit and fixed a pct exec PATH gotcha (see below). Proton Drive remote, dtoro/rclone tracked repo + webhook, and the secrets/protondrive.yaml escrow remain operator-run follow-ups (credentialed steps — Proton password/2FA, repo creation). Restic-on-USB backup deprecated in the same change.