Two silent stalls hit in LXC 132's first 24h of real traffic: rclone's own --timeout didn't catch a protondrive-specific hang (transfer at 100%, zero bytes/errors/retries for hours). Added a 5-min watchdog timer that restarts rclone-backup.service if transferred bytes are frozen for 15+ min. Also found and fixed a monitoring bug in the runner (wrong stats-group key) that made a healthy sync look falsely stalled for 22h in its own log.
188 lines
12 KiB
Markdown
188 lines
12 KiB
Markdown
# 132 — `rclone`
|
|
|
|
Off-host backup appliance. Mirrors selected `/mnt/library` folders to **Proton Drive**
|
|
with a plain `rclone sync` (monthly), and serves rclone's Web GUI on the LAN for browsing
|
|
and ad-hoc runs. **Replaces** the disabled restic-on-USB job — see [backups](../infrastructure/backups.md).
|
|
|
|
Provisioned 2026-07-01. (LXC 131 was already taken by an undocumented `teddycloud` container,
|
|
so this landed on **132**.)
|
|
|
|
## At a glance
|
|
|
|
- **Hostname:** `rclone`
|
|
- **IP:** `192.168.8.214` (static, set in PVE `net0` config — same pattern as grimmory/authentik)
|
|
- **Privilege:** privileged (root in-container = host root → reads every `/mnt/library` subtree,
|
|
incl. `homecloud/` and `documents/`, regardless of owner)
|
|
- **Resources:** 1 core / 1 GiB RAM / 8 GiB rootfs (Debian 13)
|
|
- **Mounts:** `/mnt/library` **read-only** (`mp0: /mnt/library,mp=/mnt/library,ro=1`) — a backup
|
|
job must never be able to write into the library
|
|
- **Public hostname:** none — the UI is **LAN-only, no auth** (by design)
|
|
|
|
## Service / port map
|
|
|
|
| Service | Listen | Notes |
|
|
|---------|--------|-------|
|
|
| rclone Web GUI (`rcd`) | `192.168.8.214:5572` | `rclone-rcd.service`, **`--rc-no-auth`**, LAN-only. Browse `/mnt/library` + `proton:`, run ad-hoc syncs, live job status |
|
|
| monthly mirror | — | `rclone-backup.service` + `.timer` (`OnCalendar=*-*-01 03:00`) |
|
|
|
|
## Backup design
|
|
|
|
- **Mode:** plain mirror — `rclone sync` (Proton mirrors local; deletions propagate; **no versioning**).
|
|
- **Encryption:** Proton Drive's built-in E2E only (no rclone `crypt` overlay → files stay
|
|
browsable in Proton's web UI).
|
|
- **Selected set:** `/etc/rclone-backup/folders.list` — one absolute source path per line
|
|
(`#`/blank ignored). This file *is* the picked set the monthly timer mirrors. Extensible to other
|
|
disks once bind-mounted into this LXC.
|
|
- **Path mapping:** source `S` → `proton:library-backup/<S without leading slash>`
|
|
(e.g. `/mnt/library/notes` → `proton:library-backup/mnt/library/notes`).
|
|
- **Runner:** `/usr/local/sbin/rclone-backup.sh [folder ...]` (Python, despite the `.sh` name — kept
|
|
the path stable) — no arg = every enabled line. Submits each folder as an **async job through the
|
|
rclone rc API** served by `rclone-rcd.service` (the same daemon backing the Web GUI on `:5572`),
|
|
so scheduled/ad-hoc runs show up live in the GUI's **Jobs panel**, not just in logs. Gentle on
|
|
Proton's rate limits (`Transfers=4, TPSLimit=8, FastList=true` via the rc `_config` payload). The
|
|
rc API here requires **POST for every call** including `job/status` and `core/stats` — GET with
|
|
query params 404s.
|
|
- **Logs / "past runs":** per-run logs in `/var/log/rclone-backup/<safe>-<ts>.log`; one-line
|
|
JSON summary per run appended to `/var/log/rclone-backup/runs.jsonl`.
|
|
- **Failure notify:** `OnFailure=rclone-backup-notify@%n.service` → logs to journal today;
|
|
**TODO** wire to Hermes `send_message` (Matrix) per [backups](../infrastructure/backups.md).
|
|
|
|
## rclone + Proton Drive
|
|
|
|
- **rclone** installed from the official binary (not apt) so the `protondrive` backend is present
|
|
(`rclone v1.74.3`).
|
|
- Remote **`proton:`** (type `protondrive`). Config at `/root/.config/rclone/rclone.conf`, mode 600.
|
|
**This file is a secret** (holds the obscured Proton password + TOTP secret + session) — **never
|
|
commit it.** Escrow the Proton account creds in the password manager.
|
|
- **Config gotchas** (from rclone docs/forum):
|
|
- Log into Proton via a **browser at least once** first, or key generation fails.
|
|
- For unattended runs, store the **TOTP _secret_** (not a 6-digit code) so rclone self-generates
|
|
codes; obscure with `rclone obscure`.
|
|
- Passwords with **extended-ASCII** characters are known to break auth.
|
|
- Proton's API is rate-limited → keep `--transfers`/`--tpslimit` conservative (baked into the runner).
|
|
- **DR escrow (pending):** store the Proton creds as sops secret `secrets/protondrive.yaml`, granted
|
|
to this LXC's age key, so the remote can be rebuilt after a re-provision.
|
|
|
|
## The UI (rclone Web GUI)
|
|
|
|
`rclone rcd --rc-web-gui --rc-no-auth --rc-addr 0.0.0.0:5572` (assets auto-downloaded on first
|
|
start). Reach it at **http://192.168.8.214:5572** on the LAN.
|
|
|
|
> **Security note:** `--rc-no-auth` exposes *full* rclone control — including deleting remote data —
|
|
> to anyone on the LAN (accepted per the design choice). The container has only a LAN NIC, so it is
|
|
> not publicly reachable. Harden later by adding `--rc-user/--rc-pass` or fronting it with Authentik.
|
|
|
|
## Tracked config (deferred)
|
|
|
|
**Not yet tracked.** The runner, systemd units, and `folders.list` currently live as plain files
|
|
directly on the LXC — fully functional, just not version-controlled or auto-deployed. A
|
|
`dtoro/rclone` gitea repo package (runner, units, `install.sh`, webhook receiver) is pre-built and
|
|
staged at `/root/rclone-repo` on the LXC for whenever this gets tracked (Shape A, like
|
|
[caddy](121-caddy.md)). Gitea `ALLOWED_HOST_LIST` already includes `192.168.8.214` in anticipation.
|
|
See [auto-deploy](../infrastructure/auto-deploy.md).
|
|
|
|
**Selected folders (live in `/etc/rclone-backup/folders.list`):** `/mnt/library/cloud` (287G),
|
|
`/mnt/library/documents` (249M), `/mnt/library/repos` (83M). `/mnt/library/notes` was synced once as
|
|
a connectivity test (not in the recurring set). Proton quota checked: 2 TiB plan, ~1.65 TiB free
|
|
after this set.
|
|
|
|
## Enrollment gotcha: `pct exec` PATH
|
|
|
|
`pct exec` (lxc-attach) does **not** source `/etc/environment` or run a login shell, so
|
|
`/usr/local/bin` (where bootstrap installs `sops`) isn't on `$PATH` by default — bootstrap's own
|
|
`command -v sops` post-install check failed under `pct exec` even though the binary installed fine.
|
|
Fixed by symlinking `/usr/local/bin/{sops,homelab}` into `/usr/bin` (always on the minimal PATH),
|
|
rather than relying on `/etc/environment`. Same category as the documented [`pct exec` no-initgroups
|
|
gotcha](../infrastructure/media-permissions.md#gotchas) — worth adding to
|
|
[agent-enrollment.md troubleshooting](../operations/agent-enrollment.md#troubleshooting) if it recurs
|
|
on future LXC bootstraps.
|
|
|
|
## Known issue: silent protondrive stalls + watchdog
|
|
|
|
The protondrive backend has been observed (twice in the first 24h) to hang a transfer at 100% (or
|
|
mid-percentage) with **zero bytes, zero errors, zero retries** for hours — no timeout ever fires,
|
|
including `--timeout 5m`/`--contimeout 30s` set via the rc `_config` (confirmed applied via
|
|
`options/get`, made no difference). Signature: `core/stats` `transferring` list shows the exact same
|
|
byte counts across repeated polls; `journalctl -u rclone-rcd.service` goes completely silent (no
|
|
new lines at all) once it happens. The only reliable fix found is **killing and restarting**
|
|
(`systemctl restart rclone-backup.service`) — rclone sync is idempotent on resume, so already-
|
|
uploaded bytes aren't re-transferred.
|
|
|
|
**`rclone-backup-watchdog.timer`** (every 5 min) → `rclone-backup-watchdog.sh`: if
|
|
`rclone-backup.service` is active but total transferred bytes (global `core/stats` on the rc API)
|
|
haven't moved for 15 minutes, it restarts the service automatically. State kept in
|
|
`/var/lib/rclone-backup/watchdog-state.json`, cleared whenever the service isn't running so a stale
|
|
timestamp doesn't cause a false trigger next time it starts.
|
|
|
|
**Runner monitoring bug (fixed 2026-07-03):** the runner's per-job progress polling queried
|
|
`core/stats` under group key `job/<jobid>`, but rclone actually tracks stats under whatever `_group`
|
|
name the job was submitted with. This silently returned all-zero stats the entire time, making a
|
|
perfectly healthy sync look stalled at 0 bytes for 22+ hours in the per-run log — the real progress
|
|
was only visible via `core/stats` with no group filter (global) or a `job/list`+`job/status` cross-
|
|
check. Fixed by using the same `group` variable throughout. **Lesson: distrust the per-run log's
|
|
"progress bytes=" line at a glance during this incident window; cross-check with unfiltered
|
|
`core/stats` before concluding a stall is real.**
|
|
|
|
## Related
|
|
|
|
- [Backups](../infrastructure/backups.md) — this job supersedes the disabled restic-on-USB backup
|
|
- [Hubris host](../hosts/hubris.md) — owns `/mnt/library`
|
|
- [Media permissions](../infrastructure/media-permissions.md) — read-only consumer of `/mnt/library`
|
|
- [Containers index](index.md)
|
|
|
|
## Changelog
|
|
|
|
### 2026-07-03 — two silent protondrive stalls hit; watchdog added; monitoring bug fixed
|
|
|
|
The `cloud` sync stalled silently twice in its first ~24h (see "Known issue" above) — once
|
|
pre-timeout-fix (~6.5h with zero progress before being caught), once post-fix (~22h, but that
|
|
second one turned out to be **partly a false alarm**: a bug in the runner's stats-group key made a
|
|
healthy, actively-transferring sync (real progress 48.6G → 62.4G confirmed via unfiltered
|
|
`core/stats`) look completely frozen in its own log. Fixed the group-key bug, then caught and
|
|
confirmed a **second, genuine** stall (zero rcd log activity for 15+ min, frozen byte counts) and
|
|
restarted again. Added `rclone-backup-watchdog.timer`/`.service` (5-min interval, 15-min stall
|
|
threshold) so future stalls auto-recover without manual intervention. Total transferred as of this
|
|
entry: ~62.4 GB of the ~145 GB selected set.
|
|
|
|
### 2026-07-02 — runner rewritten to submit jobs via the rc API (GUI job visibility)
|
|
|
|
The original runner (`rclone sync` invoked as a standalone CLI subprocess) was invisible to the Web
|
|
GUI's Jobs panel — the GUI only tracks work submitted through its own `rcd` process. Rewrote
|
|
`/usr/local/sbin/rclone-backup.sh` in Python, submitting each folder via `POST /sync/sync` with
|
|
`_async: true` against `http://127.0.0.1:5572` (the running `rclone-rcd.service`), then polling
|
|
`POST /job/status` + `POST /core/stats` (both **must be POST** — GET-with-querystring 404s on this
|
|
rc API) until finished, logging periodic progress snapshots and the same `runs.jsonl` summary line
|
|
as before. Verified live: submitted job visible in `POST /job/list`'s `runningIds` while running,
|
|
completed cleanly (`success: true`) once done. Deployed via atomic rename (write-then-`mv`) rather
|
|
than truncating in place, specifically so it wouldn't risk corrupting the still-running original
|
|
`cloud`+`documents`+`repos` sync mid-flight (verified after the fact: that sync's bash process was
|
|
unaffected, kept running to completion under the old in-memory script content). The already-running
|
|
scheduled sync from before this change is a standalone process and won't retroactively appear in the
|
|
GUI; every run after this point will.
|
|
|
|
### 2026-07-02 — Proton Drive auth fixed; real folder set enabled; first live sync
|
|
|
|
Initial `rclone config` failed 2FA (`422 ... auth/v4/2fa`) because a live 6-digit TOTP code was
|
|
entered instead of the TOTP secret — reconfigured with the secret, auth now works
|
|
(`rclone lsd proton:` lists the Drive). Verified end-to-end with a real sync of `/mnt/library/notes`
|
|
(219 objects, 5.964 MiB, exit 0) — confirmed files land as plain, browsable objects on Proton (not
|
|
an opaque archive), matching the plain-mirror + Proton-E2E design. Checked Proton quota (2 TiB
|
|
plan, 1.945 TiB free) before enabling a large folder. `folders.list` set to the real selection:
|
|
`cloud` (287G), `documents` (249M), `repos` (83M); a full sync of that set was kicked off via the
|
|
actual `rclone-backup.service` unit (not an ad-hoc call) to validate the real monthly path early
|
|
rather than waiting for the Aug 1 timer. Tracked-repo step (`dtoro/rclone` on gitea) deferred by
|
|
choice — runner/units/`folders.list` remain plain files on the LXC for now; the repo package stays
|
|
staged at `/root/rclone-repo` for later.
|
|
|
|
### 2026-07-01 — provisioned; enrolled
|
|
|
|
LXC 132 created (Debian 13, privileged, `192.168.8.214`, `/mnt/library` read-only). rclone v1.74.3
|
|
installed from the official binary (`protondrive` backend present). Runner + monthly timer +
|
|
`folders.list` deployed; rclone Web GUI (`rcd`, LAN-only no-auth) live on `:5572`. Enrolled into
|
|
homelab-context (`--no-mesh`, LAN-only issuance): age key issued, inventory finalized, shared
|
|
secrets granted, `homelab whoami` + `homelab secret hello` verified. Gitea `ALLOWED_HOST_LIST`
|
|
updated to include `192.168.8.214`. Hit and fixed a `pct exec` PATH gotcha (see below). Proton Drive
|
|
remote, `dtoro/rclone` tracked repo + webhook, and the `secrets/protondrive.yaml` escrow remain
|
|
operator-run follow-ups (credentialed steps — Proton password/2FA, repo creation). Restic-on-USB
|
|
backup deprecated in the same change.
|