Files
oikos/plans/done/2026-07-08-signal-triggers.md

388 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 2026-07-08 — Signal triggers: host health checks
**Status:** Done — Phases 1-5 complete
## Goal
Add a comprehensive set of signal triggers for host-level monitoring —
network reachability, CPU/memory/disk pressure, thermal state, pending
updates, disk health, ZFS pool status, and process liveness. Every check
raises properly deduped signals through the existing `UpsertSignal` path
and feeds the OODA pipeline (observe → classify → decide → act).
---
## 1. New check kinds
| Kind | What it measures | Signal kind | Severity mapping |
|------|-----------------|-------------|------------------|
| `ping` | ICMP reachability + RTT | `ping-unreachable` | down → critical |
| `cpu` | Usage % + thermal (Linux only) | `cpu-pressure`, `cpu-thermal` | >90% → warning, >95% → critical; temp >85°C → critical |
| `memory` | RAM usage % | `memory-pressure` | >90% → warning, >95% → critical |
| `load` | Load avg / CPU count | `load-pressure` | >CPU×2 → warning, >CPU×4 → critical |
| `swap` | Swap usage % | `swap-pressure` | >50% → warning, >80% → critical |
| `disk-usage` | Already exists as `disk` — enhance with inode %, per-mountpoint | `disk-full` | >85% → warning, >95% → critical |
| `disk-smart` | SMART pre-failure attributes | `disk-smart-fail` | any fail → critical |
| `updates` | Pending apt updates (security, critical) | `updates-pending` | security >0 → warning, critical-reboot >0 → critical |
| `zfs` | Pool health + scrub status | `zfs-degraded`, `zfs-scrub-overdue` | degraded → critical, scrub >30d → warning |
| `process` | Process/service running | `process-down` | not running → critical |
| `uptime` | Detect unexpected reboots | `uptime-bounce` | < previous → warning |
---
## 2. Execution model
Two tiers based on where the check runs:
### Tier A: Local (scheduler host)
`ping`, `http`, `tcp`, `cert-expiry` — run directly from the scheduler process. Already implemented for `http`/`tcp`/`disk`/`cert-expiry`. Add `ping` here.
### Tier B: Remote via SSH (`ssh-script`)
`cpu`, `memory`, `load`, `swap`, `disk-smart`, `updates`, `zfs`, `process`, `uptime` — run on the target host via SSH. The existing `ssh-script` kind (defined in OpenAPI, not implemented in scheduler) is the one generic mechanism.
### Why one `ssh-script` kind instead of separate kinds per metric
The check runs a small allowlisted shell snippet on the target host via SSH.
The script outputs JSON with `health`, `signalKind`, `evidence`, and optional
`metrics` (name→value map). This keeps the scheduler simple — one code path for
all remote checks — while the check definition's `config.script` field encodes
what to run.
### SSH config shape
```json
{
"host": "192.168.30.10",
"port": 22,
"user": "root",
"script": "cpu_check.sh",
"timeout_s": 10
}
```
Scripts live in `/opt/oikos/checks/` on each host, deployed by the Homelab sync
timer alongside the AGENTS.md context. They are allowlisted — the scheduler only
executes scripts whose names match `^[a-z][a-z0-9_-]+\.sh$` and that exist in the
check directory.
---
## 3. Check scripts (per host, `/opt/oikos/checks/`)
### `cpu_check.sh`
```sh
#!/usr/bin/env bash
set -euo pipefail
USAGE=$(top -bn1 | awk '/^%Cpu/ {print 100 - $8}')
CORES=$(nproc)
TEMP=""
if [ -f /sys/class/thermal/thermal_zone0/temp ]; then
TEMP=$(echo "scale=1; $(cat /sys/class/thermal/thermal_zone0/temp) / 1000" | bc)
fi
echo "{\"health\":\"healthy\",\"metrics\":{\"cpu_pct\":$USAGE,\"cpu_temp\":$TEMP}}"
```
Signal raised by scheduler logic when `cpu_pct > threshold_pct` or `cpu_temp > threshold_temp`.
### `memory_check.sh`
```sh
#!/usr/bin/env bash
set -euo pipefail
MEMINFO=$(awk '/MemTotal|MemAvailable|SwapTotal|SwapFree/ {printf "\"%s\":%d,", tolower($1), $2}' /proc/meminfo | sed 's/,$//')
USED_PCT=$(python3 -c "print(round((1 - ${memavailable:-0}/${memtotal:-1})*100, 1))")
echo "{\"health\":\"healthy\",\"metrics\":{\"mem_pct\":$USED_PCT}}"
```
Actual implementation would precompute from `/proc/meminfo` in bash directly (avoid python dependency).
### `load_check.sh`
```sh
#!/usr/bin/env bash
set -euo pipefail
LOAD=$(awk '{print $1}' /proc/loadavg)
CORES=$(nproc)
echo "{\"health\":\"healthy\",\"metrics\":{\"load1\":$LOAD,\"cores\":$CORES}}"
```
Scheduler computes `load_pct = load1 / cores` and thresholds on that.
### `swap_check.sh`
Reports swap used / swap total from `/proc/meminfo`.
### `disk_smart_check.sh`
```sh
#!/usr/bin/env bash
set -euo pipefail
FAIL=0
for dev in $(lsblk -ndo NAME,TYPE | awk '$2=="disk"{print "/dev/"$1}'); do
smartctl -H "$dev" | grep -q "PASSED" || { FAIL=1; break; }
done
[ $FAIL -eq 0 ] && echo '{"health":"healthy"}' || echo '{"health":"degraded","signalKind":"disk-smart-fail","evidence":"SMART health check failed"}'
```
### `updates_check.sh`
```sh
#!/usr/bin/env bash
set -euo pipefail
apt update -qq >/dev/null 2>&1
SECURITY=$(apt list --upgradable 2>/dev/null | grep -c '\-security' || true)
REBOOT=$(test -f /var/run/reboot-required && echo 1 || echo 0)
echo "{\"health\":\"healthy\",\"metrics\":{\"security_updates\":$SECURITY,\"reboot_required\":$REBOOT}}"
```
Scheduler thresholds: `security_updates > 0` → warning signal, `reboot_required > 0` → critical.
### `zfs_check.sh`
```sh
#!/usr/bin/env bash
set -euo pipefail
STATUS=$(zpool status -x 2>&1)
if echo "$STATUS" | grep -q "all pools are healthy"; then
echo '{"health":"healthy"}'
else
echo "{\"health\":\"degraded\",\"signalKind\":\"zfs-degraded\",\"evidence\":\"$(echo $STATUS | head -1)\"}"
fi
```
### `process_check.sh`
Runs `systemctl is-active <service>` for a service name passed in config.
Signal on `inactive`/`failed`.
### `uptime_check.sh`
Reads current uptime, compares to last stored value (in a temp file or via metric).
Signal if uptime went backward (reboot detected).
---
## 4. Scheduler changes
### `internal/scheduler/scheduler.go`
Add two new cases to `executeCheck()`:
```go
case "ping":
return checkPing(ctx, cd)
case "ssh-script":
return checkSSHScript(ctx, cd)
```
Additional changes:
- **Metric recording**: stop hardcoding `probe_latency_ms`. Each check returns
a `map[string]float64` of metrics, and the scheduler writes all of them.
Change `executeCheck` signature from `(health, signalKind, evidence, err)`
to also return `metrics map[string]float64`.
- **Threshold evaluation for metric-based checks** (`cpu`, `memory`, `load`,
`swap`, `updates`): the scheduler reads `config.thresholds` from check config
JSONB and evaluates metrics against them.
### `checkPing()`
Uses `golang.org/x/net/icmp` + `golang.org/x/net/ipv4` for non-privileged
ICMP echo (or falls back to `net.DialTimeout("ip4:icmp", ...)`).
On macOS, `ping -c 1 -t <timeout>` via exec as fallback (ICMP raw sockets
require root on macOS). Returns `probe_latency_ms` metric.
Config:
```json
{
"host": "192.168.30.10",
"count": 1,
"timeout_s": 5
}
```
### `checkSSHScript()`
Connects via `golang.org/x/crypto/ssh` using agent forwarding or key file.
Executes the allowlisted script path, validates JSON output, returns metrics
and health.
Config:
```json
{
"host": "192.168.30.10",
"port": 22,
"user": "root",
"script": "cpu_check.sh",
"thresholds": {
"cpu_pct": {"warn": 90, "crit": 95},
"cpu_temp": {"crit": 85}
},
"timeout_s": 10
}
```
### Threshold config schema
Each check kind with metric-based thresholds adds a `thresholds` key:
```json
{
"thresholds": {
"<metric_name>": {"warn": <float>, "crit": <float>}
}
}
```
Missing thresholds → no signal raised; metrics still recorded.
---
## 5. DB and migration
No schema migration needed. `check_defs.config` is JSONB — new check kinds
use it with their own config shapes. `signals` table handles any `kind` string.
One small addition: add the new signal kinds to the OpenAPI `CheckKind` enum
and the generated code (`api/openapi.yaml` line 2247, `internal/httpapi/gen/api.gen.go` line 84).
---
## 6. Seed data: default checks per host
Add to `seeds/inventory.yaml` or a new `seeds/checks.yaml` — a set of default
checks per entity type:
- Every `machine`, `workstation`, `proxmox-host`, `lxc` gets:
- `cpu` (via ssh-script)
- `memory` (via ssh-script)
- `load` (via ssh-script)
- `swap` (via ssh-script)
- `disk-usage` (local or ssh-script for remote)
- `updates` (via ssh-script)
- `uptime` (via ssh-script)
- Every `machine`, `proxmox-host`, `standalone-server` additionally gets:
- `disk-smart` (via ssh-script)
- `zfs` (via ssh-script) if `storage-pool` edges exist
- Every `service` gets:
- `process` (via ssh-script, `systemctl is-active`)
Default intervals:
- `cpu`, `memory`, `load`: 60s
- `disk-usage`, `swap`: 300s
- `ping`: 30s
- `updates`: 3600s (hourly)
- `disk-smart`, `zfs`: 86400s (daily)
- `process`: 30s
---
## 7. Policy integration
New approval rule for `ssh-script` checks:
```yaml
# ssh-script execution is read_only on the target — it only reads metrics
- entity_type: machine
action: health-check
risk_class: read_only
autonomy: auto
```
The `ping` kind is `read_only` (no mutation). All new checks are `read_only`
— they observe state, no action taken automatically. The *response* to signals
(restart service, clear cache, apt upgrade) goes through the existing
classification → approval → execution pipeline separately.
---
## 8. Other common important signals (per user request)
Beyond the core checks above, these are worth including:
| # | Trigger | Why important |
|---|---------|---------------|
| 11 | **Inode exhaustion** | Filesystem can be "full" with free space but zero inodes (Docker overlay, mail queues). Distinct from disk-usage. |
| 12 | **OOM kills** | `dmesg | grep -i 'out of memory'` count since last boot. Indicates memory pressure beyond usage %. |
| 13 | **Journal errors** | `journalctl -p err -S -1h --no-pager | wc -l`. Catches kernel panics, segfaults, service failures. |
| 14 | **Docker/container health** | `docker ps --filter health=unhealthy`. Catches containers in unhealthy state. |
| 15 | **Caddy/nginx error rate** | Parse access logs for 5xx rate over last 5min. Expensive, do at 300s interval. |
| 16 | **Time drift** | `chronyc tracking | grep 'System time'`. NTP offset > 1s → warning (affects TLS, auth, DB). |
| 17 | **Open file descriptors** | `/proc/sys/fs/file-nr` ratio used/total. >80% → warning (service exhaustion). |
| 18 | **Backup freshness** | Check timestamp of last backup file. >schedule+grace → critical. |
All of these (1118) are implemented as `ssh-script` checks with
corresponding scripts in `/opt/oikos/checks/`.
---
## 9. Implementation order
### Phase 1: Foundation (2-3 days)
1. **Refactor `executeCheck`** to return metrics map. Update all existing check
functions (`http`, `tcp`, `disk`, `cert-expiry`).
2. **Add `ping` kind** — ICMP reachability on the scheduler host.
3. **Add `ssh-script` kind** — SSH execution engine, script allowlisting,
JSON output parsing.
4. **Add threshold evaluation** — scheduler reads `config.thresholds` and
compares against returned metrics to decide health/signal.
### Phase 2: Host check scripts (1-2 days)
5. Write and test each script in `/opt/oikos/checks/`:
`cpu_check.sh`, `memory_check.sh`, `load_check.sh`, `swap_check.sh`,
`disk_smart_check.sh`, `updates_check.sh`, `zfs_check.sh`,
`process_check.sh`, `uptime_check.sh`, `oom_check.sh`, `journal_check.sh`,
`time_check.sh`, `fd_check.sh`.
6. Add `tools/setup-checks.sh` to auto-deploy scripts via the sync timer
(same pattern as `tools/setup-caveman.sh`).
### Phase 3: Seed data + API (1 day)
7. Add `seeds/checks.yaml` with default check definitions per entity type.
8. Add new check kinds to OpenAPI spec and regenerate Go types.
9. Add a `/api/v1/checks/defaults/{entity_type}` endpoint that returns
recommended checks for a given entity type (convenience for operators).
### Phase 4: Policy + observability (1 day)
10. Add `read_only` policy rules for `ping` and `ssh-script` actions.
11. Add per-check-kind metric recording (not just `probe_latency_ms`).
12. Wire `cpu_temp`, `mem_pct`, `load_pct`, etc. into `query_metrics` and
the MCP `get_trend` tool.
### Phase 5: Docker + service signals (1 day)
13. `docker_health_check.sh``docker ps --filter health=unhealthy`.
14. `caddy_error_rate.sh` — parse Caddy JSON logs for 5xx.
15. `backup_freshness.sh` — check last backup timestamp.
---
## 10. Dependencies
| Dependency | For | Risk |
|------------|-----|------|
| `golang.org/x/crypto/ssh` | SSH client in scheduler | Already in `go.mod` (used by deployer) |
| `golang.org/x/net/icmp` + `ipv4` | ICMP ping | New dep; macOS needs root for raw sockets → exec fallback |
| `smartmontools` on hosts | `disk_smart_check.sh` | Already installed on Proxmox; add to LXCs |
| `zfsutils-linux` on hosts | `zfs_check.sh` | Already on Proxmox; add to LXCs with ZFS |
| SSH key on scheduler | `ssh-script` to all hosts | Already deployed (Homelab sync SSH keys) |
---
## 11. Risks and mitigations
| Risk | Mitigation |
|------|------------|
| `ssh-script` is a remote exec vector | Scripts are allowlisted by name (`^[a-z][a-z0-9_-]+\.sh$`), deployed via git (auditable), and read-only (no mutation). SSH key restricted to a dedicated `oikos-check` user with sudo only for `systemctl is-active`. |
| ICMP requires root on macOS | Fallback to `ping` CLI via `os/exec`. Production runs on Linux (strong) where raw sockets work. |
| Metrics cardinality explosion | Metrics are per-check-definition, not per-script-output. The script returns a fixed set of known metric names. TimescaleDB handles the volume. |
| Check script drift between hosts | Scripts deploy via `tools/setup-checks.sh` in the sync timer — same mechanism that keeps AGENTS.md in sync. Checksum validation before execution. |
---
## 12. Verification
- **Unit tests**: each check function (`checkPing`, `checkSSHScript`) tested
with mock SSH server and mock ICMP responses.
- **Integration test**: deploy to `strong` (macOS scheduler host), define
checks for `hubris` (Proxmox), `dns` (LXC), `caddy` (LXC), run scheduler
with `--check-interval 10s`, verify signals appear in `get_signal_history`.
- **MCP smoke test**: `search_knowledge("signal triggers")`, `get_signal_history`,
`query_metrics(metric=["cpu_pct", "mem_pct"])`, `get_trend`.
- **Policy smoke test**: `preflight(service:caddy, restart)` still returns
correct risk class despite new check kinds in the DB.