Files
oikos/plans/done/2026-07-08-signal-triggers.md

14 KiB
Raw Blame History

2026-07-08 — Signal triggers: host health checks

Status: Done — Phases 1-5 complete

Goal

Add a comprehensive set of signal triggers for host-level monitoring — network reachability, CPU/memory/disk pressure, thermal state, pending updates, disk health, ZFS pool status, and process liveness. Every check raises properly deduped signals through the existing UpsertSignal path and feeds the OODA pipeline (observe → classify → decide → act).


1. New check kinds

Kind What it measures Signal kind Severity mapping
ping ICMP reachability + RTT ping-unreachable down → critical
cpu Usage % + thermal (Linux only) cpu-pressure, cpu-thermal >90% → warning, >95% → critical; temp >85°C → critical
memory RAM usage % memory-pressure >90% → warning, >95% → critical
load Load avg / CPU count load-pressure >CPU×2 → warning, >CPU×4 → critical
swap Swap usage % swap-pressure >50% → warning, >80% → critical
disk-usage Already exists as disk — enhance with inode %, per-mountpoint disk-full >85% → warning, >95% → critical
disk-smart SMART pre-failure attributes disk-smart-fail any fail → critical
updates Pending apt updates (security, critical) updates-pending security >0 → warning, critical-reboot >0 → critical
zfs Pool health + scrub status zfs-degraded, zfs-scrub-overdue degraded → critical, scrub >30d → warning
process Process/service running process-down not running → critical
uptime Detect unexpected reboots uptime-bounce < previous → warning

2. Execution model

Two tiers based on where the check runs:

Tier A: Local (scheduler host)

ping, http, tcp, cert-expiry — run directly from the scheduler process. Already implemented for http/tcp/disk/cert-expiry. Add ping here.

Tier B: Remote via SSH (ssh-script)

cpu, memory, load, swap, disk-smart, updates, zfs, process, uptime — run on the target host via SSH. The existing ssh-script kind (defined in OpenAPI, not implemented in scheduler) is the one generic mechanism.

Why one ssh-script kind instead of separate kinds per metric

The check runs a small allowlisted shell snippet on the target host via SSH. The script outputs JSON with health, signalKind, evidence, and optional metrics (name→value map). This keeps the scheduler simple — one code path for all remote checks — while the check definition's config.script field encodes what to run.

SSH config shape

{
  "host": "192.168.30.10",
  "port": 22,
  "user": "root",
  "script": "cpu_check.sh",
  "timeout_s": 10
}

Scripts live in /opt/oikos/checks/ on each host, deployed by the Homelab sync timer alongside the AGENTS.md context. They are allowlisted — the scheduler only executes scripts whose names match ^[a-z][a-z0-9_-]+\.sh$ and that exist in the check directory.


3. Check scripts (per host, /opt/oikos/checks/)

cpu_check.sh

#!/usr/bin/env bash
set -euo pipefail

USAGE=$(top -bn1 | awk '/^%Cpu/ {print 100 - $8}')
CORES=$(nproc)
TEMP=""
if [ -f /sys/class/thermal/thermal_zone0/temp ]; then
  TEMP=$(echo "scale=1; $(cat /sys/class/thermal/thermal_zone0/temp) / 1000" | bc)
fi

echo "{\"health\":\"healthy\",\"metrics\":{\"cpu_pct\":$USAGE,\"cpu_temp\":$TEMP}}"

Signal raised by scheduler logic when cpu_pct > threshold_pct or cpu_temp > threshold_temp.

memory_check.sh

#!/usr/bin/env bash
set -euo pipefail
MEMINFO=$(awk '/MemTotal|MemAvailable|SwapTotal|SwapFree/ {printf "\"%s\":%d,", tolower($1), $2}' /proc/meminfo | sed 's/,$//')
USED_PCT=$(python3 -c "print(round((1 - ${memavailable:-0}/${memtotal:-1})*100, 1))")
echo "{\"health\":\"healthy\",\"metrics\":{\"mem_pct\":$USED_PCT}}"

Actual implementation would precompute from /proc/meminfo in bash directly (avoid python dependency).

load_check.sh

#!/usr/bin/env bash
set -euo pipefail
LOAD=$(awk '{print $1}' /proc/loadavg)
CORES=$(nproc)
echo "{\"health\":\"healthy\",\"metrics\":{\"load1\":$LOAD,\"cores\":$CORES}}"

Scheduler computes load_pct = load1 / cores and thresholds on that.

swap_check.sh

Reports swap used / swap total from /proc/meminfo.

disk_smart_check.sh

#!/usr/bin/env bash
set -euo pipefail
FAIL=0
for dev in $(lsblk -ndo NAME,TYPE | awk '$2=="disk"{print "/dev/"$1}'); do
  smartctl -H "$dev" | grep -q "PASSED" || { FAIL=1; break; }
done
[ $FAIL -eq 0 ] && echo '{"health":"healthy"}' || echo '{"health":"degraded","signalKind":"disk-smart-fail","evidence":"SMART health check failed"}'

updates_check.sh

#!/usr/bin/env bash
set -euo pipefail
apt update -qq >/dev/null 2>&1
SECURITY=$(apt list --upgradable 2>/dev/null | grep -c '\-security' || true)
REBOOT=$(test -f /var/run/reboot-required && echo 1 || echo 0)
echo "{\"health\":\"healthy\",\"metrics\":{\"security_updates\":$SECURITY,\"reboot_required\":$REBOOT}}"

Scheduler thresholds: security_updates > 0 → warning signal, reboot_required > 0 → critical.

zfs_check.sh

#!/usr/bin/env bash
set -euo pipefail
STATUS=$(zpool status -x 2>&1)
if echo "$STATUS" | grep -q "all pools are healthy"; then
  echo '{"health":"healthy"}'
else
  echo "{\"health\":\"degraded\",\"signalKind\":\"zfs-degraded\",\"evidence\":\"$(echo $STATUS | head -1)\"}"
fi

process_check.sh

Runs systemctl is-active <service> for a service name passed in config. Signal on inactive/failed.

uptime_check.sh

Reads current uptime, compares to last stored value (in a temp file or via metric). Signal if uptime went backward (reboot detected).


4. Scheduler changes

internal/scheduler/scheduler.go

Add two new cases to executeCheck():

case "ping":
    return checkPing(ctx, cd)
case "ssh-script":
    return checkSSHScript(ctx, cd)

Additional changes:

  • Metric recording: stop hardcoding probe_latency_ms. Each check returns a map[string]float64 of metrics, and the scheduler writes all of them. Change executeCheck signature from (health, signalKind, evidence, err) to also return metrics map[string]float64.
  • Threshold evaluation for metric-based checks (cpu, memory, load, swap, updates): the scheduler reads config.thresholds from check config JSONB and evaluates metrics against them.

checkPing()

Uses golang.org/x/net/icmp + golang.org/x/net/ipv4 for non-privileged ICMP echo (or falls back to net.DialTimeout("ip4:icmp", ...)). On macOS, ping -c 1 -t <timeout> via exec as fallback (ICMP raw sockets require root on macOS). Returns probe_latency_ms metric.

Config:

{
  "host": "192.168.30.10",
  "count": 1,
  "timeout_s": 5
}

checkSSHScript()

Connects via golang.org/x/crypto/ssh using agent forwarding or key file. Executes the allowlisted script path, validates JSON output, returns metrics and health.

Config:

{
  "host": "192.168.30.10",
  "port": 22,
  "user": "root",
  "script": "cpu_check.sh",
  "thresholds": {
    "cpu_pct": {"warn": 90, "crit": 95},
    "cpu_temp": {"crit": 85}
  },
  "timeout_s": 10
}

Threshold config schema

Each check kind with metric-based thresholds adds a thresholds key:

{
  "thresholds": {
    "<metric_name>": {"warn": <float>, "crit": <float>}
  }
}

Missing thresholds → no signal raised; metrics still recorded.


5. DB and migration

No schema migration needed. check_defs.config is JSONB — new check kinds use it with their own config shapes. signals table handles any kind string.

One small addition: add the new signal kinds to the OpenAPI CheckKind enum and the generated code (api/openapi.yaml line 2247, internal/httpapi/gen/api.gen.go line 84).


6. Seed data: default checks per host

Add to seeds/inventory.yaml or a new seeds/checks.yaml — a set of default checks per entity type:

  • Every machine, workstation, proxmox-host, lxc gets:
    • cpu (via ssh-script)
    • memory (via ssh-script)
    • load (via ssh-script)
    • swap (via ssh-script)
    • disk-usage (local or ssh-script for remote)
    • updates (via ssh-script)
    • uptime (via ssh-script)
  • Every machine, proxmox-host, standalone-server additionally gets:
    • disk-smart (via ssh-script)
    • zfs (via ssh-script) if storage-pool edges exist
  • Every service gets:
    • process (via ssh-script, systemctl is-active)

Default intervals:

  • cpu, memory, load: 60s
  • disk-usage, swap: 300s
  • ping: 30s
  • updates: 3600s (hourly)
  • disk-smart, zfs: 86400s (daily)
  • process: 30s

7. Policy integration

New approval rule for ssh-script checks:

# ssh-script execution is read_only on the target — it only reads metrics
- entity_type: machine
  action: health-check
  risk_class: read_only
  autonomy: auto

The ping kind is read_only (no mutation). All new checks are read_only — they observe state, no action taken automatically. The response to signals (restart service, clear cache, apt upgrade) goes through the existing classification → approval → execution pipeline separately.


8. Other common important signals (per user request)

Beyond the core checks above, these are worth including:

# Trigger Why important
11 Inode exhaustion Filesystem can be "full" with free space but zero inodes (Docker overlay, mail queues). Distinct from disk-usage.
12 OOM kills `dmesg
13 Journal errors `journalctl -p err -S -1h --no-pager
14 Docker/container health docker ps --filter health=unhealthy. Catches containers in unhealthy state.
15 Caddy/nginx error rate Parse access logs for 5xx rate over last 5min. Expensive, do at 300s interval.
16 Time drift `chronyc tracking
17 Open file descriptors /proc/sys/fs/file-nr ratio used/total. >80% → warning (service exhaustion).
18 Backup freshness Check timestamp of last backup file. >schedule+grace → critical.

All of these (1118) are implemented as ssh-script checks with corresponding scripts in /opt/oikos/checks/.


9. Implementation order

Phase 1: Foundation (2-3 days)

  1. Refactor executeCheck to return metrics map. Update all existing check functions (http, tcp, disk, cert-expiry).
  2. Add ping kind — ICMP reachability on the scheduler host.
  3. Add ssh-script kind — SSH execution engine, script allowlisting, JSON output parsing.
  4. Add threshold evaluation — scheduler reads config.thresholds and compares against returned metrics to decide health/signal.

Phase 2: Host check scripts (1-2 days)

  1. Write and test each script in /opt/oikos/checks/: cpu_check.sh, memory_check.sh, load_check.sh, swap_check.sh, disk_smart_check.sh, updates_check.sh, zfs_check.sh, process_check.sh, uptime_check.sh, oom_check.sh, journal_check.sh, time_check.sh, fd_check.sh.
  2. Add tools/setup-checks.sh to auto-deploy scripts via the sync timer (same pattern as tools/setup-caveman.sh).

Phase 3: Seed data + API (1 day)

  1. Add seeds/checks.yaml with default check definitions per entity type.
  2. Add new check kinds to OpenAPI spec and regenerate Go types.
  3. Add a /api/v1/checks/defaults/{entity_type} endpoint that returns recommended checks for a given entity type (convenience for operators).

Phase 4: Policy + observability (1 day)

  1. Add read_only policy rules for ping and ssh-script actions.
  2. Add per-check-kind metric recording (not just probe_latency_ms).
  3. Wire cpu_temp, mem_pct, load_pct, etc. into query_metrics and the MCP get_trend tool.

Phase 5: Docker + service signals (1 day)

  1. docker_health_check.shdocker ps --filter health=unhealthy.
  2. caddy_error_rate.sh — parse Caddy JSON logs for 5xx.
  3. backup_freshness.sh — check last backup timestamp.

10. Dependencies

Dependency For Risk
golang.org/x/crypto/ssh SSH client in scheduler Already in go.mod (used by deployer)
golang.org/x/net/icmp + ipv4 ICMP ping New dep; macOS needs root for raw sockets → exec fallback
smartmontools on hosts disk_smart_check.sh Already installed on Proxmox; add to LXCs
zfsutils-linux on hosts zfs_check.sh Already on Proxmox; add to LXCs with ZFS
SSH key on scheduler ssh-script to all hosts Already deployed (Homelab sync SSH keys)

11. Risks and mitigations

Risk Mitigation
ssh-script is a remote exec vector Scripts are allowlisted by name (^[a-z][a-z0-9_-]+\.sh$), deployed via git (auditable), and read-only (no mutation). SSH key restricted to a dedicated oikos-check user with sudo only for systemctl is-active.
ICMP requires root on macOS Fallback to ping CLI via os/exec. Production runs on Linux (strong) where raw sockets work.
Metrics cardinality explosion Metrics are per-check-definition, not per-script-output. The script returns a fixed set of known metric names. TimescaleDB handles the volume.
Check script drift between hosts Scripts deploy via tools/setup-checks.sh in the sync timer — same mechanism that keeps AGENTS.md in sync. Checksum validation before execution.

12. Verification

  • Unit tests: each check function (checkPing, checkSSHScript) tested with mock SSH server and mock ICMP responses.
  • Integration test: deploy to strong (macOS scheduler host), define checks for hubris (Proxmox), dns (LXC), caddy (LXC), run scheduler with --check-interval 10s, verify signals appear in get_signal_history.
  • MCP smoke test: search_knowledge("signal triggers"), get_signal_history, query_metrics(metric=["cpu_pct", "mem_pct"]), get_trend.
  • Policy smoke test: preflight(service:caddy, restart) still returns correct risk class despite new check kinds in the DB.