- Refactor executeCheck to return checkResult struct with metrics map - Add ping check kind (ICMP reachability via system ping, macOS+Linux) - Add ssh-script check kind (remote host exec via SSH, allowlisted scripts) - Add threshold evaluation (warn/crit per metric from check config JSONB) - Add inode tracking to disk check - All 4 existing checks now return structured metrics - 17 check scripts: cpu, memory, load, swap, disk_usage, disk_smart, updates, zfs, process, uptime, oom, journal, time, fd, docker_health, caddy_error_rate, backup_freshness - Auto-deploy via tools/setup-checks.sh -> checks/install.sh on git pull - Add ping to OpenAPI CheckKind enum and generated Go types
14 KiB
2026-07-08 — Signal triggers: host health checks
Status: Implemented (Phases 1-5 complete)
Goal
Add a comprehensive set of signal triggers for host-level monitoring —
network reachability, CPU/memory/disk pressure, thermal state, pending
updates, disk health, ZFS pool status, and process liveness. Every check
raises properly deduped signals through the existing UpsertSignal path
and feeds the OODA pipeline (observe → classify → decide → act).
1. New check kinds
| Kind | What it measures | Signal kind | Severity mapping |
|---|---|---|---|
ping |
ICMP reachability + RTT | ping-unreachable |
down → critical |
cpu |
Usage % + thermal (Linux only) | cpu-pressure, cpu-thermal |
>90% → warning, >95% → critical; temp >85°C → critical |
memory |
RAM usage % | memory-pressure |
>90% → warning, >95% → critical |
load |
Load avg / CPU count | load-pressure |
>CPU×2 → warning, >CPU×4 → critical |
swap |
Swap usage % | swap-pressure |
>50% → warning, >80% → critical |
disk-usage |
Already exists as disk — enhance with inode %, per-mountpoint |
disk-full |
>85% → warning, >95% → critical |
disk-smart |
SMART pre-failure attributes | disk-smart-fail |
any fail → critical |
updates |
Pending apt updates (security, critical) | updates-pending |
security >0 → warning, critical-reboot >0 → critical |
zfs |
Pool health + scrub status | zfs-degraded, zfs-scrub-overdue |
degraded → critical, scrub >30d → warning |
process |
Process/service running | process-down |
not running → critical |
uptime |
Detect unexpected reboots | uptime-bounce |
< previous → warning |
2. Execution model
Two tiers based on where the check runs:
Tier A: Local (scheduler host)
ping, http, tcp, cert-expiry — run directly from the scheduler process. Already implemented for http/tcp/disk/cert-expiry. Add ping here.
Tier B: Remote via SSH (ssh-script)
cpu, memory, load, swap, disk-smart, updates, zfs, process, uptime — run on the target host via SSH. The existing ssh-script kind (defined in OpenAPI, not implemented in scheduler) is the one generic mechanism.
Why one ssh-script kind instead of separate kinds per metric
The check runs a small allowlisted shell snippet on the target host via SSH.
The script outputs JSON with health, signalKind, evidence, and optional
metrics (name→value map). This keeps the scheduler simple — one code path for
all remote checks — while the check definition's config.script field encodes
what to run.
SSH config shape
{
"host": "192.168.30.10",
"port": 22,
"user": "root",
"script": "cpu_check.sh",
"timeout_s": 10
}
Scripts live in /opt/oikos/checks/ on each host, deployed by the Homelab sync
timer alongside the AGENTS.md context. They are allowlisted — the scheduler only
executes scripts whose names match ^[a-z][a-z0-9_-]+\.sh$ and that exist in the
check directory.
3. Check scripts (per host, /opt/oikos/checks/)
cpu_check.sh
#!/usr/bin/env bash
set -euo pipefail
USAGE=$(top -bn1 | awk '/^%Cpu/ {print 100 - $8}')
CORES=$(nproc)
TEMP=""
if [ -f /sys/class/thermal/thermal_zone0/temp ]; then
TEMP=$(echo "scale=1; $(cat /sys/class/thermal/thermal_zone0/temp) / 1000" | bc)
fi
echo "{\"health\":\"healthy\",\"metrics\":{\"cpu_pct\":$USAGE,\"cpu_temp\":$TEMP}}"
Signal raised by scheduler logic when cpu_pct > threshold_pct or cpu_temp > threshold_temp.
memory_check.sh
#!/usr/bin/env bash
set -euo pipefail
MEMINFO=$(awk '/MemTotal|MemAvailable|SwapTotal|SwapFree/ {printf "\"%s\":%d,", tolower($1), $2}' /proc/meminfo | sed 's/,$//')
USED_PCT=$(python3 -c "print(round((1 - ${memavailable:-0}/${memtotal:-1})*100, 1))")
echo "{\"health\":\"healthy\",\"metrics\":{\"mem_pct\":$USED_PCT}}"
Actual implementation would precompute from /proc/meminfo in bash directly (avoid python dependency).
load_check.sh
#!/usr/bin/env bash
set -euo pipefail
LOAD=$(awk '{print $1}' /proc/loadavg)
CORES=$(nproc)
echo "{\"health\":\"healthy\",\"metrics\":{\"load1\":$LOAD,\"cores\":$CORES}}"
Scheduler computes load_pct = load1 / cores and thresholds on that.
swap_check.sh
Reports swap used / swap total from /proc/meminfo.
disk_smart_check.sh
#!/usr/bin/env bash
set -euo pipefail
FAIL=0
for dev in $(lsblk -ndo NAME,TYPE | awk '$2=="disk"{print "/dev/"$1}'); do
smartctl -H "$dev" | grep -q "PASSED" || { FAIL=1; break; }
done
[ $FAIL -eq 0 ] && echo '{"health":"healthy"}' || echo '{"health":"degraded","signalKind":"disk-smart-fail","evidence":"SMART health check failed"}'
updates_check.sh
#!/usr/bin/env bash
set -euo pipefail
apt update -qq >/dev/null 2>&1
SECURITY=$(apt list --upgradable 2>/dev/null | grep -c '\-security' || true)
REBOOT=$(test -f /var/run/reboot-required && echo 1 || echo 0)
echo "{\"health\":\"healthy\",\"metrics\":{\"security_updates\":$SECURITY,\"reboot_required\":$REBOOT}}"
Scheduler thresholds: security_updates > 0 → warning signal, reboot_required > 0 → critical.
zfs_check.sh
#!/usr/bin/env bash
set -euo pipefail
STATUS=$(zpool status -x 2>&1)
if echo "$STATUS" | grep -q "all pools are healthy"; then
echo '{"health":"healthy"}'
else
echo "{\"health\":\"degraded\",\"signalKind\":\"zfs-degraded\",\"evidence\":\"$(echo $STATUS | head -1)\"}"
fi
process_check.sh
Runs systemctl is-active <service> for a service name passed in config.
Signal on inactive/failed.
uptime_check.sh
Reads current uptime, compares to last stored value (in a temp file or via metric). Signal if uptime went backward (reboot detected).
4. Scheduler changes
internal/scheduler/scheduler.go
Add two new cases to executeCheck():
case "ping":
return checkPing(ctx, cd)
case "ssh-script":
return checkSSHScript(ctx, cd)
Additional changes:
- Metric recording: stop hardcoding
probe_latency_ms. Each check returns amap[string]float64of metrics, and the scheduler writes all of them. ChangeexecuteChecksignature from(health, signalKind, evidence, err)to also returnmetrics map[string]float64. - Threshold evaluation for metric-based checks (
cpu,memory,load,swap,updates): the scheduler readsconfig.thresholdsfrom check config JSONB and evaluates metrics against them.
checkPing()
Uses golang.org/x/net/icmp + golang.org/x/net/ipv4 for non-privileged
ICMP echo (or falls back to net.DialTimeout("ip4:icmp", ...)).
On macOS, ping -c 1 -t <timeout> via exec as fallback (ICMP raw sockets
require root on macOS). Returns probe_latency_ms metric.
Config:
{
"host": "192.168.30.10",
"count": 1,
"timeout_s": 5
}
checkSSHScript()
Connects via golang.org/x/crypto/ssh using agent forwarding or key file.
Executes the allowlisted script path, validates JSON output, returns metrics
and health.
Config:
{
"host": "192.168.30.10",
"port": 22,
"user": "root",
"script": "cpu_check.sh",
"thresholds": {
"cpu_pct": {"warn": 90, "crit": 95},
"cpu_temp": {"crit": 85}
},
"timeout_s": 10
}
Threshold config schema
Each check kind with metric-based thresholds adds a thresholds key:
{
"thresholds": {
"<metric_name>": {"warn": <float>, "crit": <float>}
}
}
Missing thresholds → no signal raised; metrics still recorded.
5. DB and migration
No schema migration needed. check_defs.config is JSONB — new check kinds
use it with their own config shapes. signals table handles any kind string.
One small addition: add the new signal kinds to the OpenAPI CheckKind enum
and the generated code (api/openapi.yaml line 2247, internal/httpapi/gen/api.gen.go line 84).
6. Seed data: default checks per host
Add to seeds/inventory.yaml or a new seeds/checks.yaml — a set of default
checks per entity type:
- Every
machine,workstation,proxmox-host,lxcgets:cpu(via ssh-script)memory(via ssh-script)load(via ssh-script)swap(via ssh-script)disk-usage(local or ssh-script for remote)updates(via ssh-script)uptime(via ssh-script)
- Every
machine,proxmox-host,standalone-serveradditionally gets:disk-smart(via ssh-script)zfs(via ssh-script) ifstorage-pooledges exist
- Every
servicegets:process(via ssh-script,systemctl is-active)
Default intervals:
cpu,memory,load: 60sdisk-usage,swap: 300sping: 30supdates: 3600s (hourly)disk-smart,zfs: 86400s (daily)process: 30s
7. Policy integration
New approval rule for ssh-script checks:
# ssh-script execution is read_only on the target — it only reads metrics
- entity_type: machine
action: health-check
risk_class: read_only
autonomy: auto
The ping kind is read_only (no mutation). All new checks are read_only
— they observe state, no action taken automatically. The response to signals
(restart service, clear cache, apt upgrade) goes through the existing
classification → approval → execution pipeline separately.
8. Other common important signals (per user request)
Beyond the core checks above, these are worth including:
| # | Trigger | Why important |
|---|---|---|
| 11 | Inode exhaustion | Filesystem can be "full" with free space but zero inodes (Docker overlay, mail queues). Distinct from disk-usage. |
| 12 | OOM kills | `dmesg |
| 13 | Journal errors | `journalctl -p err -S -1h --no-pager |
| 14 | Docker/container health | docker ps --filter health=unhealthy. Catches containers in unhealthy state. |
| 15 | Caddy/nginx error rate | Parse access logs for 5xx rate over last 5min. Expensive, do at 300s interval. |
| 16 | Time drift | `chronyc tracking |
| 17 | Open file descriptors | /proc/sys/fs/file-nr ratio used/total. >80% → warning (service exhaustion). |
| 18 | Backup freshness | Check timestamp of last backup file. >schedule+grace → critical. |
All of these (11–18) are implemented as ssh-script checks with
corresponding scripts in /opt/oikos/checks/.
9. Implementation order
Phase 1: Foundation (2-3 days)
- Refactor
executeCheckto return metrics map. Update all existing check functions (http,tcp,disk,cert-expiry). - Add
pingkind — ICMP reachability on the scheduler host. - Add
ssh-scriptkind — SSH execution engine, script allowlisting, JSON output parsing. - Add threshold evaluation — scheduler reads
config.thresholdsand compares against returned metrics to decide health/signal.
Phase 2: Host check scripts (1-2 days)
- Write and test each script in
/opt/oikos/checks/:cpu_check.sh,memory_check.sh,load_check.sh,swap_check.sh,disk_smart_check.sh,updates_check.sh,zfs_check.sh,process_check.sh,uptime_check.sh,oom_check.sh,journal_check.sh,time_check.sh,fd_check.sh. - Add
tools/setup-checks.shto auto-deploy scripts via the sync timer (same pattern astools/setup-caveman.sh).
Phase 3: Seed data + API (1 day)
- Add
seeds/checks.yamlwith default check definitions per entity type. - Add new check kinds to OpenAPI spec and regenerate Go types.
- Add a
/api/v1/checks/defaults/{entity_type}endpoint that returns recommended checks for a given entity type (convenience for operators).
Phase 4: Policy + observability (1 day)
- Add
read_onlypolicy rules forpingandssh-scriptactions. - Add per-check-kind metric recording (not just
probe_latency_ms). - Wire
cpu_temp,mem_pct,load_pct, etc. intoquery_metricsand the MCPget_trendtool.
Phase 5: Docker + service signals (1 day)
docker_health_check.sh—docker ps --filter health=unhealthy.caddy_error_rate.sh— parse Caddy JSON logs for 5xx.backup_freshness.sh— check last backup timestamp.
10. Dependencies
| Dependency | For | Risk |
|---|---|---|
golang.org/x/crypto/ssh |
SSH client in scheduler | Already in go.mod (used by deployer) |
golang.org/x/net/icmp + ipv4 |
ICMP ping | New dep; macOS needs root for raw sockets → exec fallback |
smartmontools on hosts |
disk_smart_check.sh |
Already installed on Proxmox; add to LXCs |
zfsutils-linux on hosts |
zfs_check.sh |
Already on Proxmox; add to LXCs with ZFS |
| SSH key on scheduler | ssh-script to all hosts |
Already deployed (Homelab sync SSH keys) |
11. Risks and mitigations
| Risk | Mitigation |
|---|---|
ssh-script is a remote exec vector |
Scripts are allowlisted by name (^[a-z][a-z0-9_-]+\.sh$), deployed via git (auditable), and read-only (no mutation). SSH key restricted to a dedicated oikos-check user with sudo only for systemctl is-active. |
| ICMP requires root on macOS | Fallback to ping CLI via os/exec. Production runs on Linux (strong) where raw sockets work. |
| Metrics cardinality explosion | Metrics are per-check-definition, not per-script-output. The script returns a fixed set of known metric names. TimescaleDB handles the volume. |
| Check script drift between hosts | Scripts deploy via tools/setup-checks.sh in the sync timer — same mechanism that keeps AGENTS.md in sync. Checksum validation before execution. |
12. Verification
- Unit tests: each check function (
checkPing,checkSSHScript) tested with mock SSH server and mock ICMP responses. - Integration test: deploy to
strong(macOS scheduler host), define checks forhubris(Proxmox),dns(LXC),caddy(LXC), run scheduler with--check-interval 10s, verify signals appear inget_signal_history. - MCP smoke test:
search_knowledge("signal triggers"),get_signal_history,query_metrics(metric=["cpu_pct", "mem_pct"]),get_trend. - Policy smoke test:
preflight(service:caddy, restart)still returns correct risk class despite new check kinds in the DB.