Refetching everything on a health event was wasteful and churned the UI: one
container going degraded pulled down the entire fleet entity list (plus its
parent-grouping pass), or the whole fleet graph, to learn something the event
had already delivered.
health.changed / health.stale carry the new value in their payload, so the
views that hold the entity just patch it:
- Fleet table: patch the row. Only entity.* changes which entities exist, so
only that still refetches.
- Fleet map: patch the node AND graph.health[id] — healthOf() reads the side
map in preference to the node's own field, so patching only the nodes would
have left the rendered colour unchanged.
- Entity detail: patch the open entity. Signals still need a read (the event
says one was raised, not what the list now contains) but only the signals,
not the entity and checks alongside them.
Shared in $lib/health.ts, which returns the original array when an event does
not apply so unrelated rows keep their identity and do not re-render. Note it
matches on entity_id, never data.slug: the scheduler emits health.changed with
entity_id = the observed entity but slug = the *check's* slug.
Separately, events.ts had no reconnect. onerror was empty on the assumption
the browser retries, but EventSource only does that for a transient failure --
once it reaches CLOSED (an HTTP error on connect, e.g. the API restarting
during a deploy) it stays closed forever. A single blip silently froze every
live surface in the app with nothing on screen to say so. Now reconnects with
capped exponential backoff, and exports eventsConnected so a future indicator
can show when the stream is down.
Verified against live prod: flipping lxc:apps health recoloured the map node
and moved its counts (30 healthy -> 29, 9 down -> 10) with ZERO network
requests.
Co-Authored-By: Claude <noreply@anthropic.com>