Lab incident — observability catch

Projects · Write-up

Summary: Disk fill looks like a mystery outage until you check the gauges. Prometheus + Grafana + an alert path caught a lab guest before it looked “mysteriously dead.”

Step 1 — Name the symptom

A lab guest used for local services stopped responding cleanly. Without metrics, the first guess is always “network” or “hypervisor” — expensive wrong turns.

Step 2 — Check observability first

Check the observability path first: node exporter / guest metrics in Prometheus, a Grafana panel for filesystem use, and a simple threshold alert. Zabbix covers host-level checks in the same lab footprint.

Step 3 — Read before and after

  • Before: Symptom-first troubleshooting; disk pressure not visible on the primary dashboard.
  • After: Filesystem panel + alert; disk-fill shows as the top signal; the fix is cleanup or resize, not fabric superstition.

Step 4 — Sketch the path

[lab guests] --metrics--> [Prometheus] --> [Grafana panels]
                              -> alert path (threshold)
[lab hosts]  --checks---> [Zabbix]

Lab only — no private addressing, no employer incident data.

Step 5 — Close the loop

The catch was boring on purpose: the alert fired, the panel confirmed disk pressure, the fix was local. Transferable judgment: instrument the break you actually fear — disk fill, not mystery fabric ghosts.

Outcome

Alert → clear cause → resize/cleanup. Write-up kept employer-neutral. No dedicated repo for this note; related public samples live under /git/.

Lab · Prometheus, Grafana, Zabbix