Projects · Write-up
Summary: Disk fill looks like a mystery outage until you check the gauges. Prometheus + Grafana + an alert path caught a lab guest before it looked “mysteriously dead.”
Step 1 — Name the symptom
A lab guest used for local services stopped responding cleanly. Without metrics, the first guess is always “network” or “hypervisor” — expensive wrong turns.
Step 2 — Check observability first
Check the observability path first: node exporter / guest metrics in Prometheus, a Grafana panel for filesystem use, and a simple threshold alert. Zabbix covers host-level checks in the same lab footprint.
Step 3 — Read before and after
- Before: Symptom-first troubleshooting; disk pressure not visible on the primary dashboard.
- After: Filesystem panel + alert; disk-fill shows as the top signal; the fix is cleanup or resize, not fabric superstition.
Step 4 — Sketch the path
[lab guests] --metrics--> [Prometheus] --> [Grafana panels]
-> alert path (threshold)
[lab hosts] --checks---> [Zabbix]
Lab only — no private addressing, no employer incident data.
Step 5 — Close the loop
The catch was boring on purpose: the alert fired, the panel confirmed disk pressure, the fix was local. Transferable judgment: instrument the break you actually fear — disk fill, not mystery fabric ghosts.
Outcome
Alert → clear cause → resize/cleanup. Write-up kept employer-neutral. No dedicated repo for this note; related public samples live under /git/.