Prometheus & Grafana: Production Observability Stack
Minimum viable metrics pipeline: scrape architecture, RED/USE signals, SLO alerting, and on-call routing without alert fatigue.
Read NodePrometheus, Grafana, Loki, alerting pipelines, SLO design and NOC dashboards.
Production visibility — metrics, logs, traces, alerting and SLO governance. This category covers the stacks NOC and SRE teams rely on: Prometheus, Grafana, Loki, OpenTelemetry, PagerDuty routing and dashboard design.
We document what to monitor, how to alert without fatigue, and how to tie observability signals back to business SLOs. Distinct from SOC detection: this is platform and service health, not threat hunting.