Observability & DevOps Tooling
Production Monitoring & Alerting Stack
The Grafana and Prometheus-based monitoring and alerting layer behind how we run production service health dashboards, SLO tracking, on-call alerting, and deploy correlation.
Refresh interval
5–10s
Alert severity
Critical, Warning, Info
SLO tracking
Error budget burn rate
Alert routing
Slack + PagerDuty
Why This Exists
Software that ships without a way to see how it's behaving in production is only half-built. Observability is engineering work in its own right not an afterthought bolted on after the first outage. This stack is what we built and operate across our own production deployments.
Service Health Dashboard
Real-time visibility into every service: request throughput, latency at p50/p95/p99, error rates, CPU and memory utilisation by node, queue depth, and disk usage all on a 5–10s refresh so degradation is caught before a user has to report it.
- ▸Live request throughput and p50/p95/p99 latency by service
- ▸Per-node CPU, memory, and queue depth
- ▸Error rate and uptime tracking
- ▸Disk usage by volume with threshold-based coloring
Alerting & Incident Overview
A second dashboard focused on incident detection and resolution: MTTD and MTTR tracked over time, alert volume by severity, SLO error budget burn rate, and deploy-to-rollback correlation. On-call engineers get the context they need without digging through logs first.
- ▸Mean time to detect (MTTD) and mean time to resolve (MTTR)
- ▸Alert volume by critical, warning, and info severity
- ▸SLO budget remaining and error budget burn rate
- ▸Deploy frequency plotted against rollback rate
- ▸Alertmanager routing to Slack and PagerDuty by severity
Tech Stack
Work with me
Building something similar?
I design and build production-grade backend systems, APIs, and cloud infrastructure. If you're working on a problem in this space, let's talk.
Get in touch