New AnnouncementOpenSRE’s SRE Agent is now Open Source

Monitoring CI pipelines with Prometheus, StatsD, and Grafana

Instrument your CI/CD stack with Prometheus, StatsD, and Grafana to catch failures before they block releases.

const metadata = ; Most teams know when CI is broken, the merge queue stops, Slack fills up, and someone eventually types "looking." What they lack is continuous visibility into pipeline health before things go red. Prometheus, StatsD, and Grafana are a proven stack for that. StatsD gives you a lightweight way to emit metrics from CI scripts and runners. Prometheus stores them as time series. Grafana turns them into dashboards your platform team can actually use. This guide walks through a practical setup for CI/CD, not batch orchestrators, not data pipelines, focused on the metrics that predict broken main. What to measure in CI Before wiring tools, decide what matters. These metrics cover most reliability questions: | Metric | Type | Why it matters | | | | | | ci_build_duration_seconds | Histogram | Detect slow pipelines and regressions | | ci_build_total | Counter (labels: status, branch) | Track pass/fail volume | | ci_queue_wait_seconds | Histogram | Spot runner capacity problems | | ci_step_duration_seconds | Histogram (labels: step) | Find which stage dominates runtime | | ci_cache_hit_total | Counter | Validate cache effectiveness | | ci_flaky_rerun_total | Counter | Quantify test distrust | | ci_runner_active | Gauge | Monitor pool utilization | The goal isn't perfect coverage on day one. Start with build duration, failure rate, and queue wait: those three explain most "why is main red?" escalations. Architecture The data flow is straightforward: ` CI runner / workflow script │ UDP (StatsD) ▼ Prometheus StatsD Exporter ──scrape──▶ Prometheus ◀──query── Grafana ` CI jobs emit metrics over UDP during execution. The StatsD Exporter translates dot-notation names into Prometheus format. Prometheus scrapes the exporter on a fixed interval. Grafana queries Prometheus for dashboards and alerts. This works whether you run GitHub Actions self-hosted runners, GitLab CI, Jenkins agents, or Buildkite workers, anything that can run a script or sidecar can emit StatsD. Deploy the StatsD Exporter Run the official Prometheus StatsD Exporter with a mapping config so hierarchical metric names become labeled Prometheus metrics: `bash docker run -d \ -p 9125:9125/udp \ -p 9102:9102 \ -v $PWD/statsd_mapping.yml:/tmp/statsd_mapping.yml \ prom/statsd-exporter --statsd.mapping-config=/tmp/statsd_mapping.yml ` UDP 9125 receives StatsD packets. HTTP 9102 exposes /metrics for Prometheus to scrape. Mapping rules for CI metrics Without mapping, a metric like ci.build.main.unit_tests.duration becomes an unusable single Prometheus name. Define rules that extract labels: `yaml mappings: - match: "ci.build...duration" match_metric_type: observer name: "ci_build_duration_seconds" labels: branch: "$1" status: "$2" - match: "ci.step...duration" match_metric_type: observer name: "ci_step_duration_seconds" labels: pipeline: "$1" step: "$2" - match: "ci.cache.*" match_metric_type: counter name: "ci_cache_hit_total" labels: cache_key: "$1" ` Add Prometheus scrape config: `yaml scrape_configs: - job_name: "ci-metrics" static_configs: - targets: ["statsd-exporter:9102"] ` Emit metrics from CI jobs Any step can push StatsD metrics with a small script. Example from a GitHub Actions workflow: `bash After a build step completes DURATION_MS=$(( $(date +%s%3N) - START_MS )) echo "ci.build.$.$.duration:$|ms" | nc -u -w1 statsd-exporter 9125 echo "ci.build.$.$:1|c" | nc -u -w1 statsd-exporter 9125 ` For Python-based tooling, use the statsd client directly: `python import statsd client = statsd.StatsClient("statsd-exporter", 9125, prefix="ci") with client.timer("step.unit_tests.duration", pipeline="backend"): run_tests() client.incr("build.main.success") ` Practical tips: - Emit metrics at step boundaries, not every line of log output - Always label by branch, pipeline, and step, you'll filter on these constantly - Track re-runs explicitly (ci_flaky_rerun_total) so flake rate is visible, not hidden Grafana dashboards that matter Build two dashboards first. Resist the urge to chart everything. Pipeline health overview Panels worth having on day one: - Build success rate: rate(ci_build_total[1h]) / rate(ci_build_total[1h]) - p95 build duration: by branch and pipeline - Queue wait time: p50 and p95; spikes mean runner starvation - Active runners: gauge vs pool capacity - Failures by step: bar chart of ci_step_duration_seconds failures grouped by step label This dashboard answers: "Is CI healthy right now, and if not, where?" Merge queue / main branch view For trunk-based teams, filter everything to branch="main": - Time since last green main build - Failure rate on merge commits vs feature branches - Top failing steps in the last 24 hours - Flaky re-run count trend When main goes red, this is the screen platform pulls up first. Alerting thresholds Dashboards help humans investigate. Alerts wake them up, use sparingly. | Alert | Condition | Action | | | | | | Main branch failing | >3 consecutive failures on branch="main" | Page platform on-call | | Queue backlog | p95 queue wait > 15 min for 10 min | Scale runner pool | | Duration regression | p95 build duration +30% vs 7-day baseline | Investigate slow step | | Flake spike | ci_flaky_rerun_total rate 2× weekly average | Review quarantine policy | Wire these through Grafana Alerting or Prometheus Alertmanager. Route infra alerts to platform, not every developer. Common mistakes Metrics without labels. A single build_duration metric with no branch or step dimension is useless at scale. Scraping too infrequently. CI metrics are bursty. A 60-second scrape interval is fine; 5 minutes misses short failures. Dashboards nobody owns. Assign a DRI for the CI observability stack the same way you'd assign on-call for production. Alerting on every red build. Main will fail sometimes. Alert on patterns: consecutive failures, duration regressions, capacity saturation. Implication Prometheus, StatsD, and Grafana won't fix flaky tests or unclear ownership. What they do is make CI/CD reliability measurable: queue pressure, failure clusters, and duration drift become visible before they block every merge. Start small: three metrics, one overview dashboard, one alert. Expand once the platform team actually uses what you built.