System design interview ยท backend

Design a metrics monitoring system

"Build us something like Datadog or Prometheus: ingest 10 billion datapoints a day, keep 2 years of history, and answer 'show me p99 checkout latency for the last 30 days' in under a second."

โšก The takeaway first

Separate the write path from the read path, and roll data up aggressively at ingest time. Nobody queries raw per-second values from 6 months ago โ€” they query trends. So: ingest everything, immediately pre-compute 1-minute / 5-minute / 1-hour rollups, keep raw data for ~24 hours on fast disks, and serve old queries from tiny pre-aggregated buckets. The whole system is a funnel: wide at the top, narrow at the bottom.

๐Ÿ“‹ Requirements

Functional

Non-functional

๐Ÿšซ Common misconception

"We need to store every raw datapoint forever." You don't. A dashboard showing 90 days on a 1200-pixel-wide screen can only display 1200 points โ€” one per ~1.8 hours. Storing per-second resolution for that range is paying for precision nobody can see. Store raw briefly, rollups forever.

๐Ÿงฎ Back-of-the-envelope math

Think of it like a warehouse: every box that arrives must be shelved, and shelf space costs money. Let's count the boxes.

AssumptionValue
Unique time series100,000 (hosts ร— services ร— endpoints)
Sample interval1 sample / 10 sec per series
Write rate100,000 รท 10 = 10,000 writes/sec
Bytes per datapoint (timestamp + float)16 bytes
Raw volume10k ร— 16 B ร— 86,400 โ‰ˆ 13.8 GB/day

Now the funnel

TierResolutionRetentionSize
๐Ÿ”ฅ Hot (SSD / memory)raw (10s)24 h~14 GB
๐ŸŒก๏ธ Warm (cheap disk / object store)1-min rollup90 days100k ร— 1,440 ร— 16 B ร— 90 โ‰ˆ 207 GB
โ„๏ธ Cold (object store, Parquet)1-hour rollup2 years100k ร— 24 ร— 730 ร— 16 B โ‰ˆ 28 GB

Total โ‰ˆ 250 GB before compression. Time-series compressors (Gorilla / delta-of-delta) shrink this ~10ร—, so real disk is on the order of tens of GB โ€” a laptop could hold it. The hard part was never storage; it's write throughput and query fan-out.

Go deeper: the Gorilla compression trick

Facebook's Gorilla paper showed timestamps and floats compress beautifully: store the delta of deltas for timestamps (usually 0 bits when samples are regular) and XOR floats with the previous value (leading zeros dominate). That's why 13.8 GB/day of raw points often lands under 1.5 GB on disk. Interviewers love this reference โ€” it shows you know time-series data is special, not just "rows in a DB."

๐Ÿ—๏ธ Architecture

The centerpiece: one funnel, three temperatures.

flowchart TD
    A[App hosts & services
100k series] --> B[Ingest gateway
auth, batching, validation] B --> C[Stream processor
compute 1m / 5m / 1h rollups] C --> D[๐Ÿ”ฅ Hot store
raw, 24h - SSD] C --> E[๐ŸŒก๏ธ Warm store
1-min rollups, 90d] C --> F[โ„๏ธ Cold store
1-hour rollups, 2y - object store] G[Query API] --> D G --> E G --> F H[Dashboards & alerts] --> G I[Compactor
downsamples hot to warm to cold] --> E I --> F D --> I

Read the funnel like a kitchen: the ingest gateway is the receiving dock (checks the delivery), the stream processor is the prep cook (chops everything into standard portion sizes immediately), and the query API is the waiter who knows exactly which shelf each dish lives on based on how old the order is.

๐Ÿ” Component deep-dives

Write path โ€” batch, don't trickle

sequenceDiagram
    participant H as Host agent
    participant G as Ingest gateway
    participant S as Stream processor
    participant T as Hot store
    H->>H: batch 10s of samples
    H->>G: POST /write (1 batch, ~100 series)
    G->>G: validate + hash by series key
    G->>S: route partitions
    S->>S: update 1m/5m/1h rollup windows
    S->>T: append chunk (t,v pairs)
    T-->>H: 202 Accepted

Key insight: agents batch locally so the gateway sees ~1k requests/sec, not 10k. Partition by hash(series key) so one series always lands on the same worker โ€” rollups stay correct without cross-talk.

Read path โ€” fan out by age

sequenceDiagram
    participant U as Dashboard
    participant Q as Query API
    participant W as Warm store
    participant C as Cold store
    U->>Q: GET /query?metric=latency&from=-30d&step=1h
    Q->>Q: split range: [-30d,-1d] cold, [-1d,now] warm
    par fetch in parallel
        Q->>C: scan 1h buckets
    and
        Q->>W: scan 1m buckets, re-roll to 1h
    end
    Q->>Q: merge + sort by time
    Q-->>U: 720 points, p99 320ms

The query planner is the clever bit: it picks the coarsest tier that can answer the question. A 30-day query at 1-hour steps never touches raw data โ€” it reads 720 pre-computed points per series instead of 2.5 million.

๐Ÿ”Œ API + data model

API

POST /v1/write
  [{ metric:"http_latency_ms", labels:{svc:"checkout",pod:"a1"},
     points:[[t1,v1],[t2,v2],...] }]

GET /v1/query?metric=http_latency_ms&svc=checkout
        &from=2026-09-03&to=2026-10-03&step=1h&agg=p99

POST /v1/alerts   { metric, threshold, window, notify:"pagerduty" }

Data model

โš–๏ธ Trade-offs

DecisionOption AOption BPick
CollectionPush (agents send)Pull (server scrapes)Pull for services (Prometheus-style: dead targets are visible), push for ephemeral jobs
Precision vs costKeep raw 30 daysRaw 24h, rollups afterB โ€” dashboards can't render the difference
Query enginePre-aggregate everythingScan raw on demandPre-aggregate the common windows; allow raw scans only on the hot tier
Cardinality guardReject new seriesCharge/rate-limit per tenantBoth: hard cap per tenant + alert before hitting it

๐Ÿ”ฅ Failure modes

๐Ÿ› ๏ธ What I'd actually build

Opinionated and concrete: I would not build the storage engine. I'd run VictoriaMetrics cluster (or Mimir if the team knows Prometheus deeply) โ€” it already does the hot/warm split, Gorilla-style compression, and downsampling. Agents: Prometheus remote-write or OpenTelemetry Collector. Long-term cold: VictoriaMetrics' native historic support backed by S3 with Parquet. Query: Grafana on top, alerts via Alertmanager โ†’ PagerDuty. The genuinely custom code is thin: the ingest gateway (auth + tenant cardinality caps) and the query planner that routes by age. Everything else is bought, not built.

๐ŸŽค Interview tips

Go deeper: when raw data actually matters

Two cases break the "rollups forever" rule: forensics ("what exactly happened at 14:03:22 during the outage?") and ML training on raw signals. The honest answer: keep a sampled raw archive (e.g. 1% of series, or raw only for a golden set of critical metrics) in cold storage. Full raw forever is a cost conversation, not an engineering one โ€” put the number on the table: 13.8 GB/day ร— 730 days โ‰ˆ 10 TB raw, ~1 TB compressed. Sometimes the business says yes.

๐ŸŽฎ Interactive widget: watch raw points roll up

Raw datapoints stream in on the left. Watch them get swallowed into 1-minute buckets in real time โ€” this is exactly what the stream processor does. Then flip retention tiers and drag the downsampling slider.

Retention tiers โ€” pick one, watch the math

Downsampling visualizer โ€” one day of 1-minute data

Same 24 hours of data, re-bucketed. Drag the slider: fewer points, same shape โ€” until it isn't.

1 min โ†’ 1,440 pts

v2026.10.03-01