"Build us something like Datadog or Prometheus: ingest 10 billion datapoints a day, keep 2 years of history, and answer 'show me p99 checkout latency for the last 30 days' in under a second."
(metric name, labels, timestamp, value) from thousands of hosts."We need to store every raw datapoint forever." You don't. A dashboard showing 90 days on a 1200-pixel-wide screen can only display 1200 points โ one per ~1.8 hours. Storing per-second resolution for that range is paying for precision nobody can see. Store raw briefly, rollups forever.
Think of it like a warehouse: every box that arrives must be shelved, and shelf space costs money. Let's count the boxes.
| Assumption | Value |
|---|---|
| Unique time series | 100,000 (hosts ร services ร endpoints) |
| Sample interval | 1 sample / 10 sec per series |
| Write rate | 100,000 รท 10 = 10,000 writes/sec |
| Bytes per datapoint (timestamp + float) | 16 bytes |
| Raw volume | 10k ร 16 B ร 86,400 โ 13.8 GB/day |
| Tier | Resolution | Retention | Size |
|---|---|---|---|
| ๐ฅ Hot (SSD / memory) | raw (10s) | 24 h | ~14 GB |
| ๐ก๏ธ Warm (cheap disk / object store) | 1-min rollup | 90 days | 100k ร 1,440 ร 16 B ร 90 โ 207 GB |
| โ๏ธ Cold (object store, Parquet) | 1-hour rollup | 2 years | 100k ร 24 ร 730 ร 16 B โ 28 GB |
Total โ 250 GB before compression. Time-series compressors (Gorilla / delta-of-delta) shrink this ~10ร, so real disk is on the order of tens of GB โ a laptop could hold it. The hard part was never storage; it's write throughput and query fan-out.
Facebook's Gorilla paper showed timestamps and floats compress beautifully: store the delta of deltas for timestamps (usually 0 bits when samples are regular) and XOR floats with the previous value (leading zeros dominate). That's why 13.8 GB/day of raw points often lands under 1.5 GB on disk. Interviewers love this reference โ it shows you know time-series data is special, not just "rows in a DB."
The centerpiece: one funnel, three temperatures.
flowchart TD
A[App hosts & services
100k series] --> B[Ingest gateway
auth, batching, validation]
B --> C[Stream processor
compute 1m / 5m / 1h rollups]
C --> D[๐ฅ Hot store
raw, 24h - SSD]
C --> E[๐ก๏ธ Warm store
1-min rollups, 90d]
C --> F[โ๏ธ Cold store
1-hour rollups, 2y - object store]
G[Query API] --> D
G --> E
G --> F
H[Dashboards & alerts] --> G
I[Compactor
downsamples hot to warm to cold] --> E
I --> F
D --> I
Read the funnel like a kitchen: the ingest gateway is the receiving dock (checks the delivery), the stream processor is the prep cook (chops everything into standard portion sizes immediately), and the query API is the waiter who knows exactly which shelf each dish lives on based on how old the order is.
sequenceDiagram
participant H as Host agent
participant G as Ingest gateway
participant S as Stream processor
participant T as Hot store
H->>H: batch 10s of samples
H->>G: POST /write (1 batch, ~100 series)
G->>G: validate + hash by series key
G->>S: route partitions
S->>S: update 1m/5m/1h rollup windows
S->>T: append chunk (t,v pairs)
T-->>H: 202 Accepted
Key insight: agents batch locally so the gateway sees ~1k requests/sec, not 10k. Partition by hash(series key) so one series always lands on the same worker โ rollups stay correct without cross-talk.
sequenceDiagram
participant U as Dashboard
participant Q as Query API
participant W as Warm store
participant C as Cold store
U->>Q: GET /query?metric=latency&from=-30d&step=1h
Q->>Q: split range: [-30d,-1d] cold, [-1d,now] warm
par fetch in parallel
Q->>C: scan 1h buckets
and
Q->>W: scan 1m buckets, re-roll to 1h
end
Q->>Q: merge + sort by time
Q-->>U: 720 points, p99 320ms
The query planner is the clever bit: it picks the coarsest tier that can answer the question. A 30-day query at 1-hour steps never touches raw data โ it reads 720 pre-computed points per series instead of 2.5 million.
POST /v1/write
[{ metric:"http_latency_ms", labels:{svc:"checkout",pod:"a1"},
points:[[t1,v1],[t2,v2],...] }]
GET /v1/query?metric=http_latency_ms&svc=checkout
&from=2026-09-03&to=2026-10-03&step=1h&agg=p99
POST /v1/alerts { metric, threshold, window, notify:"pagerduty" }
metric name + sorted label set โ hashed to a 64-bit ID. Everything (storage, routing, rollups) keys off this ID.(timestamp, value) pairs for one series, Gorilla-compressed. Immutable once sealed โ the unit of storage and replication.| Decision | Option A | Option B | Pick |
|---|---|---|---|
| Collection | Push (agents send) | Pull (server scrapes) | Pull for services (Prometheus-style: dead targets are visible), push for ephemeral jobs |
| Precision vs cost | Keep raw 30 days | Raw 24h, rollups after | B โ dashboards can't render the difference |
| Query engine | Pre-aggregate everything | Scan raw on demand | Pre-aggregate the common windows; allow raw scans only on the hot tier |
| Cardinality guard | Reject new series | Charge/rate-limit per tenant | Both: hard cap per tenant + alert before hitting it |
user_id label; series count jumps 100k โ 50M overnight. Mitigation: per-tenant series caps, alerting on new-series rate.Opinionated and concrete: I would not build the storage engine. I'd run VictoriaMetrics cluster (or Mimir if the team knows Prometheus deeply) โ it already does the hot/warm split, Gorilla-style compression, and downsampling. Agents: Prometheus remote-write or OpenTelemetry Collector. Long-term cold: VictoriaMetrics' native historic support backed by S3 with Parquet. Query: Grafana on top, alerts via Alertmanager โ PagerDuty. The genuinely custom code is thin: the ingest gateway (auth + tenant cardinality caps) and the query planner that routes by age. Everything else is bought, not built.
Two cases break the "rollups forever" rule: forensics ("what exactly happened at 14:03:22 during the outage?") and ML training on raw signals. The honest answer: keep a sampled raw archive (e.g. 1% of series, or raw only for a golden set of critical metrics) in cold storage. Full raw forever is a cost conversation, not an engineering one โ put the number on the table: 13.8 GB/day ร 730 days โ 10 TB raw, ~1 TB compressed. Sometimes the business says yes.
Raw datapoints stream in on the left. Watch them get swallowed into 1-minute buckets in real time โ this is exactly what the stream processor does. Then flip retention tiers and drag the downsampling slider.
Same 24 hours of data, re-bucketed. Drag the slider: fewer points, same shape โ until it isn't.