Data retention
How far back you can query, what is kept at what resolution, and what leaves your cluster.
Collection intervals
| What | Interval |
|---|---|
| Metrics agent snapshot | 30s |
| eBPF flow aggregation window | 30s |
Both are configurable — see Agent configuration. The retention figures below assume the 30 second default.
Retention by resolution
Data is kept in three tiers. Hourly and daily rollups are computed as samples arrive, not when raw data ages out, so long-range history stays available at lower resolution after the raw samples are gone.
| Tier | Resolution | Kept for | What you can do with it |
|---|---|---|---|
| Raw | one sample per scrape (30s) | 15 days | inspect a specific spike, debug a short-lived pod |
| Hourly | aggregated per hour | 90 days | per-hour cost and utilization trends |
| Daily | aggregated per day | 400 days | month-over-month and year-over-year cost history |
Because a rollup is written at the moment its samples are ingested, a raw sample expiring never leaves a hole in longer-range history. Totals stay correct across tier boundaries.
The hourly tier keeps each pod, on the node it ran on, for the full 90 days. That is what lets the workload and pod drawers show 7 and 30 days of history, and what lets per-workload cost attribution cover 90 days: a pod's requests are priced against the instance it ran on. In both drawers a range of a day or less is drawn per scrape; anything longer is drawn per hour.
Other data
The three tiers above describe node and pod metrics. Everything else has its own window:
| Data | Kept for |
|---|---|
| Kubernetes events | 15 days |
| Persistent volume capacity and usage | 15 days |
| Individual network flows | 1 day |
| Network flows per 5 minutes | 15 days |
| Network flows per hour | 90 days |
| Network flows per day | 400 days |
| Resolved hostnames (DNS answers) | 180 days |
| RDS / ElastiCache CloudWatch metrics | 15 days |
| Raw query windows | 48 hours |
| Query hourly rollups | 15 days |
| Query fingerprint text | 30 days from when the fingerprint was first seen |
Sample statement text (full mode only) | 48 hours |
| Query coverage metadata | 15 days |
| Daily usage history (the console's Usage page) | 400 days |
Events, persistent volume metrics and database CloudWatch metrics have no rollup, so their 15 days is a hard limit rather than the start of a coarser tier.
Individual flow records are kept for a day because nothing past a day needs them one by one. The flow rollups keep the service port, protocol and node alongside the addresses, so per-peer, per-port and Internet-destination views read the rollups and reach back the full 90 days. What stays on individual records is the flow list itself, which shows the last 24 hours whatever range the page is set to, and says so.
Query text is deliberately the shortest-lived data in the store. A fingerprint carries your schema names and a full-mode sample carries real statement text, so both expire well before the aggregates computed from them.
What this means for queries
A 7 day cost query is fully detailed — you can drill into an individual hour.
A 6 month query reads daily aggregates: the totals are accurate, but there is no per-hour detail to open up. A window that reaches further back than the data kept for that question is refused with an error that states how many days are retained, rather than answered with a silently shorter range.
Per-workload cost attribution reaches 90 days. A longer request attributes the most recent 90 days and reconciles them against the cluster total for those same 90 days — compute, storage and network alike — and marks the response's coverage as partial, rather than counting the older cluster spend as unallocated.
If a chart looks coarser than you expected, the window has crossed a tier boundary.
Why raw data expires at 15 days
Volume. At a 30 second interval a 100-node cluster produces on the order of 288,000 node samples per day, plus pod samples scaling with pod count. Individual 30-second samples stop being useful almost immediately — once a day has passed, the questions people ask are all aggregate ones. Keeping them forever would grow the store without bound to answer questions nobody asks.
Node and workload cost inputs are preserved in the rollups, so compute cost history survives raw samples expiring. Persistent volume and database metrics are the exception: they have no rollup and are kept for 15 days.
Self-hosted tuning
Retention windows are fixed in the current release rather than configurable. If you self-host and need different windows, open an issue describing the requirement.
What leaves your cluster
Self-hosted: nothing. Agents talk to your own ingest endpoint and the data stores are yours.
SaaS: the agents send
- resource metrics (CPU, memory, storage capacity and usage)
- object metadata — names, namespaces, owner references, labels
- Kubernetes events
- with the eBPF add-on: network flow records (addresses, service port, protocol, byte and connection counts) attributed to a pod. The client's ephemeral port is stored as 0 unless
flowDetail: connectionis set - with the eBPF add-on: DNS answers — the hostnames pods resolved, kept for 180 days. This is longer than the flow records themselves, and deliberately so: the table is what turns an address in a network chart into a name you recognise, so it has to outlive the traffic it explains. 180 days is twice the 90-day hourly flow rollup, which is the longest-lived data that cites it
- with query capture enabled: SQL and Redis query fingerprints — statements with every literal replaced by
?. A fingerprint is not anonymised: table, column and schema names are transmitted unchanged. - under
DB_CAPTURE_MODE=fullonly: one redacted sample of the real statement text per fingerprint, redacted only as thoroughly as yourredactPatternsdescribe
The agents never send Secret or ConfigMap contents, environment variables, or log contents. Application payloads never leave the cluster except in full mode, where the retained sample is drawn from real statement text. That is why the mode requires node-side consent (dbCapture.allowFull) on top of being switched on, and why the sample expires after 48 hours.