DB Optimizer
DB Optimizer extends KubeSpend's cost visibility from Kubernetes compute into the managed data stores your workloads depend on — Amazon RDS (MySQL, PostgreSQL) and Amazon ElastiCache (Redis, Valkey).
What it does
Two independent evidence streams feed the panel:
- Infrastructure metrics read from Amazon CloudWatch, and the instance inventory from the RDS and ElastiCache APIs. These calls must run with credentials in your AWS account, so who makes them depends on where KubeSpend runs (see Where AWS is called from).
- Query telemetry (optional) captured by the eBPF agent, correlating database queries with the workloads that issued them.
Where AWS is called from
| Deployment | Discovery and CloudWatch | Status |
|---|---|---|
| Self-hosted / Private Cloud — kubespend-core runs in your AWS account | kubespend-core, with its IRSA role and the policy in Setup | Available |
| KubeSpend Cloud — kubespend-core runs in KubeSpend's account | The KubeSpend agent in your cluster, with an IAM role you grant it | Not available yet |
KubeSpend Cloud never calls AWS on your behalf. Its server runs in KubeSpend's own AWS account, where an RDS or CloudWatch call would describe the wrong account's databases, so the console says so instead of attempting discovery. Collection from your account by the agent is the planned path; until it ships, query capture (below) works on KubeSpend Cloud today, because the eBPF agent needs no AWS access.
Together they produce recommendations that neither stream can give alone: a specific query fingerprint issued by a named workload, matched to the CPU, connection and cache pressure it causes on a named database instance.
Architecture
┌─────────────────┐ CloudWatch ┌───────────────────┐
│ Amazon RDS │◄────GetMetricData───│ KubeSpend server │
│ Amazon │ │ (IAM role) │
│ ElastiCache │ └───────────────────┘
└─────────────────┘ ▲
│
┌─────────────────┐ query fingerprints │
│ eBPF agent │───(relayed by kubespend-agent)┘
│ (optional) │
└─────────────────┘The diagram is the self-hosted layout, where the KubeSpend server runs in your account.
Key points:
- Self-hosted: CloudWatch is scraped by the KubeSpend server, with its own IAM role. No agent-side configuration needed.
- KubeSpend Cloud: core makes no AWS call for discovery or metrics (
DB_OPTIMIZER_AWS_SOURCE=agent). - The eBPF agent adds optional query capture for the correlation layer.
- Only registered instances are polled — you control CloudWatch API cost directly.
- Metrics are kept for 15 days, the same as raw node and pod metrics.
Monitoring depth
Each registered instance has a monitoring depth setting:
| Depth | What it provides | Requires |
|---|---|---|
| CloudWatch Only | CPU, memory, connections, IOPS, storage, replication | IAM permissions on core |
| CloudWatch + Query Capture | All above + query fingerprints, execution counts, per-query latency | IAM, plus every precondition below — and one extra setting for TLS-encrypted databases |
Query capture preconditions
All of these must hold. Miss one and capture produces nothing — in most cases silently.
DB_CAPTURE_MODE=fingerprint(orfull) on the eBPF agent. The default isoff, which loads no query program at all.dbCapture.namespacesnon-empty. This is a second, independent opt-in and is empty by default, so setting the mode alone captures nothing: no pod is in scope until a namespace is listed.- A kubespend-agent with relay query forwarding, in the default relay mode. Query batches ride each flow window to the kubespend-agent, which forwards them to the ingest endpoint with its own key, so the eBPF DaemonSet still holds no key. Against an older kubespend-agent the eBPF agent logs that it is discarding query statistics and counts them.
mode: directalso works and sends them straight to the ingest endpoint. POSTGRES_DSNconfigured on core. Without it the query ingest RPC is not registered, so there is nothing to upload to.- The
db_optimizerentitlement on the org key. This is checked instead ofnetwork, not as well as it, so an org entitled to network flows alone has its query batches refused.
Encrypted databases need one more setting
The preconditions above enable capture from the wire, which works only on unencrypted connections. If your database enforces TLS — rds.force_ssl=1, or ElastiCache with in-transit encryption — the wire carries ciphertext and those settings produce no statements no matter how correct they are.
For those, add:
DB_UPROBE_MODE=captureon the eBPF agent, withCAP_SYS_PTRACEandhostPID. This reads the statement inside the client process, before it is encrypted. Without the capability the scan sees roughly 7% of the node and none of your application containers, which looks like "no databases here" rather than a misconfiguration.
What it covers today: Go clients on amd64 nodes. Other runtimes need an OpenSSL-level probe that is not built yet, and on arm64 (Graviton) the path is compiled out. Java is not addressable this way at all.
What it adds beyond statements: per-query latency on encrypted connections.
| Engine | How latency is measured | Coverage |
|---|---|---|
PostgreSQL (pgx, database/sql) | the client library call enclosing the query is timed end to end, so connection-pool wait is included | all executions |
| Redis / ElastiCache (go-redis) | the encrypted write is paired with the read that answers it | most executions |
Redis coverage is partial by design. A client can pipeline several commands into one write, and the round trip that returns cannot be attributed to any single one of them — so no per-command latency is recorded for those, rather than a divided or duplicated guess. The console shows the percentage of executions a latency average covers whenever it is not all of them.
Next steps
- Setup & IAM permissions — what to add to your IRSA role
- RDS metrics reference — which CloudWatch metrics are collected
- ElastiCache metrics reference — cache-specific metrics