Skip to content

DB Optimizer ​

DB Optimizer extends KubeSpend's cost visibility from Kubernetes compute into the managed data stores your workloads depend on — Amazon RDS (MySQL, PostgreSQL) and Amazon ElastiCache (Redis, Valkey).

What it does ​

Two independent evidence streams feed the panel:

  1. Infrastructure metrics read from Amazon CloudWatch, and the instance inventory from the RDS and ElastiCache APIs. These calls must run with credentials in your AWS account, so who makes them depends on where KubeSpend runs (see Where AWS is called from).
  2. Query telemetry (optional) captured by the eBPF agent, correlating database queries with the workloads that issued them.

Where AWS is called from ​

DeploymentDiscovery and CloudWatchStatus
Self-hosted / Private Cloud — kubespend-core runs in your AWS accountkubespend-core, with its IRSA role and the policy in SetupAvailable
KubeSpend Cloud — kubespend-core runs in KubeSpend's accountThe KubeSpend agent in your cluster, with an IAM role you grant itNot available yet

KubeSpend Cloud never calls AWS on your behalf. Its server runs in KubeSpend's own AWS account, where an RDS or CloudWatch call would describe the wrong account's databases, so the console says so instead of attempting discovery. Collection from your account by the agent is the planned path; until it ships, query capture (below) works on KubeSpend Cloud today, because the eBPF agent needs no AWS access.

Together they produce recommendations that neither stream can give alone: a specific query fingerprint issued by a named workload, matched to the CPU, connection and cache pressure it causes on a named database instance.

Architecture ​

┌─────────────────┐     CloudWatch      ┌───────────────────┐
│  Amazon RDS     │◄────GetMetricData───│  KubeSpend server │
│  Amazon         │                     │  (IAM role)       │
│  ElastiCache    │                     └───────────────────┘
└─────────────────┘                               ▲
                                                  │
┌─────────────────┐   query fingerprints          │
│  eBPF agent     │───(relayed by kubespend-agent)┘
│  (optional)     │
└─────────────────┘

The diagram is the self-hosted layout, where the KubeSpend server runs in your account.

Key points:

  • Self-hosted: CloudWatch is scraped by the KubeSpend server, with its own IAM role. No agent-side configuration needed.
  • KubeSpend Cloud: core makes no AWS call for discovery or metrics (DB_OPTIMIZER_AWS_SOURCE=agent).
  • The eBPF agent adds optional query capture for the correlation layer.
  • Only registered instances are polled — you control CloudWatch API cost directly.
  • Metrics are kept for 15 days, the same as raw node and pod metrics.

Monitoring depth ​

Each registered instance has a monitoring depth setting:

DepthWhat it providesRequires
CloudWatch OnlyCPU, memory, connections, IOPS, storage, replicationIAM permissions on core
CloudWatch + Query CaptureAll above + query fingerprints, execution counts, per-query latencyIAM, plus every precondition below — and one extra setting for TLS-encrypted databases

Query capture preconditions ​

All of these must hold. Miss one and capture produces nothing — in most cases silently.

  • DB_CAPTURE_MODE=fingerprint (or full) on the eBPF agent. The default is off, which loads no query program at all.
  • dbCapture.namespaces non-empty. This is a second, independent opt-in and is empty by default, so setting the mode alone captures nothing: no pod is in scope until a namespace is listed.
  • A kubespend-agent with relay query forwarding, in the default relay mode. Query batches ride each flow window to the kubespend-agent, which forwards them to the ingest endpoint with its own key, so the eBPF DaemonSet still holds no key. Against an older kubespend-agent the eBPF agent logs that it is discarding query statistics and counts them. mode: direct also works and sends them straight to the ingest endpoint.
  • POSTGRES_DSN configured on core. Without it the query ingest RPC is not registered, so there is nothing to upload to.
  • The db_optimizer entitlement on the org key. This is checked instead of network, not as well as it, so an org entitled to network flows alone has its query batches refused.

Encrypted databases need one more setting ​

The preconditions above enable capture from the wire, which works only on unencrypted connections. If your database enforces TLS — rds.force_ssl=1, or ElastiCache with in-transit encryption — the wire carries ciphertext and those settings produce no statements no matter how correct they are.

For those, add:

  • DB_UPROBE_MODE=capture on the eBPF agent, with CAP_SYS_PTRACE and hostPID. This reads the statement inside the client process, before it is encrypted. Without the capability the scan sees roughly 7% of the node and none of your application containers, which looks like "no databases here" rather than a misconfiguration.

What it covers today: Go clients on amd64 nodes. Other runtimes need an OpenSSL-level probe that is not built yet, and on arm64 (Graviton) the path is compiled out. Java is not addressable this way at all.

What it adds beyond statements: per-query latency on encrypted connections.

EngineHow latency is measuredCoverage
PostgreSQL (pgx, database/sql)the client library call enclosing the query is timed end to end, so connection-pool wait is includedall executions
Redis / ElastiCache (go-redis)the encrypted write is paired with the read that answers itmost executions

Redis coverage is partial by design. A client can pipeline several commands into one write, and the round trip that returns cannot be attributed to any single one of them — so no per-command latency is recorded for those, rather than a divided or duplicated guess. The console shows the percentage of executions a latency average covers whenever it is not all of them.

Next steps ​

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.