Skip to content

Architecture ​

  your cluster                                  KubeSpend server
┌──────────────────────────┐                 ┌────────────────────────────┐
│                          │                 │                            │
│  kubespend-agent         │   gRPC + API    │  ingest (gRPC)             │
│  ├─ metrics-server ──────┼── key, TLS ────▶│  ├─ snapshots              │
│  ├─ Kubernetes API       │    outbound,    │  └─ flows (entitled only)  │
│  ├─ relay + aggregation  │    once per     │                            │
│  └─ 30s snapshots        │    cluster      │  pricing                   │
│        ▲                 │                 │  ├─ AWS Price List API ────┼──▶ AWS
│        │ SA token        │                 │  └─ EC2 spot history       │    (read
│  kubespend-ebpf-agent    │                 │                            │     only)
│  ├─ TC/TCX BPF programs  │                 │  cost + savings engine     │
│  ├─ per-pod enrichment   │                 │                            │
│  └─ 30s flow windows     │                 │                            │
│                          │                 │                            │
└──────────────────────────┘                 │  HTTP API ──▶ console      │
                                             └──────┬─────────────────────┘
                                                    │
                                       time-series ─┴─ relational
                                     (metrics, flows)  (auth, tenancy)

Data path ​

  1. Collect. The metrics agent takes a snapshot every 30 seconds: node capacity and usage, pod requests and usage, owner references, PV/PVC inventory, HPA config, and recent events.
  2. Relay (eBPF only). Where the eBPF add-on is installed, each DaemonSet pod aggregates its flows over a 30 second window and submits them to the metrics agent inside the cluster, authenticating with an audience-scoped projected ServiceAccount token that the relay verifies by TokenReview. The DaemonSet holds no API key, so the ingest credential never reaches a node.
  3. Ship. The metrics agent folds relayed flows into its own snapshot and makes a single outbound push over gRPC with a bearer API key — one connection per cluster, not one per node. Large snapshots are split into batches; the agent estimates the wire size before marshalling so it can split without serializing twice.
  4. Resolve identity. The server resolves org_id from the API key and overwrites whatever the payload claimed. A client cannot assert which tenant it belongs to.
  5. Store. Samples land in the time-series store, scoped to the resolved tenant. Raw 30-second samples are kept 15 days, hourly aggregates 90 days, and daily aggregates 400 days; individual network flow records are kept for one day. See Data retention.
  6. Price. On a cost query, the server resolves each node group's hourly rate from cache, falling back to the AWS Price List API or spot history on a miss.
  7. Attribute. Node cost is split across CPU and memory, then across the pods that requested capacity on that node in that hour.
  8. Serve. The HTTP API returns cost breakdowns, attribution, and recommendations to the console.

Why the server holds the AWS credentials ​

The alternative — asking each customer cluster for pricing permissions — is worse in every direction. The server only ever calls pricing:GetProducts and ec2:DescribeSpotPriceHistory, both of which return public list prices. There is no customer-specific data in either response, so one set of read-only credentials serves every tenant, and no customer has to grant KubeSpend an IAM role.

Consequence: KubeSpend prices against list rates. An org-level Enterprise Discount Program setting is not available yet; Reserved Instances and Savings Plans are not modelled. See Pricing sources.

Why two agents ​

The free agent needs no special privileges: it reads metrics-server and the Kubernetes API, and runs as a normal Deployment. Network measurement genuinely requires kernel-level packet observation, which means a DaemonSet with BPF capabilities on every node. Bundling the two would force every user to accept the privileged footprint to get CPU and memory numbers.

Everything except network works without the eBPF agent. The console opens the network surfaces for an org with the network add-on or with a trial flow allowance. For an org with neither, they say so rather than rendering zeros.

Storage split ​

Two stores, chosen for two different access patterns:

  • A column-oriented time-series store for everything that arrives on a schedule — node and pod metrics, events, volume inventory, network flows — plus the hourly and daily rollups that back long-range cost history. Every query is tenant-scoped, and data ages out on the schedule in Data retention.
  • A relational store for accounts and tenancy: organisations, users, clusters, API keys, and entitlements. Small, transactional, no expiry.

The server itself is stateless — one binary, horizontally scalable, with all state in those two stores. That is what lets you run more than one replica behind a load balancer without coordination.

Self-hosting ​

The server and console can run inside your own VPC; agents then point at your ingest endpoint instead of the SaaS one. The console's API URL is inlined at build time by Vite, so a self-hosted console needs to be built with your own VITE_*_API_URL values rather than configured at runtime.

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.