Skip to content

FAQ ​

General ​

What does KubeSpend actually do? ​

Runs a small agent in your cluster, prices the resources it observes against real published AWS rates, and shows you where the money goes plus what to change. See What is KubeSpend.

Does it need access to my cloud bill? ​

No. No Cost and Usage Report, no billing export, no cross-account billing role. Prices come from the public AWS Price List API and EC2 spot price history, fetched by the KubeSpend server with its own read-only credentials. The only two permissions involved are pricing:GetProducts and ec2:DescribeSpotPriceHistory, both of which return public list prices.

Which clouds are supported? ​

AWS today. The cost model is cloud-agnostic but the pricing sources are AWS-specific.

Can I self-host? ​

Yes — the server and console run inside your own VPC, agents point at your ingest endpoint, and both data stores are yours. One thing to know: the console's API URL is inlined at build time by Vite, so a self-hosted console must be built with your own VITE_*_API_URL values rather than configured at runtime.

What is the pricing model? ​

Flat, based on cluster size. Not a percentage of savings. The metrics agent is free; network data from the eBPF agent is included in the free trial up to a flow-record allowance, and uncapped with the paid network add-on.

Accuracy ​

Why does a cost show as N/A instead of a number? ​

Because it could not be derived from a real price. KubeSpend never substitutes a fallback or an estimate. Common causes: the server has no AWS credentials (pricingAvailable: false), the region is outside the 16 supported ones, or an instance type could not be resolved — in which case it appears in coverage.unpriced. See Coverage and confidence.

Why is KubeSpend's figure lower than my AWS bill? ​

Most often one of:

  • coverage.partial is true — the agent was not reporting for part of the window, so cost is understated
  • coverage.unpriced is non-empty — those node-hours are excluded rather than guessed
  • network transfer is only partly priced: cross-AZ and internet egress are, NAT gateway and inter-region are not, and bytes whose peer zone could not be resolved are excluded
  • Reserved Instances and Savings Plans are not modelled; pricing is list-rate

Does it account for Reserved Instances or Savings Plans? ​

No. Figures are on-demand and spot list prices. RI and SP amortization is not implemented, and an org-level Enterprise Discount Program percentage is not configurable yet.

Why 9:1 for the CPU-to-memory split? ​

It is AWS's own ratio, taken from Fargate's published pricing and matching AWS split cost allocation data for EKS. See The cost model.

Are pods charged on requests or usage? ​

Requests. A request is what actually consumes the cluster — reserved capacity cannot be given to another pod whether it is used or not. Usage drives recommendations, not billing. See Cost attribution.

Network and eBPF ​

Does KubeSpend show network cost per workload? ​

Not per workload. Network dollars are cluster-level and daily: with the eBPF agent installed, flows are classified by zone and two buckets are priced — cross-AZ (intra-region) transfer at a flat per-GB rate, and internet egress on tiered rates, egress direction only — then folded into the cluster's daily cost. Per workload you get bytes split by zone, not dollars.

Read the figure as a floor. Same-AZ traffic is free so it contributes nothing, but bytes whose peer had no resolvable zone are not priced either, and NAT gateway and inter-region transfer are not modelled at all. coverage.networkUnclassifiedBytes tells you how much was left out. networkUsd is null, not 0, when there are no flow records or when no region could be resolved. See What we measure.

Then why bother with eBPF? ​

Because the measurement is the hard part, and no other mechanism gets it. AWS knows the bytes but not the workload; Kubernetes knows the workload but not the bytes. A TC program sees both, via the pod's veth interface. See Why eBPF for cost.

What is eBPF, in one paragraph? ​

A way to run small, kernel-verified programs inside Linux at defined hooks, without a kernel module or a reboot. Every program is proven safe before it loads — bounded loops, in-bounds memory access only. See What is eBPF, and ebpf.io for the wider ecosystem.

What kernel do I need? ​

6.6 or newer for the eBPF agent, which attaches via TCX. The free metrics agent has no kernel requirement. eBPF is Linux-only.

Does the eBPF agent need to be privileged? ​

Not on modern kernels. It runs with BPF, PERFMON and NET_ADMIN capabilities rather than privileged: true. Below kernel 5.11 add SYS_RESOURCE; below 5.8 add SYS_ADMIN.

What traffic does it miss? ​

IPv6, ICMP, SCTP, and QinQ-tagged frames. IPv4 TCP and UDP only.

Operations ​

How much overhead does the agent add? ​

The metrics agent is a single small Deployment reading metrics-server and the Kubernetes API on a 30 second cycle — requests around 50m CPU / 64Mi memory.

The eBPF agent runs one pod per node and accounts traffic per packet, which is its main overhead and an active area of work. On very busy nodes, check the logs for a rising drop count — that means the node is generating events faster than the agent drains them.

In every configuration, the connection is outbound only and originates from a single pod per cluster. You never open an inbound port, grant an IAM role, or peer a VPC with us.

What data leaves my cluster? ​

Resource metrics, object metadata (names, namespaces, owner references, labels), and Kubernetes events. With the eBPF agent, also:

  • network conversations with byte and connection counts, attributed to a pod: addresses, service port, protocol and direction. The client's ephemeral port is not kept by default (flowDetail: conversation); flowDetail: connection keeps one record per 5-tuple
  • DNS answers — resolved hostnames, kept for 180 days. Longer than the flows themselves, because a name has to outlive the traffic it explains — see Data retention
  • SQL and Redis query fingerprints, only when query capture is enabled. A fingerprint is a statement with every literal replaced by ?. It is not anonymised: table, column and schema names are transmitted unchanged.

Never Secret or ConfigMap contents, environment variables, or log contents. Application payloads never leave the cluster except under DB_CAPTURE_MODE=full, which additionally retains one redacted sample of the real statement text per fingerprint — redacted only as thoroughly as your redactPatterns describe. That mode also requires node-side consent (dbCapture.allowFull), so a console-side change alone cannot start recording statement text.

How long is data kept? ​

Raw 30-second samples 15 days, hourly rollups 90 days, daily rollups 400 days. Individual network flow records are kept for 1 day; the flow rollups keep ports and peers for 90 days hourly and 400 days daily. Rollups are written as data arrives rather than from expiring raw rows, so long-range history survives the raw TTL. See Data retention.

Why does my cluster show as a UID instead of a name? ​

Cluster ID auto-detection fell back to the kube-system namespace UID because no friendly node label was found. Set --set clusterId="my-cluster", or rename it in the console.

My byte counts look about double. Why? ​

The veth and node interface regexes are overlapping, so each pod byte is counted on both the veth and the uplink. See Agent configuration.

Can I restrict the agent to one namespace? ​

Yes, -namespace / NAMESPACE. Cluster-level cost will then be incomplete by design.

Why can the savings total look too good? ​

The headline figure sums workload rightsizing, instance recommendations and consolidation plans, which overlap by construction — rightsizing a workload and consolidating away the node it runs on can describe the same dollar twice. Individual findings are reliable; treat the grand total as an upper bound. See Rightsizing.

Security ​

How do agents authenticate? ​

The metrics agent uses a bearer API key over gRPC. Keys are per-org and stored hashed. Anything a client claims about its own identity — an org_id in a payload, an email in a body — is ignored; the server resolves identity from the key.

The eBPF agent holds no API key. It submits flows to the metrics agent inside the cluster using an audience-scoped projected ServiceAccount token, verified by TokenReview. The metrics agent aggregates those flows into its own batches and makes the single outbound push.

Does the eBPF agent need its own API key? ​

No. In the default relay mode it holds no API key at all. What it does need is relay.service — the name of the Service fronting the metrics agent's relay port. It has no default, and the chart calls fail without it, so this is an install-time template error, not a runtime dial failure:

bash
# the metrics-agent chart names its Service after the release: kubespend-agent here
kubectl get svc -n kubespend -l app.kubernetes.io/name=kubespend-agent

helm install kubespend-ebpf-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-ebpf-agent \
  --namespace kubespend \
  --set relay.service="kubespend-agent"

No default would be right for every install, and a wrong guess would surface as a CrashLoop rather than an error, so the chart refuses to guess.

Your ingest credential stays in one Secret in one namespace rather than on every node. Rotating it is a single edit, and a compromised node yields a node-scoped token instead of an org ingest key.

Three consequences:

  • kubespend-agent must be installed and reachable in the same cluster.
  • Both charts must report the same clusterId, or the relay rejects every window.
  • The server applies the org's network allowance (trial flow records, or the uncapped network add-on) against that agent's org key.

--set mode=direct opts out, pushing from each node directly. Only then does apiKey apply.

How do users sign in? ​

Email OTP by default; no passwords are stored. Google SSO is optional and only enabled when both the client ID and client secret are configured — otherwise the route is not registered at all rather than existing in a half-configured state.

Should I pass the API key with --set apiKey=? ​

Prefer an existing Secret. --set writes the key into Helm release metadata stored in the cluster, where anyone with read access to the release can recover it. See Agent configuration.

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.