Installation
KubeSpend runs a server + console (SaaS-hosted by KubeSpend, or self-hosted in your VPC) and one or two agents in each cluster you want to monitor. This guide covers installing the agents.
Prerequisites
- Kubernetes 1.24+ and Helm 3.8+ (the charts are OCI artifacts on Amazon ECR Public at
oci://public.ecr.aws/kubespend.io/charts; nohelm repo add, no pull credentials) metrics-serverinstalled (free agent) — verify withkubectl top nodes- An API key from the KubeSpend console
- For the eBPF agent: Linux nodes, kernel 6.6+ (TCX). The free trial includes network data up to a flow-record allowance; the paid
networkadd-on removes the cap
1. kubespend-agent (free)
Reports node/pod CPU + memory and Kubernetes events. Runs as a single non-privileged Deployment: one pod for the whole cluster, whatever the node count, because it reads every node and pod from metrics-server and the Kubernetes API.
helm install kubespend-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-agent \
--version 0.1.0 \
--namespace kubespend --create-namespace \
--set apiKey="<YOUR_API_KEY>"--version pins the chart release, so the install is reproducible. It also skips the tag-listing call Helm otherwise makes to find the newest release. Drop it to install the latest.
The cluster name is auto-detected — no clusterId needed. Detection uses a friendly node label (e.g. eksctl alpha.eksctl.io/cluster-name) if present, otherwise a stable id from the kube-system namespace UID. Override with --set clusterId="my-cluster", or rename in the console later.
Using an existing Secret (recommended for GitOps)
Avoid putting the key in values. Create a Secret and reference it:
kubectl -n kubespend create secret generic kubespend-creds \
--from-literal=api-key="<YOUR_API_KEY>"
helm install kubespend-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-agent \
--version 0.1.0 \
--namespace kubespend --create-namespace \
--set clusterId="my-cluster" \
--set existingSecret="kubespend-creds" \
--set existingSecretKey="api-key"Key values
| Value | Default | Description |
|---|---|---|
clusterId | (auto-detected) | logical cluster name; auto-detected if unset |
apiKey | — (required) | API token (or use existingSecret) |
existingSecret | — | pre-created Secret name |
serverAddr | ingest.kubespend.io:443 | ingest endpoint (override if self-hosted) |
scrapeInterval | 30s | scrape/push cadence |
namespace | (all) | restrict to one namespace |
resources | 50m/64Mi → 200m/128Mi | requests/limits |
2. kubespend-ebpf-agent (network measurement)
Per-node DaemonSet (one pod on every node, control plane included, by tolerating every taint) that captures network flows via eBPF and attributes them to pods. Runs with fine-grained capabilities (BPF, PERFMON, NET_ADMIN) rather than privileged: true on kernel 6.6+.
What the network figure covers
Cross-AZ transfer and internet egress are priced and added to the cluster's daily cost. Three limits: the dollars are cluster-level daily only (per workload you get bytes split by zone), bytes whose peer zone could not be resolved are not priced, and NAT gateway and inter-region transfer are not modelled. Read a network figure as a floor — see What we measure.
relay.service is required: it names the Service fronting the metrics agent's relay port, has no default, and the chart calls fail without it, so a missing value breaks helm install rather than the running pod. The metrics-agent chart names its Service after its release (kubespend-agent for the release above, <release>-kubespend-agent when the release name does not contain kubespend-agent); confirm with:
kubectl get svc -n kubespend -l app.kubernetes.io/name=kubespend-agenthelm install kubespend-ebpf-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-ebpf-agent \
--version 0.1.0 \
--namespace kubespend \
--set relay.service="kubespend-agent"What a healthy install looks like
kubectl -n kubespend get pods -o wideOne kubespend-agent pod, plus one kubespend-ebpf-agent pod per node if you installed the network agent. A three-node cluster with both shows four pods. A single kubespend-agent pod is correct on any size of cluster. Running it once per node would report every node several times.
How flows reach the server
kubespend-ebpf-agent ──ServiceAccount token──▶ kubespend-agent ──API key──▶ core ingest
(DaemonSet, per node) (relay + aggregation) (outbound, once)The DaemonSet holds no API key. It ships flow records to the kubespend-agent already running in the cluster, which folds them into its own snapshot batches and performs the single outbound push. Authentication to that relay uses an audience-scoped projected ServiceAccount token, verified server-side by TokenReview.
Two things follow from that:
- Your ingest credential never lands on a node. It stays in one Secret in one namespace, so rotating it is a single edit rather than a fleet-wide rollout, and a compromised node yields a node-scoped token rather than an org ingest key.
kubespend-agentmust be installed and reachable in the same cluster. This is the chart default (mode: relay); it is not an optional topology.
clusterId must match on both charts. The relay accepts a flow window only when the cluster it names is the metrics agent's own. Both charts auto-detect the same name, so leave it unset on both or set the same value on both.
What the free trial includes
Network data is part of the free trial, up to a lifetime allowance of flow records (50 million by default; a small cluster emits roughly 350 per node every 30 seconds). It is checked server-side against the org the metrics agent's key belongs to. Once the allowance is used, new flows are dropped. Metrics keep flowing, and network data already stored stays visible. The paid network add-on removes the cap. Ask support to raise the allowance or enable the add-on.
Direct mode (no relay)
--set mode=direct makes each DaemonSet pod push to the ingest endpoint itself, and only then does it need apiKey or existingSecret. That puts the credential on every node and opens one outbound connection per node. Use it only where an in-cluster relay is not possible.
Key values
| Value | Default | Description |
|---|---|---|
clusterId | (auto-detected) | auto-detected; set to match the kubespend-agent value |
mode | relay | relay via kubespend-agent, or direct |
relay.service | — (required in relay mode) | Service fronting the metrics agent's relay port; the chart fails to render if unset |
apiKey / existingSecret | — (unused in relay mode) | direct mode only; relay mode carries no key |
vethRegex | ^(veth|lxc|cali|eni) | pod veth interfaces; direction is inverted |
nodeRegex | ^(ens|eth|en[posx]) | node uplinks; "" to skip |
window | 30s | aggregation/upload window |
flowDetail | conversation | connection keeps client ports, ~5× more records |
enableJitInitContainer | false | sets net.core.bpf_jit_enable=1 on the node; only needed where it reads 0 |
securityContext.privileged | false | true for bring-up; default uses fine-grained caps |
securityContext.capabilities.add | [BPF,PERFMON,NET_ADMIN] | add SYS_RESOURCE (<5.11), SYS_ADMIN (<5.8) |
Kernel note: the agent uses TCX links (kernel 6.6+). On older kernels a classic
tc/clsact fallback is on the roadmap. eBPF is Linux-only.
Troubleshooting
These are the failures seen most often on real installs. The console's Connect Cluster dialog lists the same ones under "If the install fails". Failures after the eBPF agent is running are on the eBPF agent page.
403 … Your authorization token has expired (or 400 Bad Request) from public.ecr.aws
The charts need no login. If Helm or Docker still holds an old login to public.ecr.aws, from pushing images or from another tool, Helm sends it anyway. ECR Public then refuses the expired token rather than falling back to anonymous access. Log out and retry:
helm registry logout public.ecr.aws
docker logout public.ecr.aws… exists and cannot be imported into the current release: invalid ownership metadata
An earlier install that did not use Helm left objects with the names the chart wants to create. Helm only adopts objects it created. The ClusterRole and ClusterRoleBinding are cluster-wide, so deleting the kubespend namespace does not remove them. Remove the leftovers the error names, then install again:
kubectl delete clusterrole,clusterrolebinding kubespend-agent kubespend-ebpf-agent --ignore-not-foundBe careful with kubectl delete ns kubespend. It removes everything else in that namespace too.
Only one kubespend-agent pod
Expected. See what a healthy install looks like. The per-node pods come from the network agent.
helm install kubespend-ebpf-agent fails before any pod exists
relay.service is unset. Find the metrics agent's Service and pass its name:
kubectl get svc -n kubespend -l app.kubernetes.io/name=kubespend-agentNo network data: flow window NOT delivered to relay … InvalidArgument
The two agents report different cluster names. The network agent logs this line every window, and the metrics agent logs relay submission rejected: cluster mismatch with declared= and expected= values. It happens when clusterId was set on kubespend-agent but not on kubespend-ebpf-agent, which then auto-detected another name. Set the same value on the network agent:
helm upgrade kubespend-ebpf-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-ebpf-agent \
--version 0.1.0 -n kubespend --reuse-values --set clusterId="<the expected= value>"The next window after the pods restart logs uploaded flow window.
Network data stops after a while on a trial
The trial's flow-record allowance is used up. The console's ingestion health panel says so for each cluster. Ask support to raise it, or enable the network add-on.
Network agent pods crash with attach ingress: not supported
That node's kernel is older than 6.6, which TCX needs. Check each node:
kubectl get nodes -o custom-columns='NODE:.metadata.name,KERNEL:.status.nodeInfo.kernelVersion'The cluster does not appear in the console
Read the metrics agent's log. A missing kubectl top nodes means metrics-server is not installed. An authentication error means the API key was revoked or mistyped.
kubectl -n kubespend logs -l app.kubernetes.io/name=kubespend-agent --tail=50Self-hosted server
Point serverAddr at your own ingest endpoint (e.g. ingest.kubespend.internal:443).
One thing catches every self-hosted deployment: the console's API URLs are inlined at build time by Vite, not read at runtime. A self-hosted console must therefore be built with your own VITE_AUTH_API_URL, VITE_CORE_API_URL and VITE_COST_API_URL; there is no environment variable to set on the running container. A build with them unset fails fast rather than shipping a console that loads and then cannot reach anything.
Run the server with KUBESPEND_EDITION=private. A self-hosted server then never refuses or throttles a request for usage: there is no snapshot allowance, no cluster limit and no rate limit on your agents or on the API. Usage is still metered, so the console's usage page shows snapshots, clusters and data received for monitoring. The per-IP limit on the login routes stays on, because it guards against password guessing rather than metering use.
For the full server-side deployment topology, contact support — that runbook is not published.
Uninstall
helm uninstall kubespend-agent -n kubespend
helm uninstall kubespend-ebpf-agent -n kubespend