Agent configuration
Both agents are configured through Helm values, which the chart renders into container environment variables. Every setting below can also be supplied as an environment variable directly if you run the container outside Helm.
Precedence: an explicit Helm value wins over the chart default.
kubespend-agent (metrics)
| Helm value | Environment variable | Default | Purpose |
|---|---|---|---|
clusterId | CLUSTER_ID | auto-detected | logical cluster name shown in the console |
apiKey / existingSecret | API_KEY | — (required) | ingest credential |
serverAddr | SERVER_ADDR | ingest.kubespend.io:443 | ingest endpoint; override when self-hosting |
ingestInsecure | INGEST_INSECURE | false | dial the ingest endpoint without TLS; local plaintext stacks only, since the API key is then sent in cleartext |
scrapeInterval | SCRAPE_INTERVAL | 30s | scrape and push cadence; must be greater than zero |
namespace | NAMESPACE | all namespaces | restrict collection to a single namespace |
resources | — | 50m / 64Mi requests | requests and limits |
Cluster ID auto-detection prefers a friendly node label (for example eksctl's alpha.eksctl.io/cluster-name), then falls back to a stable identifier derived from the kube-system namespace UID. Set it explicitly if you want a readable name — a UID works but reads badly in the console. You can also rename a cluster in the console after it registers.
Lowering scrapeInterval increases both stored volume and cost-query precision. The retention figures in Data retention assume the 30s default.
Restricting to one namespace means cluster-level cost is incomplete by design, and coverage will reflect that. See Coverage and confidence.
kubespend-ebpf-agent (network)
| Helm value | Environment variable | Default | Purpose |
|---|---|---|---|
mode | RELAY_MODE | relay | relay (via the metrics agent) or direct |
clusterId | CLUSTER_ID | auto-detected | must match the metrics agent |
apiKey / existingSecret | API_KEY | — | required in direct mode only |
vethRegex | POD_VETH_REGEX | ^(veth|lxc|cali|eni) | pod interfaces; direction is inverted |
nodeRegex | NODE_IFACE_REGEX | ^(ens|eth|en[posx]) | node uplinks; "" to skip |
ifaceWatchInterval | IFACE_WATCH_INTERVAL | 5s | how often to rescan for new interfaces |
window | FLOW_WINDOW | 30s | aggregation and upload window |
flowDetail | FLOW_DETAIL | conversation | conversation folds client ports into a connection count; connection keeps one record per 5-tuple |
enableJitInitContainer | — | false | sets net.core.bpf_jit_enable=1 on the node; only needed where it reads 0 |
securityContext.privileged | — | false | prefer the default; capabilities suffice on kernel 6.6+ |
securityContext.capabilities.add | — | [BPF, PERFMON, NET_ADMIN] | add SYS_RESOURCE below 5.11, SYS_ADMIN below 5.8 |
In the default relay mode the DaemonSet holds no API key — it authenticates to the in-cluster metrics agent with a projected ServiceAccount token, and the org key stays in a single Secret in one namespace. That is why rotating your key does not require touching every node.
Keep the two interface patterns disjoint
vethRegex and nodeRegex must not overlap. If they do, every pod byte is counted twice — once on the pod's veth and again on the node uplink. The defaults are disjoint; verify any override against ip link output on a real node.
Passing the API key
Prefer a pre-created Secret over an inline value, especially under GitOps:
kubectl -n kubespend create secret generic kubespend-creds \
--from-literal=api-key="$KUBESPEND_API_KEY"
helm install kubespend-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-agent \
--namespace kubespend --create-namespace \
--set existingSecret=kubespend-creds \
--set existingSecretKey=api-key--set apiKey=... writes the key into Helm release metadata stored in the cluster, where anyone with read access to the release can recover it. existingSecret avoids that.
Verifying a change
kubectl -n kubespend rollout status deploy/kubespend-agent
kubectl -n kubespend logs -l app.kubernetes.io/name=kubespend-agent --tail=50The agent logs a periodic summary rather than a line per scrape, so a healthy agent is quiet. Failures are always logged at warn or error regardless of level.
See also
- Installation — full install walkthrough
- Configuration — complete Helm value reference
- The eBPF agent — requirements and troubleshooting