Skip to content

The eBPF agent ​

kubespend-ebpf-agent is a DaemonSet that runs one pod per node, loads TC programs onto pod and node interfaces, and hands aggregated flows to the in-cluster kubespend-agent, which batches them alongside its own metrics and makes the single outbound push to the server.

It carries no API key of its own — it authenticates to that relay with an audience-scoped projected ServiceAccount token, so your ingest credential never reaches a node. See how flows reach the server.

The free trial includes its network data up to a lifetime flow-record allowance; the paid network add-on removes the cap. Both are checked server-side against the org the metrics agent's key belongs to. See what the free trial includes.

What it observes ​

The agent attaches traffic-control (TC) programs to both the ingress and egress paths of the interfaces it selects, using TCX links — which is where the kernel 6.6 floor comes from. TC is the only hook used: no kprobes, no XDP.

For each packet it accounts:

FieldWhy
source and destination addressidentifies the peer
source and destination portdistinguishes flows between the same pair
protocolTCP or UDP
directioningress or egress, normalised to the pod's perspective
interfacethe key that maps traffic back to a pod
byte countthe quantity you are billed on

Records are counted in the kernel and drained to the agent, which aggregates them before sending. Drops under load are counted and logged rather than silently discarded — an uncounted drop is indistinguishable from traffic that never happened, which would make a byte-accounting error impossible to detect.

Scope limits: IPv4 with TCP or UDP only. Everything else passes through unaccounted. A single VLAN tag is handled; QinQ is not.

Performance note

Per-packet accounting on a very busy node is the agent's main overhead, and reducing it is an active area of work. If you see drops reported in the logs, that node is producing more events than the agent is draining.

What runs in userspace ​

Interface selection, by two separate regexes:

RoleDefaultMeaning
Pod veth^(veth|lxc|cali|eni)traffic belongs to a pod; direction is inverted
Node uplink^(ens|eth|en[posx])traffic belongs to the node; direction as-observed

These are deliberately two patterns, not one. An earlier combined default of ^(veth|eni|ens) counted every pod byte twice — once on the veth, once on the uplink — and attributed node traffic to no pod. Interfaces are rescanned every 5 seconds so newly created pod veths are picked up.

Direction normalisation. The TC hook reports its own perspective. On the host side of a veth pair that is inverted relative to the pod: a packet the pod sends arrives at the host end as ingress. Userspace flips it for veth interfaces so "egress" always means "the pod sent this".

Pod identity enrichment. A per-node informer maps ifindex to pod (with an IP fallback), attaching pod UID, namespace and name.

Deduplication. The same packet can be seen on both a veth and the node uplink; duplicates are suppressed over a 100 ms window.

Aggregation. Flows are aggregated over a 30 second window, then pushed to the server. By default (FLOW_DETAIL=conversation) every connection between one pod, peer and service port in a window becomes one record: bytes are summed, the client's ephemeral port is reported as 0, and a connections field counts how many connections were folded. No byte or connection is dropped. FLOW_DETAIL=connection keeps one record per 5-tuple, with the client port, at about 5–6× the record volume.

Deployment shape ​

yaml
securityContext:
  privileged: false
  capabilities:
    add: [BPF, PERFMON, NET_ADMIN]

Fine-grained capabilities rather than privileged: true on modern kernels. Older kernels need more: add SYS_RESOURCE below 5.11, SYS_ADMIN below 5.8.

mountPropagation

Use HostToContainer, not Bidirectional. Bidirectional is illegal without privileged: true and a manifest specifying it alongside fine-grained capabilities cannot be created at all. HostToContainer is sufficient: pinning maps into the host's existing bpffs is visible host-wide through the hostPath.

Requirements ​

  • Linux, kernel 6.6+ for TCX attachment
  • BPF, PERFMON, NET_ADMIN capabilities
  • kubespend-agent installed in the same cluster — it is the relay this agent reports through, not merely a companion
  • relay.service set to that agent's Service name; there is no default and the chart refuses to render without one
  • The same clusterId as that metrics agent: unset on both, or set to the same value on both. Otherwise the relay rejects every flow window
  • The org must have trial flow allowance left, or hold the network add-on
  • net.core.bpf_jit_enable=1, already the default on supported node images; an opt-in init container sets it where it is not

Install ​

relay.service is required in the default relay mode. It names the Service fronting the metrics agent's relay port, has no default, and the chart calls fail without it — so a missing value is a Helm template error, not something you discover from pod logs. Find the name first:

bash
kubectl get svc -n kubespend -l app.kubernetes.io/name=kubespend-agent

The metrics-agent chart names its Service after its release, so the documented release kubespend-agent produces a Service called kubespend-agent. A release whose name does not contain kubespend-agent gets <release>-kubespend-agent. For the documented release:

bash
helm install kubespend-ebpf-agent oci://public.ecr.aws/kubespend.io/charts/kubespend-ebpf-agent \
  --namespace kubespend \
  --set relay.service="kubespend-agent"

No API key. Set clusterId only if you set it on the metrics agent, and make the two match. See Installation.

Verify ​

bash
kubectl -n kubespend rollout status ds/kubespend-ebpf-agent
kubectl -n kubespend logs -l app.kubernetes.io/name=kubespend-ebpf-agent --tail=50

Healthy logs show interfaces being attached with their role (veth or node). A rising drop count in the logs means the node is producing events faster than the agent can drain them — usually a very busy node.

Troubleshooting ​

SymptomCause
attach ingress: not supportedkernel below 6.6; TCX unavailable
operation not permitted on loadmissing capabilities, or JIT disabled
flow window NOT delivered to relay … InvalidArgumentclusterId differs from the metrics agent's (it logs relay submission rejected: cluster mismatch); set the same value on both charts
PermissionDenied on a direct flow pushnetwork data is off for this org (trial allowance set to 0, no network add-on)
ResourceExhausted on a direct flow push, or flows stop on a trialthe trial flow-record allowance is used up; support can raise it
helm install fails, no pod is createdrelay.service is unset in relay mode — the chart calls fail at render time rather than shipping a DaemonSet with nowhere to dial
Relay dial fails / connection refusedrelay.service names a Service that does not exist, or kubespend-agent is not installed, not ready, or not reachable in this namespace
Relay rejects the tokenaudience mismatch — relay.tokenAudience must match what the metrics agent expects
Agent exits immediately on startmode: direct set without apiKey
Flows with no pod identityinterface not matched by the veth regex — check your CNI's naming
Byte counts roughly doubledveth and node regexes overlapping; verify both patterns
Pod cannot be created at allmountPropagation: Bidirectional without privileged

What it does not produce ​

Per-workload dollars. The agent's own output is bytes per pod, split by zone. Pricing happens server-side and lands at the cluster level, per day: cross-AZ transfer and internet egress are charged, same-AZ is free, and bytes whose peer had no resolvable zone are not priced at all — so the cluster network figure is a floor. NAT gateway and inter-region transfer are not modelled. See What we measure and AWS network costs.

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.