Skip to content

What is eBPF ​

eBPF lets you run small, verified programs inside the Linux kernel without writing a kernel module or rebooting. You attach a program to a hook — a packet arriving on an interface, a syscall, a function entry — and it runs every time that event fires.

If you have not met it before, ebpf.io and docs.ebpf.io are the canonical references. This page covers what matters for understanding KubeSpend.

Why it exists ​

Before eBPF, observing kernel behaviour left you two bad options:

Read from userspace. Poll /proc, scrape conntrack, sample with tcpdump. Cheap to build, but you see aggregates after the fact, you miss short-lived connections between samples, and copying every packet to userspace is expensive.

Write a kernel module. Full visibility, but a bug panics the machine, and you maintain a build per kernel version. Nobody wants a third-party module on every production node.

eBPF is the third option: kernel-level visibility with userspace-level safety.

The safety story ​

Every eBPF program passes a verifier before it loads. The verifier walks all possible execution paths and rejects the program unless it can prove:

  • it terminates — bounded loops only, no unbounded iteration
  • every memory access is in bounds, with pointer arithmetic checked
  • it only calls helper functions permitted for its program type
  • it fits within instruction and complexity limits

A rejected program does not load. A loaded program cannot panic the kernel, leak memory, or read outside what it was given. This is the property that makes running third-party code in the kernel acceptable on production nodes.

Programs are typically JIT-compiled to native instructions, so the runtime cost is a handful of instructions per event rather than a context switch.

Maps: how kernel and userspace talk ​

An eBPF program cannot make syscalls or write files. It communicates through maps — typed key/value structures the kernel program writes and a userspace process reads.

Map typeShapeTypical use
HASHkey → valueper-flow counters
PERCPU_ARRAYone slot per CPUlock-free counters
RINGBUFordered event streamper-event records to userspace
LRU_HASHbounded, evicts oldesttracking with a memory cap

Choosing between them is most of the design work in an eBPF tool. Per-CPU maps let the hot path avoid cross-CPU synchronisation entirely, at the cost of userspace summing the slots on read — a good trade for counters written constantly and read rarely.

Program types and hooks ​

The hook determines what a program can see. The ones relevant to networking:

HookWhereSees
XDPdriver, before skb allocationraw frames, earliest and fastest
TC (traffic control)after skb allocation, ingress and egressfull packet plus kernel metadata
cgroup/skbper cgrouptraffic for a process group
kprobe / tracepointkernel functionsinternal state, e.g. TCP events

KubeSpend attaches at TC, on both ingress and egress. TC is the right layer for cost accounting: it is late enough that the kernel has already associated the packet with an interface — which is what makes per-pod attribution possible — and it observes both directions, which XDP alone does not.

Attachment uses TCX, a newer link type (kernel 6.6+) that supports multiple programs on the same interface with defined ordering and clean detach. The older tc/clsact API works but makes it easy for two tools to fight over one hook.

eBPF in Kubernetes ​

You are probably already running it. Cilium implements pod networking and network policy with eBPF; recent kube-proxy modes use it instead of iptables; Falco, Pixie, Parca and most modern profilers are eBPF-based.

For a cost tool the appeal is specific: a pod's traffic is visible on its veth interface, and the kernel knows which interface a packet crossed. That gives per-pod byte accounting without touching the application, injecting a sidecar, or asking anyone to change their code.

What it costs to run ​

  • Linux only. No Windows nodes, no macOS.
  • Privileges. Loading programs and attaching to interfaces needs capabilities — BPF, PERFMON, NET_ADMIN. Fine-grained capabilities work on modern kernels; older ones need broader privileges.
  • Kernel version matters. TCX needs 6.6+. Helper availability varies across versions, which is why eBPF projects invest heavily in CO-RE ("compile once, run everywhere") and test across kernels.
  • Per-event overhead is small but not free. A program that emits one event per packet does more work than one that aggregates in a kernel map first.

Next ​

External references ​

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.