Capabilities and privileges
The eBPF agent asks for four Linux capabilities and, on most clusters, only three of them. This page says what each one grants, why the agent needs it, and — just as importantly — what it deliberately does not ask for.
If you are being asked to approve this DaemonSet, this is the page to read.
The short version
securityContext:
runAsUser: 0
privileged: false # never, for the agent container
capabilities:
add: [BPF, PERFMON, NET_ADMIN]Three capabilities cover network flow accounting and cross-AZ measurement, which is what most installations run. A fourth, SYS_PTRACE, is needed only if you turn on Go/TLS query capture for the DB Optimizer.
privileged: true is not used. That distinction is the whole point of the list: a privileged container holds every capability plus an unconfined view of the host, so enumerating three is a materially smaller grant than the usual "eBPF needs privileged" shortcut.
What each capability grants
CAP_BPF
Permits the bpf() syscall — loading a program, creating maps, and pinning both into bpffs.
Before kernel 5.8 there was no such capability and all of this lived under CAP_SYS_ADMIN. CAP_BPF exists precisely so a program loader does not have to be granted full administrative power over the host, and it is the reason this agent can run unprivileged at all.
What it does not grant: it does not let the agent attach programs to network interfaces, and it does not let it read another process's memory.
CAP_PERFMON
Permits performance-monitoring access: reading kernel memory from a BPF program (bpf_probe_read_kernel), and creating the links that carry events out of the kernel.
Paired with CAP_BPF by design — the kernel's bpf_capable() and perfmon_capable() checks are separate, and a loader that can create a map but not read anything is not useful. It replaces the observability half of what CAP_SYS_ADMIN used to cover.
What it does not grant: it is read-only with respect to observation. It confers no ability to modify traffic or to write to kernel memory.
CAP_NET_ADMIN
Permits network administration on the interfaces the agent selects — specifically, attaching the traffic-control (TC) programs that do the byte accounting, via TCX links.
This is the capability that sounds the most alarming and is the most bounded in practice. The agent attaches TC programs to the ingress and egress paths of pod veths and node uplinks. Those programs count and pass: they never modify, drop, redirect or delay a packet. There is no bpf_redirect, no TC_ACT_SHOT, no header rewriting anywhere in the agent's kernel code — the only verdict it returns is "continue".
What it does not grant: it is scoped to network configuration. It does not permit loading arbitrary programs, reading process memory, or filesystem access.
CAP_SYS_PTRACE — only for Go/TLS query capture
Needed only when the DB Optimizer's uprobe path is enabled. Leave it off and the capability is unnecessary — but so is query data for every encrypted database, which in practice is most of them.
RDS instances with rds.force_ssl=1 and ElastiCache clusters with in-transit encryption both put ciphertext on the wire, so there is nothing for a TC program to read — a statement is only visible before TLS, inside the client process. This capability is what makes those databases legible at all, and it is also what produces per-query latency for them: PostgreSQL timed at the client library call enclosing each query, Redis by pairing the encrypted write with the read that answers it. The agent finds that point by reading the client binary's symbols, which means reading /proc/<pid>/exe and /proc/<pid>/root/... for processes in other containers.
That read goes through the kernel's ptrace access-mode check, and being uid 0 is not enough to pass it across a user namespace. Measured on a single node:
| Configuration | Executables readable |
|---|---|
hostPID + root, no extra capability | 13 of 188 |
hostPID + root + CAP_SYS_PTRACE | 186 of 186 |
The failure mode is why it matters: without the capability the scan does not error, it just sees about 7% of the node — and the application containers are not in that 7%. The result looks like "no databases here to instrument" rather than like a misconfiguration.
What it does not grant: SYS_PTRACE permits inspecting processes. The agent uses it to read symbol tables from executables, which is the narrowest possible use of it. It does not attach a debugger, does not write to any process, and does not read application heap memory.
What the agent deliberately does not ask for
CAP_SYS_ADMIN
Effectively root-equivalent, and refused on modern kernels.
It would be required if uprobes were attached through perf_event_open, which is the obvious way to do it and the way the agent originally did it. perf_event_open uprobe creation is gated on CAP_SYS_ADMIN regardless of CAP_PERFMON. Rather than document that as a requirement, the attach was moved to UprobeMulti — bpf(BPF_LINK_CREATE), the same syscall that already loads the program — which needs CAP_BPF + CAP_PERFMON and a kernel of 6.6 or newer.
This was confirmed by removing SYS_ADMIN entirely and checking that probes still attached and statements still decoded.
A third-party cost agent should not be asking for SYS_ADMIN on your production cluster. If a future feature appears to need it, that is a design to revisit, not a permission to request.
privileged: true
Not used for the agent container on any supported kernel.
There is one exception, and it is worth naming rather than glossing: the optional JIT init container, off by default. net.core.bpf_jit_enable is not a namespaced sysctl, so it can only be set on the host, and that init container runs privileged for the few milliseconds it takes to write one value. It runs busybox (pinned by digest), executes a single sysctl -w, and exits before the agent starts. It is also the only image in the chart that carries a shell, which is why it is opt-in.
It is off because the kernels this agent supports (6.6+) already run with the JIT on in the common node images (EKS AL2023, Bottlerocket, Ubuntu and COS, most of which build with BPF_JIT_ALWAYS_ON). Check a node:
sysctl net.core.bpf_jit_enable # 1 (or 2) means the JIT is onOnly if it reads 0, set it during node bootstrap (EC2 user data, tuned, or your AMI), or enable the init container:
enableJitInitContainer: trueThe JIT is a performance measure, not a correctness one, but running without it costs real CPU on busy nodes.
Host-level settings, and why
Two pod-level settings are as significant as the capabilities and are easy to miss:
| Setting | Why | Consequence of omitting it |
|---|---|---|
hostNetwork: true | pod veths exist in the host network namespace; TC programs must attach there | no interface to attach to |
hostPID: true | uprobe discovery must see other pods' processes in /proc | only the agent's own processes are visible |
hostPID is only used by the uprobe path, but the chart sets it unconditionally — there is no value to turn it off. SYS_PTRACE is the part you add or omit.
Go/TLS query capture is amd64 only
The uprobe path is built for linux/amd64. On arm64 nodes — Graviton — it is compiled out entirely, so SYS_PTRACE buys nothing there and encrypted-wire query capture is unavailable. Flow accounting and cross-AZ measurement work on both architectures.
mountPropagation: HostToContainer, not Bidirectional
Bidirectional is illegal without privileged: true — the API server rejects the pod outright — so a manifest combining it with fine-grained capabilities cannot be created at all.
HostToContainer is sufficient. Bidirectional exists for a container that creates new mounts the host must see; this one only pins maps and programs as files inside the bpffs the host already has mounted, which is visible host-wide through the shared hostPath either way.
Kernel version conditions
The capability set depends on your kernel, because the capabilities themselves were introduced over several releases.
| Kernel | Required set |
|---|---|
| 6.6 and newer | BPF, PERFMON, NET_ADMIN — TCX attach and uprobe.multi both available |
| 5.11 – 6.5 | the same three, but TCX is unavailable, so flow capture will not attach |
| 5.8 – 5.10 | add SYS_RESOURCE (the memlock limit was not yet accounted to the BPF memcg) |
| below 5.8 | CAP_BPF and CAP_PERFMON do not exist; requires SYS_ADMIN |
6.6 is the practical floor. TCX is how the agent attaches, and uprobe.multi is how it attaches uprobes without SYS_ADMIN; both landed in 6.6. Below that the agent reports the feature as unavailable rather than falling back to a higher privilege — see What we measure.
Amazon Linux 2023 on current EKS ships 6.12, so this is normally satisfied without action. Check with:
kubectl get nodes -o wideVerifying what the agent actually holds
Do not take a manifest's word for it — a securityContext can be overridden by a policy admission controller without the manifest changing. Read the effective set from the running process:
# From any pod on the node, or via kubectl debug:
grep Cap /proc/<agent-pid>/statusCapEff is a bitmask. Bit 38 is PERFMON, bit 39 is BPF. Also worth checking Seccomp: 0 and Uid: 0 in the same file — a seccomp profile can block a syscall the capability permits, which presents identically to a missing capability.
readlink /proc/*/exe needs privilege to find the process, but comm and status are world-readable, which is how to locate it from an unprivileged debug pod.
Summary
| Capability | Needed for | Optional? |
|---|---|---|
CAP_BPF | loading programs and maps | no |
CAP_PERFMON | reading kernel memory from BPF, event links | no |
CAP_NET_ADMIN | attaching TC programs to interfaces | no |
CAP_SYS_PTRACE | reading client binaries for Go/TLS query capture | yes — off unless you enable uprobe capture |
CAP_SYS_ADMIN | — | never requested on kernel 6.6+ |
privileged: true | — | never for the agent; optional JIT init container only |
See also The eBPF agent for what it observes, and Data retention for what leaves your cluster.