Skip to content

Rightsizing recommendations ​

Rightsizing compares what each workload requested against what it has actually used, and flags the ones where the gap is big enough to be worth acting on.

The signal ​

Usage is aggregated per snapshot first, then summarised across snapshots. Doing it in one pass would average across pods and time simultaneously, which flattens the peak: a workload whose replicas spike at different moments would look idle.

Requests come from the most recent snapshot, since that is the current spec. Averaging a request that changed mid-window would describe neither value.

Figures are per workload, not per container

Both the usage and the request are summed across the workload's pods at each snapshot. A 3-replica DaemonSet requesting 100m per container reports as 300m. The two sides share one denominator, so the ratios and the money are right — but the number shown is the workload's total, not a value to paste into a single container's resources.requests.

How a target is chosen ​

CPU and memory are sized differently, because they fail differently. Exceeding a CPU request throttles a container; exceeding a memory request kills it.

  • CPU is sized from a high percentile of observed usage, plus headroom. A percentile rather than the peak, so a single deploy-time spike cannot set a request for the rest of the week.
  • Memory is sized from the observed peak, plus a larger headroom, and is never proposed below that peak. A memory request below what a workload has already reached is an out-of-memory kill, not a saving.
  • Reductions are stepped. No single recommendation removes more than a fixed share of a request; the rest can follow once the first change has been observed to be safe.
  • Small savings are suppressed. A finding is raised only when the gap is large enough to be worth a deploy, which keeps the list short enough to act on.
  • Requests have a floor. Very small requests are not shrunk further, because below a certain size the measurement itself is noise.

CPU is never reduced for a workload that already wants more ​

If a workload's CPU regularly exceeds what it requests, no CPU reduction is proposed however comfortable the typical figure looks — the reservation is not too large, the workload routinely wants more than it asks for. Memory is a separate decision about the same workload and may still be proposed.

A CPU request is a scheduling reservation and a CFS weight, not a limit. A pod with a small request still bursts to whatever its limit allows; what it gives up is relative priority when a node is genuinely saturated.

Turning a reclaimed request into dollars ​

The console calls this figure Reclaimable rather than a saving, for the reason the warning below gives.

monthlyCPU = (savedCores × costPerCoreHour) × 730
monthlyMem = (savedGiB   × costPerGiBHour)  × 730
annual     = monthly × 12

A node's price has to be split between its CPU and its memory first, and the split is AWS's own published split cost allocation weighting — 9 per vCPU to 1 per GiB — rather than an even halving:

perVCPU = price × 9 / (9 × vCPU + GiB)
perGiB  = price × 1 / (9 × vCPU + GiB)

So on a 2 vCPU / 8 GiB node the denominator is 26, and 69% of the price is attributed to CPU and 31% to memory. On a memory-optimised shape it moves the other way, which an even split could not express. The two components always sum back to the node's actual price.

The per-core and per-GiB rates are averaged across the cluster's nodes, weighted by how long each node ran. So a cluster that is mostly cheap spot capacity produces smaller savings figures than one on on-demand, correctly.

This is a ceiling, not a guarantee

A reclaimed request only becomes money when it lets the cluster run fewer nodes. Shrinking requests on a fleet that stays the same size saves nothing directly — it creates the headroom that makes consolidation or downsizing possible. Treat the figure as the upper bound of what is unlocked.

If no node in the cluster can be priced, findings still ship — with no dollar figures attached, rather than with fabricated ones.

The four finding types ​

TypeRaised whenDollars?
Rightsize CPUthe CPU request is well above what the workload uses, with headroomYes
Rightsize Memorythe memory request is well above the observed peak, with headroomYes
Idle Workloadthe workload used no CPU at all over the windowYes
Missing Requestsno CPU and no memory request setNo

Idle Workload is the CPU rule with a different label, applied when the workload used no CPU at all.

Missing Requests carries no dollar estimate — there is nothing to reclaim — and is marked non-actionable. It is still the highest-risk configuration in a cluster: without requests the scheduler cannot place the pod sensibly and it is first to be evicted under pressure.

Impact and confidence ​

Memory findings report a higher impact than CPU by design. Exceeding a memory limit is an immediate OOM kill; exceeding CPU only throttles.

Confidence is a function of window length only:

WindowConfidence
≥ 7 daysHigh
≥ 1 dayMedium
< 1 dayLow

Kubernetes workloads have daily and weekly cycles, so a few hours of data cannot distinguish "idle" from "between peaks". Short windows still produce recommendations; they are just labelled Low.

A caution on the total ​

The report's headline monthly savings figure sums workload findings, instance recommendations, and consolidation plans. These overlap by construction — rightsizing a workload and consolidating the node it runs on can describe the same dollar twice. Read individual findings as reliable and the grand total as an optimistic upper bound.

Next ​

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.