Skip to content

Node & instance optimization ​

Rightsizing changes what workloads ask for. This changes the fleet they run on. Both prices are looked up live, so every saving here is a real delta between two real AWS list prices — not a percentage rule of thumb.

Every recommendation on this page uses the same arithmetic:

currentMonthly = currentHourlyPrice   × 730 × nodeCount
newMonthly     = candidateHourlyPrice × 730 × nodeCount
savings        = currentMonthly − newMonthly     (suppressed if ≤ 0)
savingsPercent = savings / currentMonthly × 100

The candidate price is fetched from the Price List API exactly like the current one. If the candidate cannot be priced, no recommendation is emitted.

Instance downsize ​

Fires when a node group's sustained utilization is low on both CPU and memory, recommending one size step down within the same family.

Both dimensions, not the lower of the two: a group at 20% CPU and 70% memory does not fit on half the memory, and taking the friendlier figure is how a downsize recommendation causes an out-of-memory kill on a node rather than in a pod. An unmeasured dimension is not a low one — a group whose utilization could not be measured is never offered a downsize.

Measured per group, over the window

Each group's utilization is measured from its own nodes across the observation window. It was previously a cluster-level figure apportioned across groups by capacity ratio, which gave every group of the same shape an identical number and described the cluster wearing a per-group label.

A cheaper shape still has to hold the pods ​

Utilization decides whether a downsize is worth offering; requests decide whether it is possible. Kubernetes schedules on requests, so a smaller instance that cannot hold what is already placed on a node does not run hotter — it leaves pods Pending.

Every candidate is therefore checked against the group's busiest node, on both CPU and memory, with headroom on the request side and a margin on the capacity side for the kubelet's own reservation (which grows with instance size and which the price table does not report). A candidate whose shape cannot be resolved at all is refused rather than assumed — an unchecked shape is how a group gets moved onto an instance with half the memory.

This is why a cheaper instance is sometimes correctly not offered: a smaller shape can be a third cheaper and still be refused, because the requests already placed on the busiest node would not fit on it.

Candidate sets, not fixed targets ​

Both of the next two recommendations work the same way. KubeSpend does not have a "correct" instance family in mind. It assembles a candidate set, prices every member against the live Price List for your region and size, and recommends the cheapest one that actually beats what you pay now.

That matters because newer is not cheaper on AWS:

InstanceArch$/hr (us-east-1, .large)
m6gGraviton20.0770
m7gGraviton30.0816
m6aAMD0.0864
m7aAMD0.1159

Graviton3 costs more than Graviton2, and m7a is about 34% more expensive than m6a. AWS prices newer silicon for performance, not for cost. A rule that simply moved you to the current generation would raise your bill.

Candidates always stay within a family class (m→m, c→c, r→r) so the vCPU-to-memory ratio is preserved. A candidate that cannot be priced in your region is skipped rather than guessed at — which is how families with partial regional rollout, like m8g, are handled.

Generation optimization (same architecture) ​

The cheapest family that keeps your current architecture. For x86 nodes that means the AMD variants; for Graviton nodes, a cheaper Graviton generation.

m5 → m6a is roughly 10% at zero architecture risk — same x86_64, no image rebuild, no dependency audit. If your workloads are pinned to x86 by native dependencies, vendor images, or licensing, this is the only instance-level saving available to you, and it is worth about as much as Graviton would be.

This is reported separately from Graviton migration precisely because the risk is different. Same confidence, far less work.

Graviton migration ​

Offered only for x86 nodes, targeting the cheapest priceable ARM candidate.

The saving depends heavily on what you are coming from, which is why there is no single headline percentage:

FromTo cheapest ARMSaving
m5 (Intel)m6g~19.8%
m5a (AMD)m6g~10.5%
m6a (AMD)m6g~10.9%

AMD instances are already discounted against Intel, so the Graviton prize from an AMD baseline is roughly half what it is from Intel. Any tool quoting a flat "Graviton saves ~20%" is quoting the Intel case and ignoring the rest.

Check your images first

Graviton is a different CPU architecture. Every container image on these nodes needs an arm64 build, and anything with compiled native dependencies or an x86-only base image will not start. KubeSpend prices the move; it cannot tell you whether your images are multi-arch.

Performance is not modelled

The comparison is price at the same shape. Graviton3 and Graviton4 deliver meaningfully more per vCPU than Graviton2, so if you intend to downsize after migrating, a newer generation may win overall despite the higher hourly rate. KubeSpend will not apply a performance multiplier to guess at that — it would be a fabricated number, and the right figure depends on your workload.

Family optimization ​

Classifies a node group's usage shape and suggests a better-matched family. A group that uses CPU far more heavily than memory gets compute-optimised (c) candidates; one that uses memory far more heavily than CPU gets memory-optimised (r) candidates.

A general-purpose node running genuinely balanced work produces no finding. The candidate set for the target class is then priced exactly as above, and the cheapest option that beats the current rate wins.

Candidates here keep your current architecture. Switching family and architecture in one recommendation would bundle two unrelated risks into a single dollar figure you could not decompose — so an x86 node gets x86 candidates, and the Graviton decision stays its own finding.

Spot conversion ​

Offered only for ON_DEMAND node groups, priced as the same instance type at current spot rates for the AZ.

Spot is an availability tradeoff

Spot capacity can be reclaimed with two minutes' notice. It suits stateless, replicated, interruption-tolerant workloads. The savings figure says nothing about whether your workload tolerates eviction, and the price itself moves — the quote is a point-in-time sample, cached for 15 minutes.

Node consolidation ​

The most involved analysis: can the same pods fit on fewer nodes?

A packing simulation tries to empty the most expensive nodes onto the rest of the fleet, using the pods' requests.

monthlySavings = Σ(hourly price of removed nodes) × 730

Two constraints always hold: destinations are never packed to 100% of allocatable, so a rolling update still has room for its surge pod, and the cluster never drops below two nodes, because a single node has no failure domain.

Excluded from the move: DaemonSet and Node-owned pods (immovable by definition), and Succeeded/Failed pods (terminal). Excluded as destinations: cordoned and NotReady nodes. Packing is against allocatable, not capacity.

Confidence is capped at Medium, deliberately

The simulation does not model pod affinity and anti-affinity, topology spread constraints, taints and tolerations, PodDisruptionBudgets, or volume zone affinity. A plan that is arithmetically sound can still be unschedulable. Treat it as a candidate to validate, never as a script to run.

A separate zero-dollar informational finding reports fragmentation: enough total slack across the fleet to drop a node, but too scattered to consolidate.

Ordering ​

Recommendations are sorted by monthly savings, descending. The same overlap caveat from rightsizing applies: downsizing a node and consolidating it away are alternatives, not additive.

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.