Skip to content

Coverage and confidence ​

Every cost response carries a coverage object. It answers a question most cost tools leave unanswered: how much should you trust this number?

json
{
  "coverage": {
    "totalNodeHours": 96,
    "pricedNodeHours": 96,
    "observedHours": 23,
    "expectedHours": 24,
    "observedRatio": 0.9583,
    "partial": false,
    "complete": true,
    "unpriced": [],
    "networkPricedBytes": 41203847,
    "networkFreeBytes": 918273645,
    "networkUnclassifiedBytes": 302847561
  },
  "pricingAvailable": true
}

The two independent failure modes ​

A cost figure can be wrong for two unrelated reasons, and coverage tracks them separately.

Something could not be priced. unpriced lists the instanceType/capacityType labels whose rate could not be resolved. Those node-hours are excluded from the total rather than estimated, so the total understates. pricedNodeHours versus totalNodeHours quantifies how much.

Common causes: a region outside the supported set, a brand-new instance type not yet in the Price List response, spot history returning nothing for an AZ, or a node group missing region/instance-type metadata.

The agent was not reporting. observedHours against expectedHours measures how much of the window actually produced data.

expectedHours = (days − 1) × 24 + (currentHour + 1)
observedRatio = min(1, observedHours / expectedHours)
partial       = observedRatio < 0.95

observedHours counts distinct clock hours that produced data, independent of node count — which distinguishes "one node ran for nine hours" from "four nodes ran but we only watched nine hours of the day".

Why the threshold is 95%, not 100% ​

An agent restart, a rollout, or a scrape that lands a second late will always leave a hole somewhere. Flagging every window as partial would make the flag meaningless. 95% tolerates ordinary operational noise while still catching the case that matters: an agent that was down for hours.

When partial is true the server also logs it, because the reported cost understates actual spend and that is worth knowing before someone quotes the figure in a meeting.

complete requires both halves ​

complete = (unpriced is empty)
           AND (pricedNodeHours ≥ totalNodeHours)
           AND (NOT partial)

Both conditions are necessary. Checking only the pricing half would let a cluster observed for two hours of a 24-hour day report itself complete.

pricingAvailable ​

Sits outside coverage because it is a different class of problem: the server has no AWS pricing credentials configured, so no figure in the response can be derived. The console renders this as unavailable rather than drawing a $0 chart. It applies to self-hosted deployments; SaaS is configured already.

Recommendation confidence ​

A separate axis, based only on how much history the window contains:

WindowConfidence
≥ 7 daysHigh
≥ 1 dayMedium
< 1 dayLow

Node consolidation plans are capped at Medium regardless of window, because the simulation omits scheduling constraints that can invalidate an arithmetically-valid plan. See Node optimization.

Practical reading guide ​

SymptomLikely cause
Cost lower than the AWS billpartial: true, or non-empty unpriced
Everything is unavailablepricingAvailable: false — server pricing credentials
One node group missingcheck unpriced for its instance type
Total ≠ sum of daily rowsrounding is applied once on exit, from a full-precision accumulator
Savings total looks too goodthe grand total sums overlapping recommendations; read it as an upper bound

Network coverage: a third failure mode ​

Network cost has its own way of being incomplete, so it gets its own counters rather than folding into the node-hour ones.

FieldMeaning
networkPricedBytesbytes that carried a charge — cross-AZ plus internet egress
networkFreeBytessame-AZ bytes, correctly worth $0
networkUnclassifiedBytesbytes where one end had no resolvable zone, so they were not priced

Unclassified is the one that changes how the figure reads. Internet peers, the EKS control plane and the API ClusterIP have no zone in the cluster's pod or node inventory, and the classifier refuses to assume same-AZ for an unknown end. Those bytes are neither charged nor called free.

A large unclassified share means the network cost is a floor, not a total. Compare it against networkPricedBytes before quoting the number. Two further gaps sit outside coverage entirely: NAT gateway and inter-region transfer are not modelled, and network dollars are cluster-level daily only — there is no per-workload network cost for coverage to qualify. See What we measure.

Every figure in KubeSpend traces to a real cloud price. Where we cannot measure something, we say so.