Coverage and confidence
Every cost response carries a coverage object. It answers a question most cost tools leave unanswered: how much should you trust this number?
{
"coverage": {
"totalNodeHours": 96,
"pricedNodeHours": 96,
"observedHours": 23,
"expectedHours": 24,
"observedRatio": 0.9583,
"partial": false,
"complete": true,
"unpriced": [],
"networkPricedBytes": 41203847,
"networkFreeBytes": 918273645,
"networkUnclassifiedBytes": 302847561
},
"pricingAvailable": true
}The two independent failure modes
A cost figure can be wrong for two unrelated reasons, and coverage tracks them separately.
Something could not be priced. unpriced lists the instanceType/capacityType labels whose rate could not be resolved. Those node-hours are excluded from the total rather than estimated, so the total understates. pricedNodeHours versus totalNodeHours quantifies how much.
Common causes: a region outside the supported set, a brand-new instance type not yet in the Price List response, spot history returning nothing for an AZ, or a node group missing region/instance-type metadata.
The agent was not reporting. observedHours against expectedHours measures how much of the window actually produced data.
expectedHours = (days − 1) × 24 + (currentHour + 1)
observedRatio = min(1, observedHours / expectedHours)
partial = observedRatio < 0.95observedHours counts distinct clock hours that produced data, independent of node count — which distinguishes "one node ran for nine hours" from "four nodes ran but we only watched nine hours of the day".
Why the threshold is 95%, not 100%
An agent restart, a rollout, or a scrape that lands a second late will always leave a hole somewhere. Flagging every window as partial would make the flag meaningless. 95% tolerates ordinary operational noise while still catching the case that matters: an agent that was down for hours.
When partial is true the server also logs it, because the reported cost understates actual spend and that is worth knowing before someone quotes the figure in a meeting.
complete requires both halves
complete = (unpriced is empty)
AND (pricedNodeHours ≥ totalNodeHours)
AND (NOT partial)Both conditions are necessary. Checking only the pricing half would let a cluster observed for two hours of a 24-hour day report itself complete.
pricingAvailable
Sits outside coverage because it is a different class of problem: the server has no AWS pricing credentials configured, so no figure in the response can be derived. The console renders this as unavailable rather than drawing a $0 chart. It applies to self-hosted deployments; SaaS is configured already.
Recommendation confidence
A separate axis, based only on how much history the window contains:
| Window | Confidence |
|---|---|
| ≥ 7 days | High |
| ≥ 1 day | Medium |
| < 1 day | Low |
Node consolidation plans are capped at Medium regardless of window, because the simulation omits scheduling constraints that can invalidate an arithmetically-valid plan. See Node optimization.
Practical reading guide
| Symptom | Likely cause |
|---|---|
| Cost lower than the AWS bill | partial: true, or non-empty unpriced |
| Everything is unavailable | pricingAvailable: false — server pricing credentials |
| One node group missing | check unpriced for its instance type |
| Total ≠ sum of daily rows | rounding is applied once on exit, from a full-precision accumulator |
| Savings total looks too good | the grand total sums overlapping recommendations; read it as an upper bound |
Network coverage: a third failure mode
Network cost has its own way of being incomplete, so it gets its own counters rather than folding into the node-hour ones.
| Field | Meaning |
|---|---|
networkPricedBytes | bytes that carried a charge — cross-AZ plus internet egress |
networkFreeBytes | same-AZ bytes, correctly worth $0 |
networkUnclassifiedBytes | bytes where one end had no resolvable zone, so they were not priced |
Unclassified is the one that changes how the figure reads. Internet peers, the EKS control plane and the API ClusterIP have no zone in the cluster's pod or node inventory, and the classifier refuses to assume same-AZ for an unknown end. Those bytes are neither charged nor called free.
A large unclassified share means the network cost is a floor, not a total. Compare it against networkPricedBytes before quoting the number. Two further gaps sit outside coverage entirely: NAT gateway and inter-region transfer are not modelled, and network dollars are cluster-level daily only — there is no per-workload network cost for coverage to qualify. See What we measure.