How the CostBench Score Works — CostBench
How the CostBench Score Works — CostBench
Shortlist

Methodology

How the CostBench Score works

Every product page carries a number from 0 to 100. It rates our record — how completely and how recently we can answer "what will this actually cost?" — and nothing else.

What it is not

It is not a rating of the software. We do not review products, run scored bake-offs, or aggregate other people's reviews, so a number implying we had judged the product would be a claim we cannot source. A product with a 94 is not better than one with a 71 — we simply know more about what the first one costs.

Payment never touches this score

Vendors cannot buy the score, raise it, or have it removed. Claiming a listing does not change it. Subscribing to CostBench Sync does not change it. Affiliate relationships do not change it. Not one of the inputs below is something money can move, and a test in our build asserts that adding every paid and claimed field to a record leaves its score byte-for-byte identical. Corrections are free for everyone, forever.

The five components

The weights sum to 100 and are read from the same records this page describes, so they cannot drift from the scorer that produced them.

Price transparency

30 pts

Whether the product publishes real numbers. Concrete tier prices earn full credit. "Contact us" earns partial credit — enterprise-quote-only is an honest answer for some products, and we score our record, not the vendor's choice. Usage-priced products (per token, per image, per GPU-hour) count as published: a rate card is a price.

Record confidence

25 pts

Our extraction confidence, tempered by how many independent sources agreed. Confidence alone is self-reported by the pipeline; pairing it with corroboration stops one confidently-parsed bad page from scoring like a triple-sourced one.

Freshness

20 pts

How recently we re-verified against the vendor's own pricing page. Flat inside the re-verification window, then decaying to zero at a year. SaaS pricing moves constantly — a two-year-old record is not slightly stale, it is a different price.

Cost completeness

15 pts

Whether we can show the gap between sticker price and real spend: hidden costs, contract terms, total-cost scenarios, and negotiation levers. Each is a distinct question buyers ask, so each is credited separately.

Market grounding

10 pts

Evidence that the range reflects what buyers actually paid rather than what the vendor advertises. Sample size matters — a benchmark drawn from 2,000 purchases is not the same claim as one drawn from three.

Bands

BandScoreMeaningProducts
Complete 85+ Nearly every cost question is answered and recently checked. 436
Strong 70+ The important numbers are there; something secondary is thin. 1,385
Partial 50+ Usable prices, but real gaps — often no market data. 1,447
Thin 30+ Enough to orient you, not enough to budget from. 37
Sparse 0+ We do not yet know enough. Shown so the gap is visible, not hidden. 0

Where our corpus actually stands

Across 3,305 published products the mean score is 71.1 and the median is 71. Below is the share of each component we actually earn — including the one we are worst at.

ComponentAvg earnedOfShare
Price transparency 23.9 30 79.7%
Record confidence 15.8 25 63.2%
Freshness 18.6 20 93.0%
Cost completeness 9.5 15 63.3%
Market grounding 3.2 10 32.0%

Market grounding is our weakest component and we would rather say so here than quietly weight it out of the score. It is low because independent transaction benchmarks exist for under half the catalogue — closing that gap is the main way these numbers rise. Scores last computed 2026-08-08.

Correcting a score

The score is arithmetic over the underlying record, so it has no separate appeal process: fix the record and the number follows on the next run. If a price, tier, or contract term on your product is wrong, tell us and we will correct it free, whether or not you are a customer.