Our Methodology
How we research pricing data
Accurate pricing intelligence requires rigorous methodology. Here's exactly how we collect, verify, and maintain our database of software pricing.
A different kind of pricing database
Traditional software pricing research runs on analysts and sales calls — it's slow, it's biased toward vendors who'll take the call, and it goes stale between reports. We built CostBench on a different model: a proprietary agent pipeline that monitors pricing pages continuously, triangulates across buyer communities, review platforms, and contract benchmarks, and surfaces changes the moment they happen.
Every price is timestamped, every claim is sourced, and every data point carries a confidence score you can see. When our agents aren't certain, we say so. When vendors want to correct their own data, they can — through a verified channel that doesn't touch our editorial rankings. 875 products now carry negotiated-deal benchmark coverage on top of verified list pricing.
Data Sources
Official Pricing Pages
Every product's public pricing page is our primary source — captured on a rolling cadence so list prices, tier structures, and feature breakdowns stay current.
Confidence: HighContract Benchmarks
Median contract costs, typical discount ranges, and deal volumes sourced from established procurement benchmark providers. Gives you a view of what buyers actually pay, not just list price.
Confidence: HighVendor Corrections
Verified product owners can claim their profile and flag outdated data through our vendor portal. Corrections are reviewed and applied without changing our editorial rankings.
Confidence: HighReview Platforms
Pricing-related signals from major verified-review platforms — hidden fees, surprise renewals, and satisfaction trends flagged by actual users.
Confidence: Medium-HighBuyer Communities
Public procurement discussions, hidden-fee reports, and negotiation playbooks shared by buyers in trusted online communities.
Confidence: MediumComplaint & Dispute Records
Pricing disputes and billing complaints from public consumer-protection channels — a signal for surprise fees, auto-renewal traps, and post-sale cost escalation.
Confidence: MediumHow a price actually gets into the database
The short version: a scheduled job opens the vendor's own pricing page in a real browser every day, reads the plan table, checks the result against the record it would replace, and only then writes. Here is each step.
- 1
A daily batch is selected, oldest record first
A scheduled job runs every day and sorts the catalogue purely by how long ago each record was last verified, then takes the next batch. There is no popularity weighting — the least recently checked product goes first, so coverage stays even. 3,302 of 3,305 products carry a stored pricing-page URL and are therefore eligible.
- 2
The vendor page is rendered, not just fetched
Each pricing page is loaded in a real headless Chrome session, because most modern pricing pages build their plan tables in JavaScript and a plain HTTP fetch returns an empty shell. Where a page has a monthly/annual billing toggle, the extractor finds it, clicks through both states, and records prices for each.
- 3
Plans are extracted and normalised
Plan names, prices, billing units and limits are read from the rendered page. Currencies are recorded, and plans with no public number are flagged as custom-priced rather than silently dropped — that flag is what lets us count quote-only pricing instead of mistaking it for missing data.
- 4
Every extraction is checked against the record it would replace
Before anything is written, the proposed record is diffed against the current one. A price move over 50%, a product that appears to have lost its free plan, or a plan lineup that gains or loses two or more tiers all raise flags. Low-confidence or multiply-flagged extractions are held for review instead of published, and an empty scrape can never overwrite a verified price.
- 5
The record is stamped and snapshotted
Surviving changes update the product record with a new verification date and confidence score, and a dated snapshot of the plan structure is appended to its price history. Those snapshots are what a later re-check is compared against — they are not published as a trend line, for the reason set out under Limitations.
Sample sizes and update cadence
Every figure below is computed from the live database when this page is built, so it cannot drift out of date. If you are citing an aggregate from CostBench, these are the denominators behind it.
Across 255 categories.
The denominator for every price aggregate we publish.
Plans exist, but the vendor publishes no number on any of them.
Days since the typical product was last re-checked.
634 were re-checked within the last 30.
Each is re-read from that page; the URL ships in every CSV row.
Aggregates are recomputed and republished quarterly — see Data reports. The underlying corpus is downloadable as CSV at /data/, with the source URL and verification date on every row.
Limitations — what this data cannot tell you
Anything worth citing is worth bounding. These are the constraints we would want stated if we were the ones repeating our numbers.
These are list prices, not what buyers pay
We read the vendor's published price. Negotiated enterprise contracts routinely land below it, and the gap widens with deal size. Treat our figures as the asking price, not the clearing price.
We do not publish a price-trend series
This is deliberate, and it is the most important caveat on the page. Detecting price changes by diffing snapshots of same-named plans produces mostly false positives: a promotional rate read as list, an annual rate compared against a monthly one, a plan renamed or restructured, regional price variation, per-seat versus per-account billing. When we hand-checked all 16 candidate changes from our first index against the vendors' own records, 12 were artifacts and only 2 were real. So we publish individual changes only after a human has verified them against the vendor's own record, and we publish no trend chart at all. Any source claiming precise SaaS inflation rates from scraped pricing pages is almost certainly reporting measurement noise.
Roughly a third of the market publishes nothing
1,038 products are quote-only, so every aggregate we publish describes the priced subset. That subset skews toward self-serve and mid-market software; enterprise categories are systematically under-represented in any price average, ours included.
Prices are read from a single region
Extraction runs from European infrastructure. Vendors that vary price by geography may therefore be recorded at a non-US rate. We flag currency where the page states it, but regional drift is a known source of error on a minority of records.
Freshness varies by product
The queue works oldest-first, so a long-tail product may sit at the back of the
cycle. Every record carries its own lastVerified date — use it
rather than assuming a uniform refresh, and prefer records verified within 90
days for anything load-bearing.
A confidence score is a summary, not a guarantee
Confidence reflects source count, recency and internal consistency. A high-confidence record can still be wrong if the vendor's own page was wrong or ambiguous when we read it. Corrections are welcome and are applied without changing editorial rankings.
Source Taxonomy & Licensing
"Verified" doesn't mean "ours to license out." Every price on CostBench falls into
one of four buckets below, and every field in our API carries the matching
license status ("owned" or "restricted") so
agents and integrators know what they can build on. Buckets aren't mutually
exclusive — a product can carry a vendor-public list price and a Vendr-licensed
market median at the same time.
Tiers and price ranges extracted directly from the vendor's own public pricing page. We do the extraction and verification; free to cite with attribution.
Cost bands for quote-only vendors, mined from 3+ independent third-party mentions by our own grounded-search pipeline — not purchased from a data broker.
Cost mentions aggregated from public buyer discussion (Reddit, Hacker News). Indicative only — labeled as such wherever it's shown.
Market medians licensed from Vendr for display on our pages. Not licensed for commercial reuse or resale — flagged license: "restricted" in every API response that includes it.
API license contract and free-tier terms: /api-docs. The downloadable CSV exports at /data/ carry only owned data and are published under CC BY-NC 4.0 — quoting or citing individual figures from this dataset — in articles, reports, decks, or ai assistant answers, commercial or not — is always permitted with attribution to costbench.
Confidence Scores
Not all pricing data is created equal. A price confirmed by multiple buyer contracts is more reliable than one scraped from an outdated pricing page. We quantify this uncertainty with confidence scores.
Every price in our database has an associated confidence score. This tells you how much you can trust that specific data point and helps you make better decisions.
Verified by multiple independent sources including actual buyer contracts
Confirmed by official sources and at least one independent verification
Based on official sources but not independently verified recently
Estimated or based on limited/outdated information—marked clearly
How We Stay Current
Continuous Monitoring
Our agent pipeline revisits every product on a rolling ~30-day cadence — not quarterly, not annually. New pricing surfaces in days, not months.
Intelligent Change Detection
When a pricing page changes, we re-verify. When it doesn't, we don't waste compute revalidating identical data. The result: faster, cheaper refresh than a manual review team could match.
Multi-Source Cross-Check
Enriched records cross-reference contract benchmarks, review platforms, buyer communities, and complaint records where available. Coverage varies by product: single-source data carries lower published confidence than multi-source confirmations.
Crowdsourced Corrections
Vendors and buyers can report outdated or incorrect data. Vendor corrections flow through our verified claim portal. Buyer reports route into the same review queue — every report is logged and reviewed.
Timestamp & Source Everything
Every record carries its verification timestamp and a confidence score, and enriched fields name their sources. Historical snapshots record what our extractor observed on the vendor's pricing page on a given date.
Calculating True Cost
List price is just the beginning. True cost includes everything you'll actually pay to use the software effectively. Here's what we include:
License Fees
Base subscription costs at the appropriate tier for your use case
Implementation
Professional services, data migration, and initial setup costs
Training
User training, certification, and onboarding program costs
Support Tiers
Premium support, SLAs, and dedicated account management
Required Add-ons
Features marketed separately but needed for core functionality
Integrations
Third-party connectors, API access, and custom development
Spot something off?
Pricing changes fast. If you've found outdated or incorrect information, we want to know. If you're the vendor, claim your profile and flag it through the verified channel. If you're a buyer, drop us a note and we'll route it into the same review queue.