AI Price Check — What Do AI Assistants Say Your Software Costs? | CostBench
AI Price Check — What Do AI Assistants Say Your Software Costs?
Shortlist
CostBench · AI Price Check

What do AI assistants say your product costs?

When an AI assistant answers a pricing question, it quotes a number. We sample assistant-grade AI models for that quote and score it against the verified CostBench pricing record — so you can see, per model, whether the price the AI gives is accurate, stale, or wrong.

The models we sample (named exactly, via API)
  • gemini-2.5-flash-grounded — the Gemini API model with Google Search grounding. Results do not represent the consumer Gemini or ChatGPT apps.
  • gemini-2.5-flash-memory — the same Gemini API model with no web search (parametric memory only).
  • perplexity-sonar and openai-gpt-4o-web — web-grounded assistant-grade API models, when funded.

Each result card shows the exact model id sampled. A verdict compares that model's quote to CostBench's current verified record — it is not a statement about any consumer chat app or the AI ecosystem as a whole.

Accuracy so far — by question type, never blended
“What does it cost per month?” 82% accurate · 131 of 160 answers
“What would it cost for a team of 25?” 33% accurate · 53 of 160 answers

Two different questions, two very different accuracy rates — so we never publish a single blended number. These are the gemini-2.5-flash-grounded lane over 160 products (as of July 13, 2026). Simple list-price questions land far more often than multi-seat team math, where most misses are flagged for review rather than confirmed wrong.

Longitudinal staleness tracking began 2026-07-13. Each product is re-sampled over time, so a time-series of how AI answers drift will appear here as real price changes occur — there is no back-dated history to show yet.

What the verdicts mean

  • Accurate every price the model quoted is within tolerance of the current verified record.
  • Partial some of the model’s prices match the current record and some do not.
  • Stale the quoted price matches an older value in the record’s history, not the current one.
  • Needs review the model's quote conflicts with the record — on currency, billing period, pricing axis, or by matching nothing. Flagged for review: either the model's answer or CostBench's record needs adjudication, and a human checks the vendor source before any call is made. We never auto-label a model "wrong".
  • Wrong set only by a human, after checking the vendor source confirms the model's price matches nothing current or historical — never emitted automatically.
  • No price the model declined to state a number.
  • Custom OK the product is custom / contact-sales priced, and the model said so.
It measures
  • What a specific assistant-grade AI model, sampled via API, quoted for this product’s pricing.
  • Whether that quote matches CostBench’s current verified record, which we uniquely maintain.
  • The exact model and retrieval mode sampled — labeled per model, never blended.
It does not measure
  • The consumer ChatGPT, Claude, or Gemini apps — we sample assistant-grade API models only.
  • A product’s standing, citations, or ranking across the whole AI ecosystem.
  • Anything CostBench can control in an assistant’s answer — we can’t, and don’t claim to.