What do AI assistants say your product costs?
When an AI assistant answers a pricing question, it quotes a number. We sample assistant-grade AI models for that quote and score it against the verified CostBench pricing record — so you can see, per model, whether the price the AI gives is accurate, stale, or wrong.
gemini-2.5-flash-grounded— the Gemini API model with Google Search grounding. Results do not represent the consumer Gemini or ChatGPT apps.gemini-2.5-flash-memory— the same Gemini API model with no web search (parametric memory only).perplexity-sonarandopenai-gpt-4o-web— web-grounded assistant-grade API models, when funded.
Each result card shows the exact model id sampled. A verdict compares that model's quote to CostBench's current verified record — it is not a statement about any consumer chat app or the AI ecosystem as a whole.
Two different questions, two very different accuracy rates — so we never publish a single
blended number. These are the gemini-2.5-flash-grounded lane over 160 products
(as of July 13, 2026). Simple list-price questions land far more often than multi-seat team math,
where most misses are flagged for review rather than confirmed wrong.
Longitudinal staleness tracking began 2026-07-13. Each product is re-sampled over time, so a time-series of how AI answers drift will appear here as real price changes occur — there is no back-dated history to show yet.
AI price-accuracy checks are coming soon
We’re sampling assistant-grade AI models and scoring their pricing answers against our verified record. This surface opens with the accuracy beta. In the meantime, see how often AI assistants retrieve a product’s CostBench pricing record:
Open the AI Pricing Visibility tool →Autocomplete searches every product with a CostBench pricing record.
—
Not yet sampled. We haven’t sampled AI pricing answers for this product yet — no data to show. Sampling expands over time; claim the record free and we’ll include it.
Claim this record free →We sample assistant-grade AI models via API — Gemini with search grounding, Perplexity Sonar, and GPT with web — never the consumer chat apps. Each model is named exactly as sampled. A verdict compares that model’s quote to CostBench’s current record. How we measure →
No pricing record found for that product on CostBench yet.
What the verdicts mean
- Accurate every price the model quoted is within tolerance of the current verified record.
- Partial some of the model’s prices match the current record and some do not.
- Stale the quoted price matches an older value in the record’s history, not the current one.
- Needs review the model's quote conflicts with the record — on currency, billing period, pricing axis, or by matching nothing. Flagged for review: either the model's answer or CostBench's record needs adjudication, and a human checks the vendor source before any call is made. We never auto-label a model "wrong".
- Wrong set only by a human, after checking the vendor source confirms the model's price matches nothing current or historical — never emitted automatically.
- No price the model declined to state a number.
- Custom OK the product is custom / contact-sales priced, and the model said so.
- What a specific assistant-grade AI model, sampled via API, quoted for this product’s pricing.
- Whether that quote matches CostBench’s current verified record, which we uniquely maintain.
- The exact model and retrieval mode sampled — labeled per model, never blended.
- The consumer ChatGPT, Claude, or Gemini apps — we sample assistant-grade API models only.
- A product’s standing, citations, or ranking across the whole AI ecosystem.
- Anything CostBench can control in an assistant’s answer — we can’t, and don’t claim to.