Best LLM API with a Free Tier in 2026
Most lists of “free” LLM APIs quietly mix three different things: a standing $0 tier you can build on indefinitely, a one-off credit grant that expires, and a provider with no free access at all. Those are not interchangeable. A $5 credit that expires in 30 days will not carry a side project past its first week.
We checked 18 LLM API providers against their own live pricing and rate-limit documentation on 2026-08-06. Eight have a genuine standing free tier and are ranked below with their verified rate limits, free models, context windows and whether a card is needed to sign up. The other ten are listed underneath with what they actually offer, so this page answers the question in both directions. Where a vendor does not publish a number, we say so rather than estimate it.
The best LLM API Providers tools in 2026 are Google Gemini API ($0–$18/per million tokens), Groq (free), and OpenRouter ($0–$75/per million tokens). Eight LLM API providers offer a genuine standing free tier as of August 2026: Google Gemini API (current Flash models free, no card required), Groq (30 requests/min, 131K context on gpt-oss-120b), OpenRouter (14 free models, up to 1M context, 50 requests/day), Cloudflare Workers AI (10,000 Neurons per day), Mistral AI (free mode, no card required), SambaNova Cloud (200,000 tokens/day per model), Cohere (1,000 calls/month, non-commercial use only) and Vercel AI Gateway (renewing monthly credit). Cerebras, Fireworks, Alibaba Qwen and Amazon Bedrock offer expiring credits rather than a free tier, and Together AI, DeepInfra, Moonshot Kimi, MiniMax, xAI and Perplexity charge from the first token.
Eight LLM API providers offer a genuine standing free tier as of August 2026: Google Gemini API (current Flash models free, no card required), Groq (30 requests/min, 131K context on gpt-oss-120b), OpenRouter (14 free models, up to 1M context, 50 requests/day), Cloudflare Workers AI (10,000 Neurons per day), Mistral AI (free mode, no card required), SambaNova Cloud (200,000 tokens/day per model), Cohere (1,000 calls/month, non-commercial use only) and Vercel AI Gateway (renewing monthly credit). Cerebras, Fireworks, Alibaba Qwen and Amazon Bedrock offer expiring credits rather than a free tier, and Together AI, DeepInfra, Moonshot Kimi, MiniMax, xAI and Perplexity charge from the first token.
Compare the top 3 side-by-side
Drag the seat slider, lock a tier per product, see Vendr median pricing and hidden costs for Google Gemini API, Groq, OpenRouter.
Our Rankings
Google Gemini API
Google's Gemini API is the only free tier on this list that puts current frontier models behind a $0 entry point with no card on file: Gemini 3.6 Flash, 3.5 Flash, 2.5 Flash and 2.5 Flash-Lite are all listed free of charge, and billing is something you opt into when you outgrow it rather than something you set up first. Paid usage starts at $0.10 per 1M input tokens on 2.5 Flash-Lite.- Free rate limits
- Not published — the rate-limits doc states free-tier limits "can be viewed in Google AI Studio", which requires sign-in. Google Search grounding is separately documented at 1,500 requests/day free on Gemini 2.5 and 5,000 free search requests/month on Gemini 3.x.
- Free models
- Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 2.5 Flash, Gemini 2.5 Flash-Lite
- Card to sign up
- Not required
- Cheapest paid
- Gemini 2.5 Flash-Lite — $0.10 per 1M input, $0.40 per 1M output
- Verified
- 2026-08-06 · vendor source
- Current-generation Flash models are free of charge, not just small legacy models
- No credit card required — you set up billing only to move onto a paid tier
- Google Search grounding included free for 1,500 requests/day on Gemini 2.5
- Same API surface from the free tier through to paid, so nothing is rewritten when you scale
- Google no longer publishes free-tier rate limits — they are visible only in AI Studio once signed in
- Gemini 3.1 Pro Preview and Omni Flash Preview are paid-only
Groq
Groq publishes an exact free-plan table — 30 requests/minute across its main chat models, 1,000 requests/day, and 14,400/day on llama-3.1-8b-instant — which makes it the easiest provider here to plan against before you write any code. The free tier runs gpt-oss-120b at a 131,072-token context, and paid usage starts at $0.05 per 1M input tokens.- Free rate limits
- 30 requests/min on all main chat models. Daily: 14,400 requests on llama-3.1-8b-instant, 1,000 on llama-3.3-70b-versatile, gpt-oss-120b, gpt-oss-20b and qwen3.6-27b, 250 on groq/compound. Token budgets run 100K–500K/day depending on model.
- Free models
- llama-3.1-8b-instant, llama-3.3-70b-versatile, openai/gpt-oss-120b, openai/gpt-oss-20b, qwen/qwen3.6-27b, groq/compound and compound-mini, whisper-large-v3
- Context window
- 131,072 tokens on openai/gpt-oss-120b
- Card to sign up
- Not stated by vendor
- Cheapest paid
- Llama 3.1 8B — $0.05 per 1M input, $0.08 per 1M output
- Verified
- 2026-08-06 · vendor source
- Per-model free limits published in full: RPM, RPD, TPM and TPD
- gpt-oss-120b available free at a 131,072-token context window
- llama-3.1-8b-instant allows 14,400 requests/day — the highest free request ceiling here
- Cheapest paid entry of any provider on this list at $0.05 per 1M input tokens
- Whether a card is needed to sign up is not stated publicly — the billing page requires login
- Free daily token budgets are modest (100K–500K/day depending on model)
OpenRouter
OpenRouter exposes 14 model IDs suffixed ":free" that are permanently free behind a single API key, including a 1,000,000-token-context Nemotron model. The free allowance is 20 requests/minute and 50 requests/day, rising to 1,000/day once you have bought $10 in credits at any point — a lifetime threshold, not a running balance.- Free rate limits
- 20 requests/min and 50 requests/day on :free models. The daily cap rises to 1,000 requests/day once $10 in credits has been purchased at any point (a lifetime threshold, not a required balance).
- Free models
- 14 ":free" model IDs, including nvidia/nemotron-3-ultra-550b-a55b:free, inclusionai/ling-3.0-flash:free, google/gemma-4-31b-it:free and poolside/laguna-s-2.1:free
- Context window
- 1,000,000 tokens on nvidia/nemotron-3-ultra-550b-a55b:free
- Card to sign up
- Not required
- Cheapest paid
- inclusionai/ling-2.6-flash — $0.01 per 1M input, $0.03 per 1M output
- Verified
- 2026-08-06 · vendor source
- 14 permanently free model IDs behind one key, so you can A/B models without new accounts
- The largest free context window on this list — 1,000,000 tokens on nemotron-3-ultra-550b
- No credit card needed for the 50 requests/day tier
- Paid routing starts at $0.01 per 1M input tokens
- 50 requests/day is the tightest free daily cap here until you have spent $10
- It is a router — paid prices are set by the upstream provider, not by OpenRouter
Cloudflare Workers AI
Workers AI gives every account 10,000 Neurons per day at no charge on both the Free and Paid Workers plans, resetting at 00:00 UTC. It is the only free tier here that runs inference in the same runtime as your application code, which removes a network hop for anything already deployed on Cloudflare. Beyond the daily allocation you pay $0.011 per 1,000 Neurons.- Free rate limits
- 10,000 Neurons per day at no charge, on both the Free and Paid Workers plans. The allocation resets at 00:00 UTC and requests fail once it is spent.
- Free models
- Drawn from the shared daily Neuron allocation rather than a fixed free-model list. @cf/zai-org/glm-5.2 is the one model flagged as requiring the Paid plan.
- Card to sign up
- Not stated by vendor
- Cheapest paid
- $0.011 per 1,000 Neurons beyond the free allocation
- Verified
- 2026-08-06 · vendor source
- A daily free allocation that renews indefinitely, on the free Workers plan
- Inference runs inside the Workers runtime — no separate provider round-trip
- One shared allocation across the model catalogue rather than a per-model quota
- Cheapest per-model rate on the catalogue is $0.027 per 1M input tokens
- Neurons are a compute unit, not tokens — cost per request varies by model and is harder to forecast
- Requests fail once the daily allocation is spent rather than degrading
- Cloudflare does not publish which models are free-plan eligible; glm-5.2 is flagged Paid-plan only
Mistral AI API
Mistral enables API access by default in what it calls Free mode, stating plainly that no credit card is required — the shortest path from signup to a working key on this list, and the strongest option for teams that need an EU-headquartered provider. Paid usage starts at $0.10 per 1M tokens on Ministral 3 3B.- Free rate limits
- Not published — the docs direct you to Admin Panel › API › Limits for your account’s limits.
- Card to sign up
- Not required
- Cheapest paid
- Ministral 3 3B — $0.10 per 1M input, $0.10 per 1M output
- Verified
- 2026-08-06 · vendor source
- API access enabled by default with no credit card, in the vendor’s own words
- EU-based provider, which matters where data residency is a procurement requirement
- Symmetric input/output pricing on the Ministral 3 family makes cost easy to forecast
- Paid entry at $0.10 per 1M input and output tokens
- Free-mode rate limits are not published — the docs direct you to your account Admin Panel
- Which models are available in Free mode is not documented publicly
- La Plateforme is mid-rebrand to Mistral Studio and some docs URLs have moved
SambaNova Cloud
SambaNova's Free Tier applies for as long as no payment method is linked to the account, and publishes 200,000 tokens per day per model — the largest documented free daily token budget here — across DeepSeek-V3.1, Llama-3.3-70B and gpt-oss-120b. Adding a card is what moves you onto the paid Developer tier, so the free tier cannot lapse by surprise.- Free rate limits
- 20 requests/min and 200,000 tokens/day, applied per model. The Free Tier applies for as long as no payment method is linked; linking one activates the paid Developer tier (60 requests/min, 12,000 requests/day, 20M tokens/day).
- Free models
- DeepSeek-V3.1, Meta-Llama-3.3-70B-Instruct, gpt-oss-120b, plus DeepSeek-V3.2 and gemma-4-31B-it in preview
- Card to sign up
- Not required
- Cheapest paid
- gpt-oss-120b — $0.22 per 1M input, $0.59 per 1M output
- Verified
- 2026-08-06 · vendor source
- 200,000 tokens/day per model, and the quota is per model rather than shared
- The free tier ends only when you choose to link a payment method
- Frontier open-weight models included free: DeepSeek-V3.1, Llama-3.3-70B, gpt-oss-120b
- Paid entry at $0.22 per 1M input tokens on gpt-oss-120b
- 20 requests/minute is a low ceiling for anything interactive
- Context windows are not published in the rate-limit documentation
Cohere API
Cohere issues every account a free rate-limited trial key with no stated expiry, covering the whole model range including Command A at a 256,000-token context, plus the Embed and Rerank endpoints that most RAG stacks actually need. The catch is a licence one, not a technical one: the trial key is explicitly not permitted for production or commercial use.- Free rate limits
- 1,000 API calls/month on the trial key. Per-endpoint: Chat 20 requests/min, Rerank 10/min, Embed 2,000 inputs/min, Tokenize 100/min, Audio Transcriptions 5/min.
- Free models
- All Cohere models and APIs, including Command A, Command A+, Command R+ and Command R7B
- Context window
- 256,000 tokens on Command A
- Card to sign up
- Not stated by vendor
- Cheapest paid
- Aya Expanse 8B — $0.50 per 1M input, $1.50 per 1M output
- Verified
- 2026-08-06 · vendor source
- Covers Embed and Rerank, not just chat — the pieces a RAG pipeline needs
- 256,000-token context on Command A, the largest free context here outside OpenRouter
- No stated expiry date on the trial key
- Embed allows 2,000 inputs/minute, which is generous for corpus indexing
- Explicitly not licensed for production or commercial use — evaluation and prototyping only
- 1,000 API calls/month is a hard monthly ceiling
- Whether a card is needed at signup is not stated
Vercel AI SDK
Every Vercel team account gets a free tier with a monthly credit that renews until you buy credits of your own, covering a subset of the model catalogue through one endpoint. Its distinguishing property is the billing model rather than the allowance: Vercel charges no markup and no platform fee on tokens, so you pay each provider's list price.- Free rate limits
- Rate limited per model; Vercel does not publish the numbers. Exceeding a limit returns a 429.
- Free models
- A subset of the gateway catalogue, not the full model list. Vercel does not publish which models.
- Card to sign up
- Not stated by vendor
- Cheapest paid
- Provider list price, with no markup and no platform fee on tokens
- Verified
- 2026-08-06 · vendor source
- A monthly credit that renews, rather than a one-off grant
- No markup and no platform fee on tokens — you pay provider list price
- One endpoint across many providers, so swapping models is a config change
- Bring-your-own-key is also unmarked up on the paid tier
- Vercel does not publish the dollar value of the monthly free credit
- The free tier covers a subset of models, not the full catalogue
- Free requests are rate limited per model and the limits are not published
- Buying any credits ends the monthly free credit
Also evaluated
These 10 providers were checked against the same criteria and did not qualify for the ranking. Each row states what the vendor actually offers.
Free credits, not a free tier (4)
A one-off or expiring grant. Useful for evaluation, but it runs out — these are not a standing way to keep building at $0.
| Provider | What you actually get | Verified |
|---|---|---|
| Cerebras Inference API | $5 in credits that expire 30 days after they are granted, and only once a verified payment method is on file. Cerebras states there is no permanently free tier. Trial limits are 5 requests/min and 30,000 tokens/min. | 2026-08-06 |
| Fireworks AI | $1 in free credits, spent against normal per-token rates. Accounts with no payment method are capped at 10 requests/min, and service is suspended once the credit is used up. | 2026-08-06 |
| Qwen API (Alibaba) | 1,000,000 tokens per model, valid 30–90 days from the day you activate Alibaba Cloud Model Studio. One-off and non-renewable, and limited to the international (Singapore) deployment. | 2026-08-06 |
| Amazon Bedrock | No Bedrock-specific free tier — Bedrock is not listed among AWS Free Tier services. New AWS accounts get up to $200 in account-level credits usable across services for up to 6 months. | 2026-08-06 |
No free tier (6)
Paid from the first token. Listed so this page answers the question both ways.
| Provider | What you actually get | Verified |
|---|---|---|
| Together AI | Together states it does not currently offer free trials; new accounts must buy a minimum $5 in credits before making a request. Cheapest model is $0.03 per 1M input tokens. | 2026-08-06 |
| DeepInfra | The pricing page states you must add a card or pre-pay before using the service, and no model is listed at $0. Cheapest model is $0.019 per 1M input tokens. | 2026-08-06 |
| Moonshot Kimi API | A minimum $1 top-up is required before the API will serve requests. The $5 voucher granted at $5 of cumulative spend is a rebate on money already spent, not a free tier. | 2026-08-06 |
| MiniMax API | No free tier on the text API. MiniMax publishes a free tier only for Video Generation V2, where it means 2 concurrent tasks rather than free tokens. | 2026-08-06 |
| xAI Grok API | Credits must be loaded before the API will serve requests. Tier 0 exists at $0 cumulative spend but governs rate limits only — it does not include free tokens. | 2026-08-06 |
| Perplexity API | No free allowance appears on any published pricing page. Sonar is billed per token plus a per-request search fee of $5–$12 per 1,000 requests. | 2026-08-06 |
Evaluation Criteria
- Free Tier Durability (5/5)
Whether the $0 tier is standing and indefinite, or a one-off credit grant that expires — the single distinction that decides whether a provider is ranked here at all
- Published Rate Limits (4/5)
Whether the vendor publishes free-tier requests/minute, requests/day and token budgets, so you can plan capacity before writing code
- No Credit Card Required (4/5)
Whether signup and API key issuance require a payment method on file
- Free Model Quality and Context (3/5)
Capability of the models reachable on the free tier and the context window they allow
- Paid Tier Value (2/5)
Price per million tokens when you outgrow free, so the upgrade path is not a cliff
How We Picked These
We evaluated 18 products and ranked the top 8 (last researched 2026-08-06).
Whether the $0 tier is standing and indefinite, or a one-off credit grant that expires — the single distinction that decides whether a provider is ranked here at all
Whether the vendor publishes free-tier requests/minute, requests/day and token budgets, so you can plan capacity before writing code
Whether signup and API key issuance require a payment method on file
Capability of the models reachable on the free tier and the context window they allow
Price per million tokens when you outgrow free, so the upgrade path is not a cliff
Frequently Asked Questions
01 Which LLM API has the best free tier in 2026?
Google Gemini API. It is the only free tier that includes current-generation frontier models (Gemini 3.6 Flash, 3.5 Flash, 2.5 Flash and 2.5 Flash-Lite are all listed free of charge) and requires no credit card — you set up billing only when you choose to move to a paid tier. Its one weakness is that Google no longer publishes free-tier rate limits publicly; they are visible only in AI Studio once you sign in. If you need published limits to plan against, Groq is the better choice at 30 requests/minute and up to 14,400 requests/day.
02 Which free LLM API has the highest rate limits?
It depends which limit binds first. Groq allows the most requests per day — 14,400 on llama-3.1-8b-instant. SambaNova Cloud allows the most tokens per day at 200,000 per model, and the quota is per model rather than shared. Groq also leads on requests per minute at 30, against 20 for both OpenRouter and SambaNova. OpenRouter is the most restrictive on daily requests at 50/day, though that rises to 1,000/day once you have bought $10 in credits at any point.
03 Which free LLM API tiers do not require a credit card?
Four providers state plainly that no card is needed: Google Gemini API, Mistral AI ("API access is enabled by default with no credit card required"), OpenRouter for its 50 requests/day tier, and SambaNova Cloud, where linking a card is what moves you off the free tier. Groq, Cohere, Cloudflare Workers AI and Vercel do not state either way on a public page. Cerebras does require a verified payment method, which is one reason its $5 grant is a trial rather than a free tier.
04 What is the largest context window available on a free LLM API tier?
1,000,000 tokens, on OpenRouter’s nvidia/nemotron-3-ultra-550b-a55b:free. Next is Cohere at 256,000 tokens on Command A, though the Cohere trial key is not licensed for commercial use, and then Groq at 131,072 tokens on openai/gpt-oss-120b. Several providers, including Mistral, SambaNova and Cloudflare, do not publish context windows alongside their free-tier documentation.
05 Is a free trial the same as a free tier?
No, and the difference is what makes most "free LLM API" lists misleading. A free tier is a standing $0 entry point that renews or persists indefinitely. A trial is a one-off grant that expires or exhausts: Cerebras gives $5 that expires in 30 days and requires a verified payment method first, Fireworks gives $1, Alibaba Model Studio gives 1,000,000 tokens per model valid 30–90 days, and Amazon Bedrock has no free tier of its own beyond account-level AWS credits. All four are useful for evaluation and none of them will keep a project running.
06 Which LLM APIs have no free tier at all?
Six of the providers we checked charge from the first token: Together AI (minimum $5 credit purchase, and it states it does not offer free trials), DeepInfra (card or pre-payment required), Moonshot Kimi (minimum $1 top-up), MiniMax (free tier exists only for video generation, not text), xAI (credits must be loaded first) and Perplexity (billed per token plus a per-request search fee).
07 Can I use a free LLM API tier in production?
Technically for most, but check the licence and the failure mode. Cohere’s trial key is explicitly not permitted for production or commercial use, which is a licence restriction rather than a rate limit. Elsewhere the practical constraint is behaviour at the ceiling: Cloudflare Workers AI fails requests once the daily 10,000-Neuron allocation is spent, and free daily caps on Groq and SambaNova are low enough that a modest traffic spike will hit them. Free tiers are best treated as prototyping and evaluation budgets.
08 What is the cheapest LLM API once you outgrow the free tier?
OpenRouter routes to models from $0.01 per 1M input tokens, and DeepInfra lists $0.019, though DeepInfra has no free tier to outgrow. Among providers that do have a free tier, Groq is cheapest at $0.05 per 1M input and $0.08 per 1M output on Llama 3.1 8B, followed by Google Gemini 2.5 Flash-Lite at $0.10/$0.40 and Mistral Ministral 3 3B at $0.10/$0.10.
Explore More LLM API Providers
See all LLM API Providers pricing and comparisons.
View all LLM API Providers software →