Best LLM API with a Free Tier in 2026
Most lists of “free” LLM APIs quietly mix different things: a standing $0 tier you can build on indefinitely, a limited evaluation path, and providers with no free access at all. Those are not interchangeable.
We checked 18 LLM API providers against their own live pricing and rate-limit documentation on 2026-08-06. Eight have a genuine standing free tier and are ranked below using only current tier names and pricing. The other ten are listed underneath with what they actually offer, so this page answers the question in both directions.
The best LLM API Providers tools in 2026 are Google Gemini API ($0–$18/per million tokens), Groq (free), and OpenRouter ($0–$75/per million tokens). Eight LLM API providers offer a genuine standing free tier as of August 2026: Google Gemini API (current Flash models free, no card required), Groq (30 requests/min, 131K context on gpt-oss-120b), OpenRouter (14 free models, up to 1M context, 50 requests/day), Cloudflare Workers AI (10,000 Neurons per day), Mistral AI (free mode, no card required), SambaNova Cloud (200,000 tokens/day per model), Cohere (1,000 calls/month, non-commercial use only) and Vercel AI Gateway (renewing monthly credit). Cerebras, Fireworks, Alibaba Qwen and Amazon Bedrock offer expiring credits rather than a free tier, and Together AI, DeepInfra, Moonshot Kimi, MiniMax, xAI and Perplexity charge from the first token.
Eight LLM API providers offer a genuine standing free tier as of August 2026: Google Gemini API (current Flash models free, no card required), Groq (30 requests/min, 131K context on gpt-oss-120b), OpenRouter (14 free models, up to 1M context, 50 requests/day), Cloudflare Workers AI (10,000 Neurons per day), Mistral AI (free mode, no card required), SambaNova Cloud (200,000 tokens/day per model), Cohere (1,000 calls/month, non-commercial use only) and Vercel AI Gateway (renewing monthly credit). Cerebras, Fireworks, Alibaba Qwen and Amazon Bedrock offer expiring credits rather than a free tier, and Together AI, DeepInfra, Moonshot Kimi, MiniMax, xAI and Perplexity charge from the first token.
Compare the top 3 side-by-side
Drag the seat slider, lock a tier per product, see Vendr median pricing and hidden costs for Google Gemini API, Groq, OpenRouter.
Our Rankings
Google Gemini API
Google's Gemini API puts a current $0 entry point behind its Free tier, with billing something you opt into when you outgrow it rather than something you set up first. Paid usage is split across Flash-Lite (Paid), Flash (Paid), and Pro (Paid), so teams can move from experimentation to production without changing providers.- Free rate limits
- Not published — the rate-limits doc states free-tier limits "can be viewed in Google AI Studio", which requires sign-in. Google Search grounding is separately documented at 1,500 requests/day free on Gemini 2.5 and 5,000 free search requests/month on Gemini 3.x.
- Free models
- Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 2.5 Flash, Gemini 2.5 Flash-Lite
- Card to sign up
- Not required
- Cheapest paid
- Gemini 2.5 Flash-Lite — $0.10 per 1M input, $0.40 per 1M output
- Verified
- 2026-08-06 · vendor source
- Current-generation Flash models are free of charge, not just small legacy models
- No credit card required — you set up billing only to move onto a paid tier
- Google Search grounding included free for 1,500 requests/day on Gemini 2.5
- Same API surface from the free tier through to paid, so nothing is rewritten when you scale
- Google no longer publishes free-tier rate limits — they are visible only in AI Studio once signed in
- Gemini 3.1 Pro Preview and Omni Flash Preview are paid-only
Groq
Groq publishes a clear Free tier for prototyping and evaluation, with paid production usage handled through Developer and high-volume deployments handled through Enterprise. That makes it one of the easier providers here to reason about before you write any code.- Free rate limits
- 30 requests/min on all main chat models. Daily: 14,400 requests on llama-3.1-8b-instant, 1,000 on llama-3.3-70b-versatile, gpt-oss-120b, gpt-oss-20b and qwen3.6-27b, 250 on groq/compound. Token budgets run 100K–500K/day depending on model.
- Free models
- llama-3.1-8b-instant, llama-3.3-70b-versatile, openai/gpt-oss-120b, openai/gpt-oss-20b, qwen/qwen3.6-27b, groq/compound and compound-mini, whisper-large-v3
- Context window
- 131,072 tokens on openai/gpt-oss-120b
- Card to sign up
- Not stated by vendor
- Cheapest paid
- Llama 3.1 8B — $0.05 per 1M input, $0.08 per 1M output
- Verified
- 2026-08-06 · vendor source
- Free tier available for prototyping and evaluation
- Developer tier is structured for production API usage
- Enterprise tier covers high-volume enterprise deployments
- Simple tier ladder from Free to Developer to Enterprise
- Whether a card is needed to sign up is not stated publicly — the billing page requires login
- Free daily token budgets are modest (100K–500K/day depending on model)
OpenRouter
OpenRouter offers Free Models at $0 for experimentation and development with free open-source models, then Pay-as-you-go pricing per million tokens for developers and teams that want model flexibility without multiple provider accounts.- Free rate limits
- 20 requests/min and 50 requests/day on :free models. The daily cap rises to 1,000 requests/day once $10 in credits has been purchased at any point (a lifetime threshold, not a required balance).
- Free models
- 14 ":free" model IDs, including nvidia/nemotron-3-ultra-550b-a55b:free, inclusionai/ling-3.0-flash:free, google/gemma-4-31b-it:free and poolside/laguna-s-2.1:free
- Context window
- 1,000,000 tokens on nvidia/nemotron-3-ultra-550b-a55b:free
- Card to sign up
- Not required
- Cheapest paid
- inclusionai/ling-2.6-flash — $0.01 per 1M input, $0.03 per 1M output
- Verified
- 2026-08-06 · vendor source
- Free Models tier available for experimentation and development with free open-source models
- Pay-as-you-go tier is priced per million tokens
- Useful for developers and teams wanting model flexibility without multiple provider accounts
- 50 requests/day is the tightest free daily cap here until you have spent $10
- It is a router — paid prices are set by the upstream provider, not by OpenRouter
Cloudflare Workers AI
Cloudflare Workers AI has a Free tier at $0 for prototyping and low-volume production at the edge, with Pay-as-you-go (Neurons / tokens) for latency-sensitive inference and Enterprise for high-volume production at the edge.- Free rate limits
- 10,000 Neurons per day at no charge, on both the Free and Paid Workers plans. The allocation resets at 00:00 UTC and requests fail once it is spent.
- Free models
- Drawn from the shared daily Neuron allocation rather than a fixed free-model list. @cf/zai-org/glm-5.2 is the one model flagged as requiring the Paid plan.
- Card to sign up
- Not stated by vendor
- Cheapest paid
- $0.011 per 1,000 Neurons beyond the free allocation
- Verified
- 2026-08-06 · vendor source
- Free tier is available for prototyping and low-volume production at the edge
- Inference runs inside the Workers runtime — no separate provider round-trip
- Pay-as-you-go (Neurons / tokens) is structured for latency-sensitive inference at the edge
- Enterprise tier covers high-volume production at the edge
- Neurons are a compute unit, not tokens — cost per request varies by model and is harder to forecast
Mistral AI API
Mistral enables API access through a $0 Free tier for evaluation and prototyping, making it a strong option for teams that need an EU-headquartered provider. Paid usage is organized by model family across Mistral Small, Mistral Medium, and Mistral Large.- Free rate limits
- Not published — the docs direct you to Admin Panel › API › Limits for your account’s limits.
- Card to sign up
- Not required
- Cheapest paid
- Ministral 3 3B — $0.10 per 1M input, $0.10 per 1M output
- Verified
- 2026-08-06 · vendor source
- Free tier available for evaluation and prototyping
- EU-based provider, which matters where data residency is a procurement requirement
- Mistral Small covers cost-efficient chat, classification, and code tasks
- Mistral Large is positioned for complex reasoning, agentic workflows, and vision tasks
- Free rate limits are not published in the tier data
- Which models are available in Free is not documented in the tier data
- La Plateforme is mid-rebrand to Mistral Studio and some docs URLs have moved
SambaNova Cloud
SambaNova Cloud's Free tier is aimed at testing ultra-low-latency inference, while Developer (Pay-as-you-go) is the production path for latency-critical workloads on open-weight models. Enterprise covers high-throughput production deployments where custom pricing is required.- Free rate limits
- 20 requests/min and 200,000 tokens/day, applied per model. The Free Tier applies for as long as no payment method is linked; linking one activates the paid Developer tier (60 requests/min, 12,000 requests/day, 20M tokens/day).
- Free models
- DeepSeek-V3.1, Meta-Llama-3.3-70B-Instruct, gpt-oss-120b, plus DeepSeek-V3.2 and gemma-4-31B-it in preview
- Card to sign up
- Not required
- Cheapest paid
- gpt-oss-120b — $0.22 per 1M input, $0.59 per 1M output
- Verified
- 2026-08-06 · vendor source
- Free tier available for testing ultra-low-latency inference
- Developer (Pay-as-you-go) supports latency-critical production workloads on open-weight models
- Enterprise tier covers high-throughput production deployments
- Clear path from Free tier to Developer (Pay-as-you-go)
- 20 requests/minute is a low ceiling for anything interactive
- Context windows are not published in the rate-limit documentation
Cohere API
Cohere's Trial (Free) tier gives teams a $0 path for evaluation and prototyping. Paid options are split across Pay-as-you-go (Legacy Models), Model Vault (Dedicated Instances), and Enterprise (Current Models), depending on whether you need legacy model usage, dedicated instances, or current-generation enterprise access.- Free rate limits
- 1,000 API calls/month on the trial key. Per-endpoint: Chat 20 requests/min, Rerank 10/min, Embed 2,000 inputs/min, Tokenize 100/min, Audio Transcriptions 5/min.
- Free models
- All Cohere models and APIs, including Command A, Command A+, Command R+ and Command R7B
- Context window
- 256,000 tokens on Command A
- Card to sign up
- Not stated by vendor
- Cheapest paid
- Aya Expanse 8B — $0.50 per 1M input, $1.50 per 1M output
- Verified
- 2026-08-06 · vendor source
- Trial (Free) tier available for evaluation and prototyping
- Pay-as-you-go (Legacy Models) supports existing customers on legacy Command model contracts
- Model Vault (Dedicated Instances) is built for dedicated, private, guaranteed-performance inference at scale
- Enterprise (Current Models) covers enterprises requiring current-generation models, private deployments, or AI platform access
- Trial (Free) is for evaluation and prototyping only
- Whether a card is needed at signup is not stated
Vercel AI SDK
Vercel AI SDK includes a Free Tier at $0 for individual developers and small projects experimenting with LLMs through a unified gateway. Paid (Pay-As-You-Go) covers teams and production workloads needing full access to all models at cost, while Enterprise is for stricter compliance, security, and scale requirements.- Free rate limits
- Rate limited per model; Vercel does not publish the numbers. Exceeding a limit returns a 429.
- Free models
- A subset of the gateway catalogue, not the full model list. Vercel does not publish which models.
- Card to sign up
- Not stated by vendor
- Cheapest paid
- Provider list price, with no markup and no platform fee on tokens
- Verified
- 2026-08-06 · vendor source
- Free Tier available for individual developers and small projects experimenting with LLMs through a unified gateway
- Paid (Pay-As-You-Go) supports teams and production workloads needing full access to all models at cost
- Enterprise covers strict compliance, security, and scale requirements
- Unified gateway across multiple models
- The Free Tier covers individual developers and small projects, not full production workloads
- Paid (Pay-As-You-Go) uses usage pricing rather than a fixed monthly plan
Also evaluated
These 10 providers were checked against the same criteria and did not qualify for the ranking. Each row states what the vendor actually offers.
Free credits, not a free tier (4)
A one-off or expiring grant. Useful for evaluation, but it runs out — these are not a standing way to keep building at $0.
| Provider | What you actually get | Verified |
|---|---|---|
| Cerebras Inference API | $5 in credits that expire 30 days after they are granted, and only once a verified payment method is on file. Cerebras states there is no permanently free tier. Trial limits are 5 requests/min and 30,000 tokens/min. | 2026-08-06 |
| Fireworks AI | $1 in free credits, spent against normal per-token rates. Accounts with no payment method are capped at 10 requests/min, and service is suspended once the credit is used up. | 2026-08-06 |
| Qwen API (Alibaba) | 1,000,000 tokens per model, valid 30–90 days from the day you activate Alibaba Cloud Model Studio. One-off and non-renewable, and limited to the international (Singapore) deployment. | 2026-08-06 |
| Amazon Bedrock | No Bedrock-specific free tier — Bedrock is not listed among AWS Free Tier services. New AWS accounts get up to $200 in account-level credits usable across services for up to 6 months. | 2026-08-06 |
No free tier (6)
Paid from the first token. Listed so this page answers the question both ways.
| Provider | What you actually get | Verified |
|---|---|---|
| Together AI | Together states it does not currently offer free trials; new accounts must buy a minimum $5 in credits before making a request. Cheapest model is $0.03 per 1M input tokens. | 2026-08-06 |
| DeepInfra | The pricing page states you must add a card or pre-pay before using the service, and no model is listed at $0. Cheapest model is $0.019 per 1M input tokens. | 2026-08-06 |
| Moonshot Kimi API | A minimum $1 top-up is required before the API will serve requests. The $5 voucher granted at $5 of cumulative spend is a rebate on money already spent, not a free tier. | 2026-08-06 |
| MiniMax API | No free tier on the text API. MiniMax publishes a free tier only for Video Generation V2, where it means 2 concurrent tasks rather than free tokens. | 2026-08-06 |
| xAI Grok API | Credits must be loaded before the API will serve requests. Tier 0 exists at $0 cumulative spend but governs rate limits only — it does not include free tokens. | 2026-08-06 |
| Perplexity API | No free allowance appears on any published pricing page. Sonar is billed per token plus a per-request search fee of $5–$12 per 1,000 requests. | 2026-08-06 |
Evaluation Criteria
- Free Tier Durability (5/5)
Whether the $0 tier is standing and indefinite, or a one-off credit grant that expires — the single distinction that decides whether a provider is ranked here at all
- Published Rate Limits (4/5)
Whether the vendor publishes free-tier requests/minute, requests/day and token budgets, so you can plan capacity before writing code
- No Credit Card Required (4/5)
Whether signup and API key issuance require a payment method on file
- Free Model Quality and Context (3/5)
Capability of the models reachable on the free tier and the context window they allow
- Paid Tier Value (2/5)
Price per million tokens when you outgrow free, so the upgrade path is not a cliff
How We Picked These
We evaluated 18 products and ranked the top 8 (last researched 2026-08-06).
Whether the $0 tier is standing and indefinite, or a one-off credit grant that expires — the single distinction that decides whether a provider is ranked here at all
Whether the vendor publishes free-tier requests/minute, requests/day and token budgets, so you can plan capacity before writing code
Whether signup and API key issuance require a payment method on file
Capability of the models reachable on the free tier and the context window they allow
Price per million tokens when you outgrow free, so the upgrade path is not a cliff
Frequently Asked Questions
01 Which LLM API has the best free tier in 2026?
Google Gemini API. It is the only free tier that includes current-generation frontier models (Gemini 3.6 Flash, 3.5 Flash, 2.5 Flash and 2.5 Flash-Lite are all listed free of charge) and requires no credit card — you set up billing only when you choose to move to a paid tier. Its one weakness is that Google no longer publishes free-tier rate limits publicly; they are visible only in AI Studio once you sign in. If you need published limits to plan against, Groq is the better choice at 30 requests/minute and up to 14,400 requests/day.
02 Which free LLM API has the highest rate limits?
It depends which limit binds first. Groq allows the most requests per day — 14,400 on llama-3.1-8b-instant. SambaNova Cloud allows the most tokens per day at 200,000 per model, and the quota is per model rather than shared. Groq also leads on requests per minute at 30, against 20 for both OpenRouter and SambaNova. OpenRouter is the most restrictive on daily requests at 50/day, though that rises to 1,000/day once you have bought $10 in credits at any point.
03 Which free LLM API tiers do not require a credit card?
Four providers state plainly that no card is needed: Google Gemini API, Mistral AI ("API access is enabled by default with no credit card required"), OpenRouter for its 50 requests/day tier, and SambaNova Cloud, where linking a card is what moves you off the free tier. Groq, Cohere, Cloudflare Workers AI and Vercel do not state either way on a public page. Cerebras does require a verified payment method, which is one reason its $5 grant is a trial rather than a free tier.
04 What is the largest context window available on a free LLM API tier?
1,000,000 tokens, on OpenRouter’s nvidia/nemotron-3-ultra-550b-a55b:free. Next is Cohere at 256,000 tokens on Command A, though the Cohere trial key is not licensed for commercial use, and then Groq at 131,072 tokens on openai/gpt-oss-120b. Several providers, including Mistral, SambaNova and Cloudflare, do not publish context windows alongside their free-tier documentation.
05 Is a free trial the same as a free tier?
No, and the difference is what makes most "free LLM API" lists misleading. A free tier is a standing $0 entry point that renews or persists indefinitely. A trial is a one-off grant that expires or exhausts: Cerebras gives $5 that expires in 30 days and requires a verified payment method first, Fireworks gives $1, Alibaba Model Studio gives 1,000,000 tokens per model valid 30–90 days, and Amazon Bedrock has no free tier of its own beyond account-level AWS credits. All four are useful for evaluation and none of them will keep a project running.
06 Which LLM APIs have no free tier at all?
Six of the providers we checked charge from the first token: Together AI (minimum $5 credit purchase, and it states it does not offer free trials), DeepInfra (card or pre-payment required), Moonshot Kimi (minimum $1 top-up), MiniMax (free tier exists only for video generation, not text), xAI (credits must be loaded first) and Perplexity (billed per token plus a per-request search fee).
07 Can I use a free LLM API tier in production?
Technically for most, but check the licence and the failure mode. Cohere’s trial key is explicitly not permitted for production or commercial use, which is a licence restriction rather than a rate limit. Elsewhere the practical constraint is behaviour at the ceiling: Cloudflare Workers AI fails requests once the daily 10,000-Neuron allocation is spent, and free daily caps on Groq and SambaNova are low enough that a modest traffic spike will hit them. Free tiers are best treated as prototyping and evaluation budgets.
08 What is the cheapest LLM API once you outgrow the free tier?
OpenRouter routes to models from $0.01 per 1M input tokens, and DeepInfra lists $0.019, though DeepInfra has no free tier to outgrow. Among providers that do have a free tier, Groq is cheapest at $0.05 per 1M input and $0.08 per 1M output on Llama 3.1 8B, followed by Google Gemini 2.5 Flash-Lite at $0.10/$0.40 and Mistral Ministral 3 3B at $0.10/$0.10.
Explore More LLM API Providers
See all LLM API Providers pricing and comparisons.
View all LLM API Providers software →