Lepton AI Pricing 2026
Complete pricing guide with plans, hidden costs, and cost analysis
Lepton AI uses usage-based pricing from $0.00021–$7.00 per million tokens; there is no monthly subscription price for this API tier.
Are you Lepton AI? Claim this profile
Is Lepton AI a fit?
Derived from this product's own pricing and contract record — not a review.
Best for
- Developers needing fast serverless inference for open-source models
- Teams deploying custom models with full control over GPU configuration
Not a fit if
- The sticker price is your whole budget — Tokenization Differences (and 4 more).
Lepton AI Usage Pricing
| Model | Unit | Rate |
|---|---|---|
| Llama 3.1 8B Instruct | 1M input tokens | $0.070 |
| Llama 3.1 8B Instruct | 1M output tokens | $0.070 |
| Llama 3.3 70B Instruct | 1M input tokens | $0.600 |
| Llama 3.3 70B Instruct | 1M output tokens | $0.600 |
| Mistral 7B Instruct | 1M input tokens | $0.070 |
| Mistral 7B Instruct | 1M output tokens | $0.070 |
| DeepSeek-R1 (671B) | 1M input tokens | $3.00 Full distilled model |
| DeepSeek-R1 (671B) | 1M output tokens | $7.00 Full distilled model |
| Llama 3.1 70B Instruct | 1M input tokens | $0.600 |
| Llama 3.1 70B Instruct | 1M output tokens | $0.600 |
- Rates approximate — verify at lepton.ai/pricing
- OpenAI-compatible endpoint: api.lepton.ai/api/v1
| Model | Unit | Rate |
|---|---|---|
| A10G GPU (24GB) | second | $0.00021 ~$0.75/hr |
| A100 80GB | second | $0.00056 ~$2.00/hr |
| H100 80GB | second | $0.001 ~$4.00/hr |
- GPU instance pricing approximate — verify at lepton.ai/pricing
- Price expressed per second; multiply by 3600 for hourly equivalent
Track Lepton AI pricing
Get an email when Lepton AI's pricing changes — plus the weekly SaaS Price Watch: verified price changes and deals across 3,000+ products. One-click unsubscribe.
You're on the list — first digest lands Tuesday.
Lepton AI uses usage-based pricing from $0.070 to $4.00 per million tokens as of August 2026. Usage cost depends on the selected model and generation volume.
- 9 documented hidden costs beyond list price
How much does Lepton AI cost?
Lepton AI Pricing Overview
Lepton AI uses usage-based pricing from $0.00021–$7.00 per million tokens across 13 published rates.
Lepton AI with a None minimum commitment.
There are at least 9 documented hidden costs beyond Lepton AI's list price, including implementation, training, and add-on fees.
This pricing was last verified in June 5, 2026 from 1 independent source.
Lepton AI is a cloud platform for AI workloads offering two main services: serverless LLM inference endpoints for popular open-source models, and GPU cloud instances for custom deployments. The inference API is OpenAI-compatible and supports Llama 3.x, Mistral, and other models with per-token pricing from $0.07/M tokens. GPU instances start at $0.39/hr. Lepton AI is designed for ML engineers who want a single platform for both quick API inference and custom model deployments.
How Lepton AI Pricing Compares
Compare Lepton AI pricing against top alternatives in LLM API Providers.
Compare Lepton AI vs Alternatives
Before committing to Lepton AI, compare pricing with these 3 alternatives in the same category.
What Companies Actually Pay for Lepton AI
How Lepton AI Pricing Compares
| Software | Starting Price | Top Price |
|---|---|---|
| Lepton AI | $0.070 | $4.00 per million tokens |
| Amazon Bedrock | $0.07 per million tokens | $75 per million tokens |
| Anyscale | Free | $4.9591 per million tokens |
| Baidu ERNIE API | $0.1 per million tokens | $10 per million tokens |
| Cerebras Inference API | $0.1 per million tokens | $6 per million tokens |
| Cohere API | Free | $15 per million tokens |
Detailed pricing comparisons:
Lepton AI Contract Terms
Lepton AI contracts do not auto-renew. Changes require advance notice. These terms are sourced from verified buyer experiences.
Lepton AI Pricing FAQ
01 How much does Lepton AI cost?
Lepton AI serverless inference starts at $0.07 per million tokens for small models like Llama 3.1 8B. GPU cloud instances start at approximately $0.75/hr for A10G. Pricing is pay-as-you-go with no minimum commitments.
02 What models does Lepton AI support?
Lepton AI supports Llama 3.x (8B to 70B), Mistral models, DeepSeek R1, and other popular open-source models through its serverless inference API. Custom model deployment is available on GPU cloud instances.
03 Is Lepton AI OpenAI-compatible?
Yes, Lepton AI's inference API is OpenAI-compatible. Point your OpenAI SDK to api.lepton.ai/api/v1 to use it as a drop-in replacement for compatible models.
04 Does Lepton AI have a free tier?
Lepton AI offers a free tier with limited credits for new accounts. Check lepton.ai for current free tier details and credit amounts.
Is this pricing incorrect? — we'll verify and update it.