Price check $0.00021–$7.00 per million tokens 13 published rates · pay as you go

Lepton AI Usage Pricing

Model Unit Rate
Llama 3.1 8B Instruct 1M input tokens $0.070
Llama 3.1 8B Instruct 1M output tokens $0.070
Llama 3.3 70B Instruct 1M input tokens $0.600
Llama 3.3 70B Instruct 1M output tokens $0.600
Mistral 7B Instruct 1M input tokens $0.070
Mistral 7B Instruct 1M output tokens $0.070
DeepSeek-R1 (671B) 1M input tokens $3.00 Full distilled model
DeepSeek-R1 (671B) 1M output tokens $7.00 Full distilled model
Llama 3.1 70B Instruct 1M input tokens $0.600
Llama 3.1 70B Instruct 1M output tokens $0.600
  • Rates approximate — verify at lepton.ai/pricing
  • OpenAI-compatible endpoint: api.lepton.ai/api/v1
Model Unit Rate
A10G GPU (24GB) second $0.00021 ~$0.75/hr
A100 80GB second $0.00056 ~$2.00/hr
H100 80GB second $0.001 ~$4.00/hr
  • GPU instance pricing approximate — verify at lepton.ai/pricing
  • Price expressed per second; multiply by 3600 for hourly equivalent
Pricing Alerts

Track Lepton AI pricing

Get an email when Lepton AI's pricing changes — plus the weekly SaaS Price Watch: verified price changes and deals across 3,000+ products. One-click unsubscribe.

Compare Lepton AI with alternativesAdjust seats, lock a tier, add up to 2 more products side-by-side. Shareable URL.
Quick Answer
Last verified:
Medium confidence

Lepton AI uses usage-based pricing from $0.070 to $4.00 per million tokens as of August 2026. Usage cost depends on the selected model and generation volume.

  • 9 documented hidden costs beyond list price

How much does Lepton AI cost?

Lepton AI uses usage-based pricing from $0.00021–$7.00 per million tokens. The published rate card contains 13 per-generation prices; there is no monthly subscription price for this API tier.

Lepton AI Pricing Overview

Lepton AI uses usage-based pricing from $0.00021–$7.00 per million tokens across 13 published rates.

Lepton AI with a None minimum commitment.

There are at least 9 documented hidden costs beyond Lepton AI's list price, including implementation, training, and add-on fees.

This pricing was last verified in June 5, 2026 from 1 independent source.

Lepton AI is a cloud platform for AI workloads offering two main services: serverless LLM inference endpoints for popular open-source models, and GPU cloud instances for custom deployments. The inference API is OpenAI-compatible and supports Llama 3.x, Mistral, and other models with per-token pricing from $0.07/M tokens. GPU instances start at $0.39/hr. Lepton AI is designed for ML engineers who want a single platform for both quick API inference and custom model deployments.

How Lepton AI Pricing Compares

Compare Lepton AI pricing against top alternatives in LLM API Providers.

Compare Lepton AI vs Alternatives

Before committing to Lepton AI, compare pricing with these 3 alternatives in the same category.

All Lepton AI alternatives & migration guides

What Companies Actually Pay for Lepton AI

Review scores
Third-party review aggregates, as of Jul 2026
Top pricing complaints
Hardware ReliabilityNetwork ComplexitySoftware ConfigurationMaintenance Complexity and Scalability Issues

How Lepton AI Pricing Compares

Software Starting Price Top Price
Lepton AI $0.070 $4.00 per million tokens
Amazon Bedrock $0.07 per million tokens $75 per million tokens
Anyscale Free $4.9591 per million tokens
Baidu ERNIE API $0.1 per million tokens $10 per million tokens
Cerebras Inference API $0.1 per million tokens $6 per million tokens
Cohere API Free $15 per million tokens

9 Lepton AI Hidden Costs Beyond the List Price

Beyond the listed price, Lepton AI has at least 9 documented hidden costs that can significantly increase total cost of ownership.

Watch for 9 hidden costs
  • Tokenization Differences 5.3 times more
    high 1 source
    industry "For example, one analysis found that the same input could produce 2.65 times more tokens depending on the model, and on tool-heavy workloads, one model cost 5.3 times more than another despite only a 2x difference in list prices."
  • Context Window Accumulation $1,750
    high 1 source
    industry "This means that even a short response can incur costs for the entire accumulated conversation history."
  • Input vs. Output Asymmetry 4 to 10 times more
    medium 1 source
    industry "Output Asymmetry: Most LLM API providers charge separately for input and output tokens, with output tokens typically costing 4 to 10 times more due to higher compute requirements."
  • Tiered Pricing Thresholds double the per-token cost
    high 1 source
    industry "Tiered Pricing and Context Thresholds: Some providers have tiered pricing based on context length, where crossing a certain threshold (e.g., 128K tokens) can double the per-token cost."
  • Non-Token Billing
    medium 1 source
    industry "Non-Token Billing: Costs can also arise from non-token-based billing for services like image generation (billed by image quality and resolution), video generation (per second), and fine-tuning (per-token or per-hour)."
  • Provider-Specific Model Pricing
    low 1 source
    industry "Provider-Specific Pricing for the Same Model: The same LLM model can have different prices depending on the provider offering it."
  • Marketplace Brokerage Fees
    medium 1 source
    industry "NVIDIA adds a brokerage margin on top of these underlying neocloud rates, which can make direct price comparisons challenging."
  • Storage and Data Transfer (Egress) Fees $0 to over $0.50/GB
    high 1 source
    industry "For general cloud providers, egress fees can range from $0 to over $0.50/GB, with some hyperscalers charging around $0.09 per GB for data exceeding 100GB/month."
  • Engineering Complexity and Maintenance
    high 1 source
    industry "Lepton AI itself dramatically reduced its cloud storage costs by 96.7% to 98% by optimizing its storage solution, switching from Amazon EFS to JuiceFS, which leverages object storage and providers with no data transfer fees."
Tip

Ask your Lepton AI sales rep about these costs upfront. Getting them in writing before signing can save you from surprise charges later.

Full hidden costs breakdown →

Intelligence sourced from 1 independent sources
industry
Key claims include inline source attribution. Data verified against multiple independent sources. 10 source citations total.

Lepton AI Contract Terms

Lepton AI contracts do not auto-renew. Changes require advance notice. These terms are sourced from verified buyer experiences.

Contract Terms
Minimum Commitment None
Based on 1 verified source

Lepton AI Pricing FAQ

01 How much does Lepton AI cost?

Lepton AI serverless inference starts at $0.07 per million tokens for small models like Llama 3.1 8B. GPU cloud instances start at approximately $0.75/hr for A10G. Pricing is pay-as-you-go with no minimum commitments.

02 What models does Lepton AI support?

Lepton AI supports Llama 3.x (8B to 70B), Mistral models, DeepSeek R1, and other popular open-source models through its serverless inference API. Custom model deployment is available on GPU cloud instances.

03 Is Lepton AI OpenAI-compatible?

Yes, Lepton AI's inference API is OpenAI-compatible. Point your OpenAI SDK to api.lepton.ai/api/v1 to use it as a drop-in replacement for compatible models.

04 Does Lepton AI have a free tier?

Lepton AI offers a free tier with limited credits for new accounts. Check lepton.ai for current free tier details and credit amounts.

Is this pricing incorrect? — we'll verify and update it.