Together AI Pricing 2026
Complete pricing guide with plans, hidden costs, and cost analysis
Together AI uses custom pricing — contact their sales team for a quote.
Are you Together AI? Claim this profile
Is Together AI a fit?
Derived from this product's own pricing and contract record — not a review.
Best for
- Variable-volume API inference usage
- Consistent high-volume inference on dedicated H100 GPUs
- High-throughput dedicated inference on H200 GPUs
- High-performance dedicated inference on B200 GPUs
Not a fit if
- You need to budget without talking to sales — every plan here is quote-only.
- You want to try it before paying — there is no free plan and no trial.
- The sticker price is your whole budget — Fine-tune Hosting (and 1 more).
All Together AI Plans & Pricing
| Plan | Monthly | Annual | Best For |
|---|---|---|---|
| Serverless Inference | Contact Sales | Contact Sales | Variable-volume API inference usage |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Serverless Inference Best for: Variable-volume API inference usage
| |||
| Dedicated Inference (NVIDIA HGX H100) | Contact Sales | Contact Sales | Consistent high-volume inference on dedicated H100 GPUs |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Dedicated Inference (NVIDIA HGX H100) Best for: Consistent high-volume inference on dedicated H100 GPUs
| |||
| Dedicated Inference (NVIDIA HGX H200) | Contact Sales | Contact Sales | High-throughput dedicated inference on H200 GPUs |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Dedicated Inference (NVIDIA HGX H200) Best for: High-throughput dedicated inference on H200 GPUs
| |||
| Dedicated Inference (NVIDIA HGX B200) | Contact Sales | Contact Sales | High-performance dedicated inference on B200 GPUs |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Dedicated Inference (NVIDIA HGX B200) Best for: High-performance dedicated inference on B200 GPUs
| |||
| Dedicated Inference (NVIDIA HGX B300) | Contact Sales | Contact Sales | Latest-generation dedicated inference deployments |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Dedicated Inference (NVIDIA HGX B300) Best for: Latest-generation dedicated inference deployments
| |||
| Dedicated Inference (NVIDIA GB200 NVL72) | Contact Sales | Contact Sales | Large-scale dedicated inference deployments |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Dedicated Inference (NVIDIA GB200 NVL72) Best for: Large-scale dedicated inference deployments
| |||
| Dedicated Inference (NVIDIA GB300 NVL72) | Contact Sales | Contact Sales | Large-scale latest-generation dedicated inference deployments |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Dedicated Inference (NVIDIA GB300 NVL72) Best for: Large-scale latest-generation dedicated inference deployments
| |||
| Enterprise | Contact Sales | Contact Sales | Large-scale enterprise deployments |
| Verified pricing · last checked July 2026 · 4 sources
Get this price at Together AI →
| |||
| What's included at Enterprise Best for: Large-scale enterprise deployments
| |||
View all features by plan (compare side-by-side)
Serverless Inference
- Pay-as-you-go per-token pricing
- Serverless inference across chat, vision, image, audio, video, transcription, embeddings, rerank, and moderation models
- Cached input pricing for select models
- Batch API pricing available
- Displayed prices may vary by resolution or duration for non-text modalities
Dedicated Inference (NVIDIA HGX H100)
- Single-tenant GPU instances
- Guaranteed performance
- Support for custom models
- Autoscaling and traffic spike handling
- On-demand pay-as-you-go pricing
Dedicated Inference (NVIDIA HGX H200)
- Single-tenant GPU instances
- Guaranteed performance
- Support for custom models
- Autoscaling and traffic spike handling
- Contact sales pricing
Dedicated Inference (NVIDIA HGX B200)
- Single-tenant GPU instances
- Guaranteed performance
- Support for custom models
- Autoscaling and traffic spike handling
- On-demand pay-as-you-go pricing
Dedicated Inference (NVIDIA HGX B300)
- Single-tenant GPU instances
- Guaranteed performance
- Support for custom models
- Autoscaling and traffic spike handling
- Contact sales pricing
Dedicated Inference (NVIDIA GB200 NVL72)
- Single-tenant GPU instances
- Guaranteed performance
- Support for custom models
- Autoscaling and traffic spike handling
- Contact sales pricing
Dedicated Inference (NVIDIA GB300 NVL72)
- Single-tenant GPU instances
- Guaranteed performance
- Support for custom models
- Autoscaling and traffic spike handling
- Contact sales pricing
Enterprise
- Volume discounts
- Dedicated support
- Custom SLAs
- Private deployments
Track Together AI pricing
Get an email when Together AI's pricing changes — plus the weekly SaaS Price Watch: verified price changes and deals across 3,000+ products. One-click unsubscribe.
You're on the list — first digest lands Tuesday.
Together AI uses custom pricing as of August 2026 with 8 plans available. Contact Together AI directly for a personalized quote. The median contract is $727/year based on 4 third-party buyer reports.
Use the interactive pricing calculator to estimate your exact cost based on team size and requirements.
- Free tier: No free tier available
Together AI offers 8 pricing tiers: Serverless Inference, Dedicated Inference (NVIDIA HGX H100), Dedicated Inference (NVIDIA HGX H200), Dedicated Inference (NVIDIA HGX B200), Dedicated Inference (NVIDIA HGX B300), Dedicated Inference (NVIDIA GB200 NVL72), Dedicated Inference (NVIDIA GB300 NVL72), Enterprise. The Dedicated Inference (NVIDIA HGX H100) plan is consistent high-volume inference on dedicated h100 gpus.
Compared to other llm api providers software, Together AI is positioned at the budget-friendly price point.
- Median contract: $727/yr from 4 buyer reports Directional
- 4 documented hidden costs beyond list price
How much does Together AI cost?
Together AI Pricing Overview
Together AI uses custom pricing — contact their sales team for a quote. The Serverless Inference plan requires contacting sales for a custom quote and is designed for variable-volume api inference usage. The Dedicated Inference (NVIDIA HGX H100) plan requires contacting sales for a custom quote and is designed for consistent high-volume inference on dedicated h100 gpus. The Dedicated Inference (NVIDIA HGX H200) plan requires contacting sales for a custom quote and is designed for high-throughput dedicated inference on h200 gpus. The Dedicated Inference (NVIDIA HGX B200) plan requires contacting sales for a custom quote and is designed for high-performance dedicated inference on b200 gpus. The Dedicated Inference (NVIDIA HGX B300) plan requires contacting sales for a custom quote and is designed for latest-generation dedicated inference deployments. The Dedicated Inference (NVIDIA GB200 NVL72) plan requires contacting sales for a custom quote and is designed for large-scale dedicated inference deployments. The Dedicated Inference (NVIDIA GB300 NVL72) plan requires contacting sales for a custom quote and is designed for large-scale latest-generation dedicated inference deployments. The Enterprise plan requires contacting sales for a custom quote and is designed for large-scale enterprise deployments.
Together AI with a Not applicable in the same way as fixed-term contracts. minimum commitment, requiring Not applicable in the same way as fixed-term contracts. notice to cancel.
The median Together AI customer pays $727/year based on 4 verified purchases.
There are at least 4 documented hidden costs beyond Together AI's list price, including implementation, training, and add-on fees.
This pricing was last verified in July 30, 2026 from 4 independent sources.
Together AI offers usage-based inference pricing through its Serverless tier, with dedicated GPU options (1x H100, 1x H200, 1x B200) and an Enterprise plan available at custom-quoted rates. Serverless pricing varies by model — the provider median is $0.50/1M input tokens and $1.20/1M output tokens across 23 tracked models, with individual models ranging from $0.02/1M input tokens for lightweight options up to $2.00/1M for large-scale coding models. Dedicated instances and Enterprise pricing require direct contact with Together AI's sales team.
How Together AI Pricing Compares
Compare Together AI pricing against top alternatives in LLM API Providers.
Usage-Based Rates
Per-unit pricing for Together AI API usage.
Serverless Inference
| Model | Input | Output | Cached | Per |
|---|---|---|---|---|
| deepseek-v4-pro | $1.74 | $3.48 | $0.200 | 1M tokens |
| minimax-m3 | $0.300 | $1.20 | $0.060 | 1M tokens |
| kimi-k2-7-code | $0.950 | $4.00 | $0.190 | 1M tokens |
| glm-5-2 | $1.40 | $4.40 | $0.260 | 1M tokens |
| inkling | $1.00 | $4.05 | $0.170 | 1M tokens |
| kimi-k3 | $3.00 | $15.00 | $0.300 | 1M tokens |
| gemma-4-31b | $0.390 | $0.970 | — | 1M tokens |
| nvidia-nemotron-3-ultra | $0.600 | $3.60 | $0.200 | 1M tokens |
| qwen3-7-plus | $0.320 | $1.28 | — | 1M tokens |
| lfm2-5-8b-a1b | $0.030 | $0.120 | — | 1M tokens |
| kimi-k2-6 | $1.20 | $4.50 | $0.200 | 1M tokens |
| qwen3-7-max | $1.25 | $3.75 | $0.130 | 1M tokens |
| gpt-oss-120b | $0.150 | $0.600 | — | 1M tokens |
| qwen3-5-397b-a17b | $0.600 | $3.60 | $0.350 | 1M tokens |
| qwen3-5-9b | $0.170 | $0.250 | — | 1M tokens |
| gemma-4-31b-it-pearl | $0.280 | $0.860 | — | 1M tokens |
| cogito-v2-1-671b | $1.25 | $1.25 | — | 1M tokens |
| rnj-1-instruct | $0.150 | $0.150 | — | 1M tokens |
| llama-3-3-70b | $1.04 | $1.04 | — | 1M tokens |
| gemma-3n-e4b-instruct | $0.060 | $0.120 | — | 1M tokens |
| gpt-oss-20b | $0.050 | $0.200 | — | 1M tokens |
| glm-5-1 | $1.40 | $4.40 | $0.260 | 1M tokens |
| minimax-m2-7 | $0.300 | $1.20 | $0.060 | 1M tokens |
| qwen3-6-plus | $0.500 | $3.00 | — | 1M tokens |
| qwen2-5-7b-instruct-turbo | $0.300 | $0.300 | — | 1M tokens |
| llama-3-8b-instruct-lite | $0.140 | $0.140 | — | 1M tokens |
| qwen3-235b-a22b-instruct-2507-fp8-throughput | $0.200 | $0.600 | — | 1M tokens |
- Top serverless inference models listed from pricing page
- Cached input pricing available for select models
- Batch API pricing available
Dedicated Inference (NVIDIA HGX H100)
| Model | Unit | Rate |
|---|---|---|
| h100 | hour | $5.49 |
- All prices per GPU per hour
Dedicated Inference (NVIDIA HGX B200)
| Model | Unit | Rate |
|---|---|---|
| b200 | hour | $8.99 |
- All prices per GPU per hour
Compare Together AI vs Alternatives
Before committing to Together AI, compare pricing with these 3 alternatives in the same category.
What Companies Actually Pay for Together AI
The median Together AI buyer pays $727/year based on 4 verified purchase transactions.
| Model | Input /1M | Output /1M | Blended /1M |
|---|---|---|---|
| togetherai_glm-5-1 | $1.40 | $4.40 | $2.15 |
| togetherai_qwen3-coder-480b-a35b-instruct_fp8 | $2.00 | $2.00 | $2.00 |
| togetherai_glm-5_fp4 | $1.00 | $3.20 | $1.55 |
| togetherai_deepseek-v3-0324 | $1.25 | $1.25 | $1.25 |
| togetherai_cogito-v2-1-reasoning | $1.25 | $1.25 | $1.25 |
How Together AI Pricing Compares
| Software | Starting Price | Top Price |
|---|---|---|
| Together AI | $0.03 per million tokens / hour | $15 per million tokens / hour |
| Amazon Bedrock | $0.07 per million tokens | $75 per million tokens |
| Anyscale | Free | $4.9591 per million tokens |
| Baidu ERNIE API | $0.1 per million tokens | $10 per million tokens |
| Cerebras Inference API | $0.1 per million tokens | $6 per million tokens |
| Claude API | $0.1 per million tokens | $75 per million tokens |
Detailed pricing comparisons:
Together AI Contract Terms
Together AI contracts do not auto-renew. Changes require Not applicable in the same way as fixed-term contracts.. These terms are sourced from verified buyer experiences.
Together AI Price History
Pricing changes CostBench has tracked for Together AI, verified across 2 snapshots going back to 2026Q2.
-
Together AI price increase
Highest tier increased from $10 to $15/month
+50.75% -
Together AI tier restructure
Changed from 5 to 8 pricing tiers
-
Together AI tier renamed
"Serverless" renamed to "Serverless Inference"
-
Together AI tier renamed
"Dedicated (1x H100)" renamed to "Dedicated Inference (NVIDIA HGX H100)"
-
Together AI tier renamed
"Dedicated (1x H200)" renamed to "Dedicated Inference (NVIDIA HGX H200)"
-
Together AI tier renamed
"Dedicated (1x B200)" renamed to "Dedicated Inference (NVIDIA HGX B200)"
-
Together AI tier renamed
"Enterprise" renamed to "Dedicated Inference (NVIDIA HGX B300)"
View Together AI's full price history →
Together AI Pricing FAQ
01 How much does Together AI cost?
Together AI offers serverless inference starting at $0.10 per million tokens for small models. Mid-range models cost $0.50–1.00/M tokens, and large models like DeepSeek-R1 cost $3.00/M tokens. Dedicated GPU deployments start at $3.99/hr (1x H100) or $9.95/hr (1x B200). Batch processing saves 40–50%.
02 Does Together AI have a free tier?
Together AI does not advertise a permanent free tier or free credits on their pricing page. They offer pay-as-you-go Serverless pricing with no minimum commitment, so you only pay for what you use.
03 What models does Together AI support?
Together AI supports a wide range of open-source models including Llama, DeepSeek, Qwen, Mistral, and Kimi. They also offer image generation (FLUX, Stable Diffusion), video (Google Veo 2.0), audio transcription, text-to-speech, and embedding models.
04 Together AI vs Fireworks AI: which is cheaper?
Both offer similar serverless per-token pricing starting around $0.10/M tokens for small models. Fireworks AI gives new users $1 in free credits. For dedicated GPU hosting, Together AI's H100 is $3.99/hr versus Fireworks AI's A100 at $2.90/hr, making Fireworks slightly cheaper for dedicated compute at equivalent GPU tiers.
05 What is Together AI's Dedicated GPU pricing?
Together AI's Dedicated GPU hosting starts at $3.99/hr for a 1x H100 (single-tenant) and $9.95/hr for a 1x B200 (latest generation). Dedicated deployments are best for consistent high-volume inference where you need guaranteed resources and custom model hosting.
06 What are the cheapest models available on Together AI?
Based on Artificial Analysis data, the most affordable models on Together AI's Serverless tier start at $0.02/1M input tokens (Gemma 3n E4B) and $0.03/1M input tokens (LFM2 24B A2B). The provider median blended rate across all 23 tracked models is $0.875/1M tokens.
07 Does Together AI offer dedicated GPU instances?
Yes. Together AI offers dedicated GPU instances on three hardware tiers: 1x H100, 1x H200, and 1x B200. All dedicated instance pricing is custom-quoted. An Enterprise plan is also available for larger-scale deployments requiring custom SLAs or support.
Is this pricing incorrect? — we'll verify and update it.