All Together AI Plans & Pricing

Plan Monthly Annual Best For
View all features by plan (compare side-by-side)

Serverless Inference

  • Pay-as-you-go per-token pricing
  • Serverless inference across chat, vision, image, audio, video, transcription, embeddings, rerank, and moderation models
  • Cached input pricing for select models
  • Batch API pricing available
  • Displayed prices may vary by resolution or duration for non-text modalities

Dedicated Inference (NVIDIA HGX H100)

  • Single-tenant GPU instances
  • Guaranteed performance
  • Support for custom models
  • Autoscaling and traffic spike handling
  • On-demand pay-as-you-go pricing

Dedicated Inference (NVIDIA HGX H200)

  • Single-tenant GPU instances
  • Guaranteed performance
  • Support for custom models
  • Autoscaling and traffic spike handling
  • Contact sales pricing

Dedicated Inference (NVIDIA HGX B200)

  • Single-tenant GPU instances
  • Guaranteed performance
  • Support for custom models
  • Autoscaling and traffic spike handling
  • On-demand pay-as-you-go pricing

Dedicated Inference (NVIDIA HGX B300)

  • Single-tenant GPU instances
  • Guaranteed performance
  • Support for custom models
  • Autoscaling and traffic spike handling
  • Contact sales pricing

Dedicated Inference (NVIDIA GB200 NVL72)

  • Single-tenant GPU instances
  • Guaranteed performance
  • Support for custom models
  • Autoscaling and traffic spike handling
  • Contact sales pricing

Dedicated Inference (NVIDIA GB300 NVL72)

  • Single-tenant GPU instances
  • Guaranteed performance
  • Support for custom models
  • Autoscaling and traffic spike handling
  • Contact sales pricing

Enterprise

  • Volume discounts
  • Dedicated support
  • Custom SLAs
  • Private deployments
Pricing Alerts

Track Together AI pricing

Get an email when Together AI's pricing changes — plus the weekly SaaS Price Watch: verified price changes and deals across 3,000+ products. One-click unsubscribe.

Compare Together AI with alternativesAdjust seats, lock a tier, add up to 2 more products side-by-side. Shareable URL.
Quick Answer
Last verified:
High confidence

Together AI uses custom pricing as of August 2026 with 8 plans available. Contact Together AI directly for a personalized quote. The median contract is $727/year based on 4 third-party buyer reports.

Use the interactive pricing calculator to estimate your exact cost based on team size and requirements.

  • Free tier: No free tier available

Together AI offers 8 pricing tiers: Serverless Inference, Dedicated Inference (NVIDIA HGX H100), Dedicated Inference (NVIDIA HGX H200), Dedicated Inference (NVIDIA HGX B200), Dedicated Inference (NVIDIA HGX B300), Dedicated Inference (NVIDIA GB200 NVL72), Dedicated Inference (NVIDIA GB300 NVL72), Enterprise. The Dedicated Inference (NVIDIA HGX H100) plan is consistent high-volume inference on dedicated h100 gpus.

Compared to other llm api providers software, Together AI is positioned at the budget-friendly price point.

  • Median contract: $727/yr from 4 buyer reports Directional
  • 4 documented hidden costs beyond list price

How much does Together AI cost?

Together AI uses custom pricing across 8 plans. Contact Together AI directly for a personalized quote. Plans include Serverless Inference (custom pricing), Dedicated Inference (NVIDIA HGX H100) (custom pricing), Dedicated Inference (NVIDIA HGX H200) (custom pricing), Dedicated Inference (NVIDIA HGX B200) (custom pricing), Dedicated Inference (NVIDIA HGX B300) (custom pricing), Dedicated Inference (NVIDIA GB200 NVL72) (custom pricing), Dedicated Inference (NVIDIA GB300 NVL72) (custom pricing), Enterprise (custom pricing).

Together AI Pricing Overview

Together AI uses custom pricing — contact their sales team for a quote. The Serverless Inference plan requires contacting sales for a custom quote and is designed for variable-volume api inference usage. The Dedicated Inference (NVIDIA HGX H100) plan requires contacting sales for a custom quote and is designed for consistent high-volume inference on dedicated h100 gpus. The Dedicated Inference (NVIDIA HGX H200) plan requires contacting sales for a custom quote and is designed for high-throughput dedicated inference on h200 gpus. The Dedicated Inference (NVIDIA HGX B200) plan requires contacting sales for a custom quote and is designed for high-performance dedicated inference on b200 gpus. The Dedicated Inference (NVIDIA HGX B300) plan requires contacting sales for a custom quote and is designed for latest-generation dedicated inference deployments. The Dedicated Inference (NVIDIA GB200 NVL72) plan requires contacting sales for a custom quote and is designed for large-scale dedicated inference deployments. The Dedicated Inference (NVIDIA GB300 NVL72) plan requires contacting sales for a custom quote and is designed for large-scale latest-generation dedicated inference deployments. The Enterprise plan requires contacting sales for a custom quote and is designed for large-scale enterprise deployments.

Together AI with a Not applicable in the same way as fixed-term contracts. minimum commitment, requiring Not applicable in the same way as fixed-term contracts. notice to cancel.

The median Together AI customer pays $727/year based on 4 verified purchases.

There are at least 4 documented hidden costs beyond Together AI's list price, including implementation, training, and add-on fees.

This pricing was last verified in July 30, 2026 from 4 independent sources.

Together AI offers usage-based inference pricing through its Serverless tier, with dedicated GPU options (1x H100, 1x H200, 1x B200) and an Enterprise plan available at custom-quoted rates. Serverless pricing varies by model — the provider median is $0.50/1M input tokens and $1.20/1M output tokens across 23 tracked models, with individual models ranging from $0.02/1M input tokens for lightweight options up to $2.00/1M for large-scale coding models. Dedicated instances and Enterprise pricing require direct contact with Together AI's sales team.

How Together AI Pricing Compares

Compare Together AI pricing against top alternatives in LLM API Providers.

Usage-Based Rates

Per-unit pricing for Together AI API usage.

Serverless Inference

Model Input Output Cached Per
deepseek-v4-pro $1.74 $3.48 $0.200 1M tokens
minimax-m3 $0.300 $1.20 $0.060 1M tokens
kimi-k2-7-code $0.950 $4.00 $0.190 1M tokens
glm-5-2 $1.40 $4.40 $0.260 1M tokens
inkling $1.00 $4.05 $0.170 1M tokens
kimi-k3 $3.00 $15.00 $0.300 1M tokens
gemma-4-31b $0.390 $0.970 1M tokens
nvidia-nemotron-3-ultra $0.600 $3.60 $0.200 1M tokens
qwen3-7-plus $0.320 $1.28 1M tokens
lfm2-5-8b-a1b $0.030 $0.120 1M tokens
kimi-k2-6 $1.20 $4.50 $0.200 1M tokens
qwen3-7-max $1.25 $3.75 $0.130 1M tokens
gpt-oss-120b $0.150 $0.600 1M tokens
qwen3-5-397b-a17b $0.600 $3.60 $0.350 1M tokens
qwen3-5-9b $0.170 $0.250 1M tokens
gemma-4-31b-it-pearl $0.280 $0.860 1M tokens
cogito-v2-1-671b $1.25 $1.25 1M tokens
rnj-1-instruct $0.150 $0.150 1M tokens
llama-3-3-70b $1.04 $1.04 1M tokens
gemma-3n-e4b-instruct $0.060 $0.120 1M tokens
gpt-oss-20b $0.050 $0.200 1M tokens
glm-5-1 $1.40 $4.40 $0.260 1M tokens
minimax-m2-7 $0.300 $1.20 $0.060 1M tokens
qwen3-6-plus $0.500 $3.00 1M tokens
qwen2-5-7b-instruct-turbo $0.300 $0.300 1M tokens
llama-3-8b-instruct-lite $0.140 $0.140 1M tokens
qwen3-235b-a22b-instruct-2507-fp8-throughput $0.200 $0.600 1M tokens
  • Top serverless inference models listed from pricing page
  • Cached input pricing available for select models
  • Batch API pricing available

Dedicated Inference (NVIDIA HGX H100)

Model Unit Rate
h100 hour $5.49
  • All prices per GPU per hour

Dedicated Inference (NVIDIA HGX B200)

Model Unit Rate
b200 hour $8.99
  • All prices per GPU per hour

Compare Together AI vs Alternatives

Before committing to Together AI, compare pricing with these 3 alternatives in the same category.

All Together AI alternatives & migration guides

What Companies Actually Pay for Together AI

The median Together AI buyer pays $727/year based on 4 verified purchase transactions.

What companies actually pay $727/yr Median across 4 third-party buyer reports Directional
Flagship models in this provider's catalog
Model Input /1M Output /1M Blended /1M
togetherai_glm-5-1 $1.40 $4.40 $2.15
togetherai_qwen3-coder-480b-a35b-instruct_fp8 $2.00 $2.00 $2.00
togetherai_glm-5_fp4 $1.00 $3.20 $1.55
togetherai_deepseek-v3-0324 $1.25 $1.25 $1.25
togetherai_cogito-v2-1-reasoning $1.25 $1.25 $1.25
Review scores
Third-party review aggregates, as of Jul 2026
Top pricing complaints
Limited AccessPayment IssuesPoor DocumentationTechnical Expertise Required
Source: aggregated from 4 third-party buyer reports for this quote-only vendor. Indicative market pricing — not official vendor pricing or contract-grade data.

How Together AI Pricing Compares

Software Starting Price Top Price
Together AI $0.03 per million tokens / hour $15 per million tokens / hour
Amazon Bedrock $0.07 per million tokens $75 per million tokens
Anyscale Free $4.9591 per million tokens
Baidu ERNIE API $0.1 per million tokens $10 per million tokens
Cerebras Inference API $0.1 per million tokens $6 per million tokens
Claude API $0.1 per million tokens $75 per million tokens

4 Together AI Hidden Costs Beyond the List Price

Beyond the listed price, Together AI has at least 4 documented hidden costs that can significantly increase total cost of ownership.

Watch for 4 hidden costs
  • Fine-tune Hosting $4,700 per month
    high 1 source
    industry "For instance, keeping a single H100 GPU running 24/7 for a fine-tuned model costs approximately $4,700 per month, regardless of actual usage."
  • Engineering Time
    high 1 source
    industry "Engineering Time: The effort required for integrating inference, retrieval, retries, and user interface development can be a substantial, often overlooked, cost."
  • Platform Access and Minimums $5, $4.00
    low 1 source
    industry "Platform Access and Minimums: Together AI does not offer a free trial; platform access requires a minimum $5 credit purchase."
  • Payment Friction
    medium 1 source
    industry "Payment Friction: Users have reported issues with payment methods, such as updating cards or resolving failed charges, which can lead to service interruptions if not managed with monitoring and alerts for account-level spending caps."
Tip

Ask your Together AI sales rep about these costs upfront. Getting them in writing before signing can save you from surprise charges later.

Full hidden costs breakdown →

Intelligence sourced from 1 independent sources
industry
Key claims include inline source attribution. Data verified against multiple independent sources. 6 source citations total.

Together AI Contract Terms

Together AI contracts do not auto-renew. Changes require Not applicable in the same way as fixed-term contracts.. These terms are sourced from verified buyer experiences.

Contract Terms
Cancellation Notice Not applicable in the same way as fixed-term contracts.
Minimum Commitment Not applicable in the same way as fixed-term contracts.
Based on 1 verified source

Together AI Price History

Pricing changes CostBench has tracked for Together AI, verified across 2 snapshots going back to 2026Q2.

  1. Together AI price increase

    Highest tier increased from $10 to $15/month

    +50.75%
  2. Together AI tier restructure

    Changed from 5 to 8 pricing tiers

  3. Together AI tier renamed

    "Serverless" renamed to "Serverless Inference"

  4. Together AI tier renamed

    "Dedicated (1x H100)" renamed to "Dedicated Inference (NVIDIA HGX H100)"

  5. Together AI tier renamed

    "Dedicated (1x H200)" renamed to "Dedicated Inference (NVIDIA HGX H200)"

  6. Together AI tier renamed

    "Dedicated (1x B200)" renamed to "Dedicated Inference (NVIDIA HGX B200)"

  7. Together AI tier renamed

    "Enterprise" renamed to "Dedicated Inference (NVIDIA HGX B300)"

Together AI Pricing FAQ

01 How much does Together AI cost?

Together AI offers serverless inference starting at $0.10 per million tokens for small models. Mid-range models cost $0.50–1.00/M tokens, and large models like DeepSeek-R1 cost $3.00/M tokens. Dedicated GPU deployments start at $3.99/hr (1x H100) or $9.95/hr (1x B200). Batch processing saves 40–50%.

02 Does Together AI have a free tier?

Together AI does not advertise a permanent free tier or free credits on their pricing page. They offer pay-as-you-go Serverless pricing with no minimum commitment, so you only pay for what you use.

03 What models does Together AI support?

Together AI supports a wide range of open-source models including Llama, DeepSeek, Qwen, Mistral, and Kimi. They also offer image generation (FLUX, Stable Diffusion), video (Google Veo 2.0), audio transcription, text-to-speech, and embedding models.

04 Together AI vs Fireworks AI: which is cheaper?

Both offer similar serverless per-token pricing starting around $0.10/M tokens for small models. Fireworks AI gives new users $1 in free credits. For dedicated GPU hosting, Together AI's H100 is $3.99/hr versus Fireworks AI's A100 at $2.90/hr, making Fireworks slightly cheaper for dedicated compute at equivalent GPU tiers.

05 What is Together AI's Dedicated GPU pricing?

Together AI's Dedicated GPU hosting starts at $3.99/hr for a 1x H100 (single-tenant) and $9.95/hr for a 1x B200 (latest generation). Dedicated deployments are best for consistent high-volume inference where you need guaranteed resources and custom model hosting.

06 What are the cheapest models available on Together AI?

Based on Artificial Analysis data, the most affordable models on Together AI's Serverless tier start at $0.02/1M input tokens (Gemma 3n E4B) and $0.03/1M input tokens (LFM2 24B A2B). The provider median blended rate across all 23 tracked models is $0.875/1M tokens.

07 Does Together AI offer dedicated GPU instances?

Yes. Together AI offers dedicated GPU instances on three hardware tiers: 1x H100, 1x H200, and 1x B200. All dedicated instance pricing is custom-quoted. An Enterprise plan is also available for larger-scale deployments requiring custom SLAs or support.

Is this pricing incorrect? — we'll verify and update it.