Together AI

Together AI pricing charges three meters for the same model: per token on serverless, per GPU-minute on dedicated endpoints, and per token trained on fine-tuning.

Pricing Model:

Pricing Model:

Usage, on three separate meters

Usage, on three separate meters

Usage, on three separate meters

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

Prepaid balance, no expiry, optional auto-recharge

Prepaid balance, no expiry, optional auto-recharge

Prepaid balance, no expiry, optional auto-recharge

Updated on:

Together AI pricing: three ways to buy the same model

Together AI pricing charges three meters for the same model: per token on serverless, per GPU-minute on dedicated endpoints, and per token trained on fine-tuning. A fourth option, provisioned throughput, bills $0.05 per unit per minute on a monthly commitment. Nothing is bundled, so what you pay depends less on the model you pick than on the meter you're standing on.

Key takeaways

  • Serverless runs $0.14 to $3.00 per million input tokens, and several models price cached input separately at 42% to 81% off.

  • Dedicated inference bills per GPU-hour by the minute, per ready replica, and drops to zero when a deployment scales to zero.

  • Fine-tuning bills every token processed across training and validation, from $0.34 to $40.00 per million on a LoRA supervised run, and every model carries its own minimum charge.

  • There's no free tier. Together is fully prepaid, requires a $5 minimum credit purchase, and suspends API access when the balance hits zero.

Together AI pricing in 2026

Mode

Unit

Rate

Notes

 

Serverless

Per 1M tokens

Kimi K3 $3.00 in / $15.00 out; DeepSeek V4 Flash $0.14 / $0.28

Cached input priced separately on most models

Batch

Per 1M tokens

Up to 50% off serverless

Discount applies to selected models only

Dedicated (DMI)

Per GPU-hour, billed per minute

H100 80GB $3.99 promotional to 30 Sep 2026, $5.49 list; B200 180GB $8.99; H200, B300, GB300 Custom

Per ready replica

Provisioned throughput

Per PTU-minute

$0.05

One month minimum, sales only

Fine-tuning

Per 1M tokens trained

LoRA supervised $0.34 to $40.00, DPO $0.84 to $100.00

Minimum charge per model, $4.00 to $60.00

GPU clusters

Per GPU-hour

H100 $3.99 on-demand, $1.99 preemptible

Reserved terms cut to $3.19 at 91 to 180 days

What Together AI actually meters

Together meters three distinct things and never mixes them into one bill line. On serverless it's the token, split into input, cached input and output, and cached input carries its own published rate on most models rather than a blanket multiplier. MiniMax M3 charges $0.30 input against $0.06 cached. GLM-5.2 charges $1.40 against $0.26, an 81% discount, while Qwen3.5-397B-A17B charges $0.60 against $0.35 and saves only 42%. The discount isn't uniform, so caching pays back very differently model to model.

On dedicated model inference the meter switches to hardware. Together bills per GPU-hour, measured by the minute, per replica, and only while a replica is ready to serve. Provisioning, cold starts and DEGRADED replicas don't bill. Token volume is irrelevant here: the model affects cost only through the GPU count it needs.

Fine-tuning meters tokens processed, defined as (n_epochs × training tokens) + (n_evals × validation tokens). Disable packing and it recalculates as dataset length multiplied by max_seq_length, which moves the number a long way. Cancelled jobs pay for completed steps only, and failed jobs get fully refunded.

How credits work

Together is fully prepaid, and credits are the only currency. You buy a balance, minimum $5, and every service draws from it: API calls, dedicated deployments, fine-tuning and evaluation jobs. There's no free trial and no signup grant. Credits don't expire, and Together commits to advance notice if that changes, which is more than most prepaid vendors say.

Auto-recharge tops the balance back to a target when it falls below a threshold you set, as a single transaction on your default payment method. It works only when that default is a card: setting a US bank account as default turns auto-recharge off automatically. One sharp edge, credits bought after an invoice is generated can't clear that invoice or any past due balance.

What happens when you hit the limit

The balance hitting zero is the limit, and Together suspends API access until you add credits. That's a hard stop rather than an overage charge, which is what prepaid should mean and frequently doesn't. Creating a dedicated endpoint or a fine-tuning job needs enough balance up front to cover the cost.

Rate limits behave differently. They're dynamic, applied per model rather than per account, and grow with sustained reliable traffic. The old Build Tier 1 to 5, Scale and Enterprise labels are retired, so there's no tier to buy into. Every serverless response returns headers carrying the current limit.

How Together AI pricing has changed

Date

Milestone

Source

 

10 Sep 2026

Preemptible GPU cluster compute enters public preview at a flat discount to on-demand, metered every one to two minutes

Vendor

1 Sep 2026

H100 80GB dedicated endpoint hardware drops to $3.99 per hour, down from $5.49

Vendor

25 Jun 2026

Seedance 2.0 adds a 4K tier at $0.836 per second, against $0.40 at 1080p

Vendor

9 Jun 2026

Cached input pricing added for GLM-5.1 and Qwen3.5-397B-A17B. DeepSeek V4 Pro cut from $2.10 to $1.74 input and $4.40 to $3.48 output

Vendor

29 May 2026

Three models raised. Llama 3.3 70B goes $0.88 to $1.04, Qwen3.5 9B goes $0.10 to $0.17 input

Vendor

10 Mar 2026

First cached input rate published: MiniMax M2.5 at $0.06 per 1M, 80% off standard input

Vendor

Flexprice’s Take

Together's three-meter structure is honest about what it actually costs to serve a model, and the only thing wrong with it is a promotion the docs forget to mention.

Three meters is the honest design. Tokens, GPU-minutes and training tokens are different costs, and collapsing them makes one workload shape subsidise another.

Billing dedicated replicas only while they're ready is the detail most vendors get wrong in their own favour. Together doesn't bill provisioning or cold starts.

Check the H100 rate before you budget. The docs list dedicated H100 hardware at $3.99 an hour with no conditions, the changelog calls it a cut from $5.49, the pricing page calls it a promotion valid until 30 September 2026, and the same docs page's worked example still calculates with $5.49. Budget from the docs and you may pick a rate that won't survive the month.

Batch is oversold too. "Up to 50% lower cost" runs across the marketing site while the docs list two discounted models.

Pick the mode before the model. Serverless wins on bursty traffic, dedicated wins once a replica stays busy most of the day, and that gap dwarfs the one between two similar models.

Best For

Teams whose traffic shape is known well enough to pick the right meter.

Watch Out For

Budgeting H100 dedicated capacity at the promotional rate past 30 September 2026.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling open model access and need per-model cost and margin per customer?

Flexprice meters it.

Flexprice’s Take

Together's three-meter structure is honest about what it actually costs to serve a model, and the only thing wrong with it is a promotion the docs forget to mention.

Three meters is the honest design. Tokens, GPU-minutes and training tokens are different costs, and collapsing them makes one workload shape subsidise another.

Billing dedicated replicas only while they're ready is the detail most vendors get wrong in their own favour. Together doesn't bill provisioning or cold starts.

Check the H100 rate before you budget. The docs list dedicated H100 hardware at $3.99 an hour with no conditions, the changelog calls it a cut from $5.49, the pricing page calls it a promotion valid until 30 September 2026, and the same docs page's worked example still calculates with $5.49. Budget from the docs and you may pick a rate that won't survive the month.

Batch is oversold too. "Up to 50% lower cost" runs across the marketing site while the docs list two discounted models.

Pick the mode before the model. Serverless wins on bursty traffic, dedicated wins once a replica stays busy most of the day, and that gap dwarfs the one between two similar models.

Best For

Teams whose traffic shape is known well enough to pick the right meter.

Watch Out For

Budgeting H100 dedicated capacity at the promotional rate past 30 September 2026.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling open model access and need per-model cost and margin per customer?

Flexprice meters it.

Customer
Sentiment Highlights

"They delivered a 2x reduction in latency and cut our costs by approximately a third"

Customer quoted on together.ai's own customer stories page (vendor-published)

Frequently Asked Questions

Frequently Asked Questions

How much does Together AI cost?

Does Together AI have a free tier?

How does Together AI bill dedicated endpoints?

Do Together AI credits expire?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack