Groq

Groq pricing bills per million tokens, in arrears, with no prepaid balance and no credit system.

Pricing Model:

Pricing Model:

Usage, pay as you go, billed in arrears

Usage, pay as you go, billed in arrears

Usage, pay as you go, billed in arrears

Packaging Model:

Packaging Model:

Freemium, Good / Better / Best (GBB)

Freemium, Good / Better / Best (GBB)

Freemium, Good / Better / Best (GBB)

Credit Model:

Credit Model:

None. No prepaid balance and no credit wallet

None. No prepaid balance and no credit wallet

None. No prepaid balance and no credit wallet

Updated on:

Groq pricing: the speed vendor charges the going rate

Groq pricing bills per million tokens, in arrears, with no prepaid balance and no credit system. Four chat models carry a published rate. Everything else, including both Llama models and MiniMax M2.7, says Contact Sales. Speed is the pitch, but the rates sit mid-pack, and the marketing site no longer hosts a pricing page.

Key takeaways

  • GPT-OSS 120B costs $0.15 per million input tokens and $0.60 output on Groq, exactly the rate Together AI, Amazon Bedrock and Nebius charge for the same model.

  • Of the 13 models Groq lists publicly, 10 carry a price and 3 say Contact Sales, including both Llama models that made Groq's name.

  • The free tier grants no credit. It's the paid API at tight limits: 8K tokens per minute and 200K per day on GPT-OSS 120B, against 250K per minute and no daily cap on Developer.

  • Batch cuts 50% and cached input cuts 50%, and the docs say the two discounts don't stack.

Groq pricing in 2026

Plan

Price

What's included

Metered limit

Free

$0

Same per-token rates, no card required

30 RPM, 1K RPD, 8K TPM on GPT-OSS 120B

Developer

$0 + usage

Batch, Flex tier, spend limits, chat support, 100 MB audio files

1K RPM, 500K RPD, 250K TPM, no daily token cap

Enterprise

Custom

Performance tier, Llama and MiniMax access, 99.9% availability SLA

Contracted

Model

Input / 1M

Cached input

Output / 1M

openai/gpt-oss-120b

$0.15

$0.075

$0.60

openai/gpt-oss-20b

$0.075

$0.037

$0.30

openai/gpt-oss-safeguard-20b

$0.075

$0.037

$0.30

qwen/qwen3.8-27b (preview)

$0.80

Not supported

$4.00

whisper-large-v3 / turbo

$0.111 / $0.04 per audio hour

n/a

n/a

Orpheus English / Arabic

$22 / $40 per 1M characters

n/a

n/a

What Groq actually meters

Groq meters three different units and keeps them apart. Text generation bills input and output tokens at separate rates, with a third rate for cached input on the GPT-OSS family. Speech to text bills by the hour of audio processed, not by the token. Text to speech bills per million characters, so Arabic at $40 runs 1.8 times the English rate of $22.

The published catalog is narrower than the pitch suggests. Thirteen models appear on the models page. Llama 3.1 8B, Llama 3.3 70B and MiniMax M2.7 are tagged Enterprise and show Contact Sales in both the price and the rate limit column. That leaves four general chat models with a public per-token rate, two of which are the same GPT-OSS 20B weights under different names.

Cached input is automatic, needs no code change and carries no fee. It applies to the three GPT-OSS models only, expires after two hours unused, and cached tokens don't count against rate limits. The headline discount is 50%: $0.075 against $0.15 on the 120B. On the 20B, Groq publishes $0.037 against $0.075, rounding the half-cent down in your favour.

Groq's own pricing page is gone. groq.com/pricing returns a 308 to the homepage and the plan comparison needs a login, so rates survive only in the docs.

What happens when you hit the limit

Groq blocks rather than throttles, and it bills you before the month ends. New Developer accounts run progressive billing: the card gets charged the moment lifetime usage crosses $1, $10, $100, $500 and $1,000. Past $1,000 the account settles monthly. Indian billing addresses get a different ladder, $1, then $10, then every $100 for good.

Spend limits are the real guardrail. Set a monthly cap and every key in the organisation starts returning a 400 with code blocked_api_access once you reach it. Spend tracking lags 10 to 15 minutes, so Groq says plainly you may overshoot. The limit resets on the first.

Rate limits bite before spend does: you get a 429 with a retry-after header. Flex tier raises limits tenfold at the same price and fails fast with a 498 capacity_exceeded when capacity runs out.

How Groq pricing has changed

Date

Milestone

Source

18 Apr 2026

MiniMax M2.5 and Qwen3-VL 32B ship Enterprise-only, with no published rate

Vendor

30 Jan 2026

PlayAI text to speech retires platform-wide; Orpheus replaces it at $22 and $40 per 1M characters

Vendor

29 Oct 2025

GPT-OSS-Safeguard 20B launches at $0.075 input, $0.30 output, caching on from day one

Vendor

21 Oct 2025

Prompt caching reaches GPT-OSS 120B, cached input $0.075 against $0.15

Vendor

25 Sep 2025

Prompt caching reaches GPT-OSS 20B, cached input $0.037 against $0.075

Vendor

5 Sep 2025

Kimi K2-0905 lands at $1.00 input and $3.00 output

Vendor

20 Aug 2025

Prompt caching introduced, Kimi K2 first, 50% off cached tokens

Vendor

Flexprice’s Take

Groq sells latency and charges list, so the premium you pay for LPU speed isn't in the token rate, it's in how few models carry a rate at all.

The comparison is cleaner than usual because GPT-OSS 120B runs everywhere. Groq charges $0.15 and $0.60. Together AI charges $0.15 and $0.60. Amazon Bedrock, Nebius and SiliconFlow charge $0.15 and $0.60. Cerebras, the other speed-first vendor, charges $0.35 input, which is 2.3 times Groq's. Against commodity hosts the gap is real: CoreWeave serves the same model at $0.030 input, so Groq runs 5 times that, and 3.5 times on output against $0.17. The market has priced speed at roughly a 5x floor rather than a surcharge Groq invented.

What stings is the catalog. Three of 13 models, including both Llamas, now answer Contact Sales. A serving platform whose public price list covers four chat models is a sales motion wearing an API's clothes.

Mechanically the billing is sound. No credits, no expiry rules, no wallet to reconcile, and spend limits that hard-block at the organisation level.

Best For

Latency-bound products on GPT-OSS wanting a published rate and no prepaid balance.

Watch Out For

Budgeting against a Llama rate you can't read without a sales call.

Manish Choudhary

CEO & Co-founder, Flexprice

Charging your own customers for inference and need per-model cost and margin per account?

Flexprice tracks AI cost and margin per customer.

Flexprice’s Take

Groq sells latency and charges list, so the premium you pay for LPU speed isn't in the token rate, it's in how few models carry a rate at all.

The comparison is cleaner than usual because GPT-OSS 120B runs everywhere. Groq charges $0.15 and $0.60. Together AI charges $0.15 and $0.60. Amazon Bedrock, Nebius and SiliconFlow charge $0.15 and $0.60. Cerebras, the other speed-first vendor, charges $0.35 input, which is 2.3 times Groq's. Against commodity hosts the gap is real: CoreWeave serves the same model at $0.030 input, so Groq runs 5 times that, and 3.5 times on output against $0.17. The market has priced speed at roughly a 5x floor rather than a surcharge Groq invented.

What stings is the catalog. Three of 13 models, including both Llamas, now answer Contact Sales. A serving platform whose public price list covers four chat models is a sales motion wearing an API's clothes.

Mechanically the billing is sound. No credits, no expiry rules, no wallet to reconcile, and spend limits that hard-block at the organisation level.

Best For

Latency-bound products on GPT-OSS wanting a published rate and no prepaid balance.

Watch Out For

Budgeting against a Llama rate you can't read without a sales call.

Manish Choudhary

CEO & Co-founder, Flexprice

Charging your own customers for inference and need per-model cost and margin per account?

Flexprice tracks AI cost and margin per customer.

Customer
Sentiment Highlights

"It's common for me to run workloads in Groq that cost less than $100, while the same workload can approach $1,000 on Bedrock or Gemini"

Groq user comparing provider bills, Hacker News, December 2025

"Cheap enough for now, but of all the companies selling inference at a loss, Cerebras and Groq are probably losing the most per token"

Hacker News commenter on inference economics, November 2025

Frequently Asked Questions

Frequently Asked Questions

How much does Groq cost?

Is Groq more expensive than other providers?

Does Groq have a free tier?

Does Groq sell dedicated capacity?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 500+ Builders on Slack

Join the Flexprice Community on Slack