Azure OpenAI

Azure OpenAI pricing is the only entry in this index where you can pay for a model without sending it anything.

Pricing Model:

Pricing Model:

Four models for the same models: standard, priority, provisioned, batch

Four models for the same models: standard, priority, provisioned, batch

Four models for the same models: standard, priority, provisioned, batch

Packaging Model:

Packaging Model:

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Good / Better / Best (GBB)

Credit Model:

Credit Model:

None. Azure Reservations replace prepaid credit

None. Azure Reservations replace prepaid credit

None. Azure Reservations replace prepaid credit

Updated on:

Azure OpenAI pricing: paying for capacity instead of consumption

Azure OpenAI pricing is the only entry in this index where you can pay for a model without sending it anything. Alongside the usual per-token billing, Microsoft sells Provisioned Throughput Units that reserve dedicated capacity and bill hourly whether or not a request arrives. Choosing between the four deployment types is the actual pricing decision, and it turns on traffic shape rather than volume.

Key takeaways

  • The same model carries four billing modes: standard per-token, priority per-token, provisioned per PTU-hour, and discounted batch.

  • Provisioned deployments bill from the moment they are created until they are deleted, regardless of tokens consumed.

  • Hourly provisioned pricing runs $1 per PTU on Global, $1.10 on Data Zone and $2 on Regional, with one-month and one-year reservations cutting that substantially.

  • Reservations are locked to a deployment type, so Global, Data Zone and Regional each need their own.

Azure OpenAI pricing in 2026

Deployment type

Billing

Latency SLA

Built for

 

Standard

Per token

None

Development, testing, variable traffic

Priority processing

Per token at priority rate

Defined target per model

Latency-sensitive work without a commitment

Provisioned

Per PTU per hour, or via reservation

Defined target per model

Mission-critical, high-scale, predictable load

Batch

Per token at a discounted rate

None

Bulk asynchronous processing

What Azure actually meters

Three of the four deployment types meter tokens in the familiar way, separated only by rate: standard, a priority tier that costs more and carries a latency target, and batch that costs less and returns results asynchronously.

The fourth meters capacity. A Provisioned Throughput Unit is a slice of dedicated model processing throughput held exclusively for your deployment. Microsoft is explicit that a provisioned deployment holds that capacity whether or not requests are being made, and bills at an hourly rate per PTU deployed regardless of the number of tokens consumed. The meter starts when the deployment is created and stops when it is deleted.

Each model publishes its own PTU-to-tokens-per-minute ratio and its own minimum PTU count, so the smallest viable deployment differs by model. PTU quota is a separate concept from capacity: quota is a policy ceiling on how many PTUs you may deploy per subscription, per region and per deployment type, and it carries no cost of its own. Having quota does not guarantee capacity is available in the region.

How reservations work

Provisioned deployments bill two ways. Hourly is the flexible mode, useful for benchmarking a model or scaling up for a short event. Microsoft's own documentation advises against treating hourly as a scaling strategy, for two stated reasons: capacity may not be available when you try to scale back up, and sustained hourly billing at high utilisation typically costs more than a reservation.

Reservations commit to one month or one year and bill matching usage at a discounted rate instead of the hourly one. Three constraints shape how they are bought:

  • Reservations are not interchangeable across deployment types. Global, Data Zone and Regional each need their own, and a Global reservation will not cover Data Zone usage.

  • Global reservations are not region-specific. One Global reservation can cover deployments across several regions provided you reserved enough units.

  • Exchanges reset the term. You can swap region, deployment type, term or payment option, but the clock restarts.

Payment runs upfront or monthly.

What happens when you hit the limit

Nothing runs out, because nothing is allocated. Standard and batch deployments bill every token at list rate with no included allowance to exceed. Provisioned deployments have the opposite problem: the capacity is fixed, so exceeding it means requests queue or fail rather than costing more.

The real constraint is quota. If you need more PTUs than your subscription allows in a region, you request a quota increase, and separately confirm the region actually has capacity to give.

How Azure OpenAI pricing has changed

Date

Milestone

Source

 

18 Nov 2025

Azure AI Foundry becomes Microsoft Foundry at Ignite. Docs move to /azure/foundry/ and rename PTUs to Foundry Provisioned Throughput

Vendor

14 Aug 2024

Provisioned Reservations added in one-month and one-year terms, quoted at up to 82% and 85% below the hourly rate

Vendor

14 Aug 2024

Self-service provisioned deployments introduced at a flat $2 per PTU-hour, replacing sales-negotiated capacity

Vendor

Flexprice’s Take

Azure sells capacity reservation properly, and it is the only vendor here doing so.

Reserved throughput answers a real problem: a workload that can't tolerate variable latency. Pricing it per hour with one-month and one-year commitment tiers is how compute has been sold for two decades.

Microsoft also documents the traps. Its own docs tell you not to scale provisioned deployments with traffic, because capacity may not be there when you scale back up.

Four billing modes for one model still shifts real work onto you. Getting it wrong costs in both directions: over-provision and you pay hourly for idle PTUs, under-provision and requests queue or fail.

Quota isn't capacity, and reservations don't cross deployment types: Global, Data Zone and Regional each need their own.

The missing rate changelog is the weakest part. Structural changes get dated blog posts, but per-model rates move without announcement on a service this widely resold, and that's a problem if your own pricing sits on top.

Best For

Steady high-volume production traffic with a hard latency requirement.

Watch Out For

Reserving the wrong deployment type, since Global, Data Zone and Regional do not cover each other.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model capacity & need to track cost and margin per customer?

Flexprice reports it down to the model.

Flexprice’s Take

Azure sells capacity reservation properly, and it is the only vendor here doing so.

Reserved throughput answers a real problem: a workload that can't tolerate variable latency. Pricing it per hour with one-month and one-year commitment tiers is how compute has been sold for two decades.

Microsoft also documents the traps. Its own docs tell you not to scale provisioned deployments with traffic, because capacity may not be there when you scale back up.

Four billing modes for one model still shifts real work onto you. Getting it wrong costs in both directions: over-provision and you pay hourly for idle PTUs, under-provision and requests queue or fail.

Quota isn't capacity, and reservations don't cross deployment types: Global, Data Zone and Regional each need their own.

The missing rate changelog is the weakest part. Structural changes get dated blog posts, but per-model rates move without announcement on a service this widely resold, and that's a problem if your own pricing sits on top.

Best For

Steady high-volume production traffic with a hard latency requirement.

Watch Out For

Reserving the wrong deployment type, since Global, Data Zone and Regional do not cover each other.

Manish Choudhary

CEO & Co-founder, Flexprice

Reselling model capacity & need to track cost and margin per customer?

Flexprice reports it down to the model.

Customer
Sentiment Highlights

The cheapest option is one of the GPT models, at a minimum of $10k/month

juliangoldsmith, on provisioned throughput units, Hacker News, September 2025

"it was just a farce to get you using provisioned throughput."

7thpower, Hacker News, April 2026

Frequently Asked Questions

Frequently Asked Questions

How much does Azure OpenAI cost?

What is a PTU in Azure OpenAI?

Do Azure OpenAI reservations cover every deployment type?

Is provisioned throughput cheaper than pay-as-you-go?

Launch usage-based billing this week, not next quarter

Launch usage-based billing this week, not next quarter

Get Instant Feedback on Your Pricing | Join the Flexprice Community with 400+ Builders on Slack

Join the Flexprice Community on Slack