Agent Shortlist

Reference dataset

AI Model API Pricing (2026).

Side-by-side per-token pricing for 32 frontier and value-tier AI models in 2026 — Claude, GPT-5, Gemini, Grok, DeepSeek, Llama, and more. Filter by vendor or tier, verify against the source, plug into the cost calculator. Pricing is verified daily against vendor sources, with the dataset published as an open JSON file (CC-BY-4.0) you can use in your own tools. For the full breakdown of what an AI agent actually costs to build and run, see how much does it cost to build an AI agent.

Verified: 2026-07-26·32 models·9 vendors·Open dataset on GitHub ↗
Always verify pricing directly with the vendor before committing budget. Vendors change prices, ship new tiers, and apply per-account discounts that a public dataset can't track. Each vendor block below links to its official pricing page.
On release dates: a clickable date (↗) links to the vendor announcement that displays it. A plain-text date means the vendor doesn't publish an explicit release date publicly, so we use the date the model first appeared on public API marketplaces or, where multiple credible third-party reports converge, the consensus date.

2026 AI API pricing at a glance.

Three tiers, picked by the cost-per-task math. Frontier-tier reasoning runs $5-$10 per million input tokens; the balanced tier ($1-$3 per million) handles ~90% of production agent work; the value tier ($0.10-$1 per million) covers high-volume mechanical tasks like classification and extraction.

TierUse caseLowest input costFrontier example

Frontier

Hardest reasoning, ambiguous decisions, long-context synthesis$1.25/M (Grok 4.3)$10/M (Claude Fable 5)

Balanced

Most production agent work — drafting, planning, multi-step workflows$0.5/M (Gemini 3 Flash Preview)$3/M (Claude Sonnet 4.6)

Value

Classification, extraction, routing, high-volume mechanical work$0.0938/M (DeepSeek V4 Flash)$1/M (Claude Haiku 4.5)

Lowest input cost across all tracked models in each tier as of 2026-07-26. Output tokens are 2-6× more expensive than input across most vendors. Full per-model breakdown in the filterable table below.

Vendor

Tier

Showing 32 of 32 models
ModelInputper 1M tokensOutputper 1M tokensTier
Claude Fable 5Released

Anthropic's premium creative-writing and roleplay model. Priced 2x above Opus 5. Different design target than Opus — reach for it when character consistency, prose quality, or long-form fiction matters more than raw reasoning throughput.

$10$50frontier
Claude Mythos 5

Limited-availability frontier model (Glasswing program). Same pricing as Fable 5. Not generally accessible — access via Anthropic's Glasswing program.

$10$50frontier
Claude Opus 5

Anthropic's current flagship reasoning model. Same $5/$25 pricing as the entire Opus 4.x line — no price bump for the generation upgrade. Best pick for the hardest agentic decisions where reasoning quality compounds.

$5$25frontier
Claude Opus 5 Fast

Low-latency variant of Opus 5 — 2x the price for guaranteed throughput and faster time-to-first-token. Right pick for production agentic loops where latency matters more than cost.

$10$50frontier
Claude Opus 4.8Released

Previous flagship — superseded by Opus 5 at the same $5/$25 price. Still accessible via API for backwards compatibility. No reason to pick over Opus 5 unless you've pinned to this version.

$5$25frontier
Claude Opus 4.8 FastReleased

Low-latency variant of Opus 4.8. Superseded by Opus 5 Fast at the same price. Right pick only if you've pinned to Opus 4.8 for other reasons.

$10$50frontier
Claude Opus 4.7Released

Previous flagship — superseded by Opus 4.8 (2026-05-28) at the same price. Still accessible via API for backwards compatibility.

$5$25frontier
Claude Sonnet 5

Anthropic's current balanced-tier flagship. Introductory pricing of $2 input / $10 output per million tokens is in effect through Aug 31, 2026; standard pricing of $3/$15 kicks in Sept 1, 2026. Either way, the best price-quality point in the Claude lineup for the bulk of production agent workloads.

$2$10balanced
Claude Sonnet 4.6Released

Previous balanced-tier model. Superseded by Sonnet 5 at introductory $2/$10 pricing (Sonnet 5 will match Sonnet 4.6 at $3/$15 after Aug 31, 2026). Still fully accessible via API for backwards compatibility.

$3$15balanced
Claude Haiku 4.5Released

Fast and cheap. Best for high-volume classification and short replies.

$1$5value
ModelInputper 1M tokensOutputper 1M tokensTier
GPT-5.5Released

OpenAI's frontier line. Marketed for coding and professional work.

$5$30frontier
GPT-5.4Released

The full GPT-5.4 — not the mini. Comparable pricing to Claude Sonnet with OpenAI's tool-use and computer-use stack.

$2.5$15balanced
GPT-5.4 miniReleased

OpenAI's value tier. Strong on coding, computer use, and subagent workloads.

$0.75$4.5value
GPT-5.4 nanoReleased

OpenAI's cheapest model. ~4x cheaper than GPT-5.4 mini. Best for high-volume mechanical work — classification, extraction, format conversion — where Haiku-tier output is sufficient.

$0.20$1.25value
GPT-5.3 Codex

OpenAI's coding-specific model — powers the Codex CLI and ChatGPT cloud-runner agent. Tuned for software engineering tasks; competitive with Claude Sonnet on coding-only benchmarks at lower input cost.

$1.75$14balanced
ModelInputper 1M tokensOutputper 1M tokensTier
Grok 4.5Released

xAI's flagship as of July 2026 — the model xAI now points everyone to for everything, including code. It costs more than Grok 4.3 ($2/$6 vs $1.25/$2.50) and drops to a 500k context (half of 4.3's 1M), the trade for what xAI bills as its most intelligent and fastest model. Reach for 4.3 instead when you need the 1M window or the lower price, and for cheap agentic coding, Grok Build 0.1 is still the value pick. xAI also ships separately-priced Voice and Imagine APIs — see the Voice & Media APIs section below.

$2$6frontier
Grok 4.3Released

The previous xAI flagship, still worth it: a 1M context — double Grok 4.5's 500k — at half the price ($1.25/$2.50), which makes it the better call for long-document or high-volume work where 4.5's extra reasoning isn't worth 2–3× the token cost. Live web access via X integration, agentic tool calling, non-reasoning mode by default. xAI also lists dated Grok 4.20 reasoning, non-reasoning, and multi-agent snapshots at this same price; we track the headline models here.

$1.25$2.5frontier
Grok Build 0.1Released

xAI's coding-specific model — trained for agentic coding workflows. xAI now steers general users to Grok 4.5 for code, but Build 0.1 stays the cheaper, faster option: a fraction of 4.5's output price with a 256k context (enough for most repo work). Direct competitor to GPT-5 mini and Claude Haiku at the value tier for coding tasks.

$1$2value
ModelInputper 1M tokensOutputper 1M tokensTier
Gemini 2.5 ProReleased

Largest context window on the list. Pricing shown is for prompts ≤200k tokens; rises to $2.50/$15 above that.

$1.25$10balanced
Gemini 2.5 FlashReleased

Hybrid reasoning model with 1M context and tunable thinking budgets.

$0.30$2.5value
Gemini 3.5 FlashReleased

Google's current Flash flagship — 'most intelligent model built for speed' per their docs. Combines frontier intelligence with native Google Search and Maps grounding (5k free prompts/month, then $14 per 1k queries). Context caching at $0.15/M.

$1.5$9balanced
Gemini 3.1 Pro PreviewReleased

Multimodal, agentic, strong on coding. Pricing shown is for prompts ≤200k tokens; rises to $4/$18 above that.

$2$12balanced
Gemini 3 Flash PreviewReleased

Google's high-speed thinking model for agentic workflows. Strong on coding and multi-turn chat. Top 5 on public model marketplaces this week.

$0.50$3balanced
Gemini 3.1 Flash-Lite PreviewReleased

Google's most cost-efficient model. Optimised for high-volume agentic tasks and simple data processing.

$0.25$1.5value

Meta / Together AI

Vendor pricing page ↗
ModelInputper 1M tokensOutputper 1M tokensTier
Llama 3.3 70B (Together AI)Released

Open weights. The price shown reflects Together AI; other inference providers host the same weights at different rates.

$0.88$0.88value
ModelInputper 1M tokensOutputper 1M tokensTier
DeepSeek V4 FlashReleased

DeepSeek's value tier. Cache-hit pricing drops input to $0.0028/M on repeat queries.

$0.09$0.19value
DeepSeek V3.2Released

DeepSeek's capable open-weight model. Extraordinary value — frontier-class performance at $0.25/M input. Top 3 on public model marketplaces this week.

$0.21$0.32value
ModelInputper 1M tokensOutputper 1M tokensTier
Kimi K2.7 CodeReleased

Moonshot AI's coding-specific model — trained for agentic coding workflows. Value-tier pricing with 256k context, direct competitor to Grok Build 0.1.

$0.75$3.5value
Kimi K2.6Released

Moonshot AI's general-purpose flagship. Strong on long-context and agentic work. For coding-specific tasks, K2.7 Code is the better pick.

$0.65$2.72balanced
ModelInputper 1M tokensOutputper 1M tokensTier
GLM 5.2Released

z.ai's current flagship — successor to GLM 5.1. ~50% price bump from 5.1 traded for 5x context (1M tokens) and stronger reasoning/tool-use benchmarks.

$0.68$2.1428balanced
GLM 5.1Released

Previous z.ai flagship — superseded by GLM 5.2 (2026-06-13). Still available; cheaper than 5.2 if 200k context is enough.

$0.97$3.036balanced
ModelInputper 1M tokensOutputper 1M tokensTier
Mistral Large 2.1Released

European frontier model. Stronger data residency story than US providers — useful for EU compliance-heavy workflows.

$2$6balanced

Beyond chat models

Voice & Media APIs.

Voice agents, text-to-speech, transcription, image and video generation — APIs that price on hours, characters, images, or seconds rather than tokens. We seed this section with vendors we have verified against the source. More vendors as we confirm their pricing pages.

xAI

8 products
ProductCategoryPrice

xAI Voice Agent

Real-time conversational voice agent. Bundled speech-to-text, LLM, and TTS in one priced unit.

Voice agent$3.00 / hour

xAI TTS

Standalone text-to-speech endpoint. Pricing is per character of input text, not per second of audio output.

Text-to-speech$15.00 / 1M chars

xAI STT (batch)

Asynchronous transcription — submit audio, poll for the result. Cheapest of xAI's STT options.

Speech-to-text$0.10 / hour

xAI STT (streaming)

Real-time streaming transcription. 2x the batch rate; the right pick for live captioning or voice agent backends.

Speech-to-text$0.20 / hour

xAI Imagine — Image

Image generation and editing at 1K or 2K resolution. Industry-leading speed per xAI's marketing. The standard tier — the higher-fidelity quality tier is 2.5× the price.

Image generation$0.02 / image

xAI Imagine — Image (quality)

The higher-fidelity image tier, 2.5× the standard $0.02 rate. Worth it when output quality matters more than per-image cost.

Image generation$0.05 / image

xAI Imagine — Video

Video generation at 480p or 720p. Priced per second of output video, not per second of compute. The 1.5 model costs more per second for higher quality.

Video generation$0.05 / second

xAI Imagine — Video 1.5

The newer, higher-quality video model at $0.08/second — 60% more than the base Imagine video rate. Priced per second of output video.

Video generation$0.08 / second

Spot an inaccuracy or want a vendor added? Voice and media APIs don't have a single aggregator the way chat models do, so verification is per-vendor. Open an issue on the ai-agent-pricing repo with a screenshot of the vendor pricing page and we'll add it.

Run your own numbers

Pricing is just the start. The real question is what your workflow actually costs.

The cost calculator combines this pricing data with realistic per-task token estimates across ten builder workflows — from ticket classification to code review — so you can see the monthly bill at your real volume, not just the headline rate.

Open the cost calculator →

Open dataset

Use this pricing data in your own product.

The full dataset is published as pricing.json in a public GitHub repo. CC-BY-4.0. No API key, no rate limit, no auth. Updated automatically when prices change.

curl

curl -O https://raw.githubusercontent.com/lucaspowell8020/ai-agent-pricing/main/pricing.json

JavaScript

const data = await fetch(
  "https://raw.githubusercontent.com/lucaspowell8020/ai-agent-pricing/main/pricing.json"
).then((r) => r.json());

const opus = data.models.find((m) => m.slug === "claude-opus-4-8");
console.log(opus.inputPricePerMillion, opus.outputPricePerMillion);

Python

import json, urllib.request

data = json.loads(urllib.request.urlopen(
    "https://raw.githubusercontent.com/lucaspowell8020/ai-agent-pricing/main/pricing.json"
).read())

opus = next(m for m in data["models"] if m["slug"] == "claude-opus-4-8")
print(opus["inputPricePerMillion"], opus["outputPricePerMillion"])

Common questions

What builders ask before they pick a model.

How much does the Claude API cost?

Claude API pricing depends on the model. Claude Opus 4.8 is $5 per million input tokens and $25 per million output tokens. Claude Sonnet 4.6 is $3 input and $15 output. Claude Haiku 4.5 is $1 input and $5 output. Anthropic cut Opus pricing 66% in early 2026, making frontier-tier reasoning much more accessible. Prompt caching can drop input costs another 50–90% on repeat queries.

How much does GPT-5 cost?

GPT-5 is $1.25 per million input tokens and $10 per million output tokens. GPT-5 mini is $0.25 input and $2 output — about 5× cheaper. OpenAI's frontier line was rebranded from GPT-4 to GPT-5 in 2025. Cached input pricing is roughly 50% of standard input.

What is the cheapest LLM API?

DeepSeek V4 Flash is currently the cheapest credible model at $0.14 per million input tokens and $0.28 per million output tokens — roughly 35× cheaper than Claude Sonnet 4.6 for input. Gemini 2.5 Flash and Claude Haiku 4.5 are the cheapest options from major US vendors. For open-weight models, Llama 3.3 70B via Together AI is competitive at the value tier. The right cheap model depends on your task — value-tier models handle classification and short replies well but struggle with complex reasoning.

How do you keep this pricing data accurate?

An automated audit runs daily. Small drift (under 25% on both input and output) is auto-applied and published the same day. Larger changes — re-pricings, model deprecations, tier consolidations — go through human editorial review against the live vendor pricing page before publication. Vendor pages remain the canonical source of truth; this dataset is a clean machine-readable mirror, not a substitute. The exact verification date is shown on this page and embedded in the public dataset's JSON.

Can I use this pricing data in my own product?

Yes. The full dataset is published as pricing.json in a public GitHub repo (lucaspowell8020/ai-agent-pricing) under CC-BY-4.0. You can fetch it directly from the raw GitHub URL, no API key required. Use it in commercial products, comparison tools, calculators, or blog posts — attribution to agentshortlist.com is appreciated but not required by the licence.

Why not just check the vendor pricing pages directly?

You can, and you should verify before committing budget. The reason this dataset exists: vendor pricing pages disagree on units (some quote per token, some per thousand, some per million), use struck-through old prices for SEO, and change format frequently. Aggregating into one normalised file makes side-by-side comparison and programmatic use possible. We treat the vendor pages as authoritative — this dataset is a clean, machine-readable mirror.

How does AI API pricing actually work?

Nearly every LLM API charges per token — small pieces of text the model reads (input) and writes (output). A token is roughly 3/4 of an English word. Pricing is usually quoted per million tokens, split between input and output rates. Output tokens are 2–6× more expensive than input across most vendors. A simple chat turn (1,000 input + 200 output tokens) on Claude Sonnet costs about $0.006. Subscription tiers (Claude Pro, ChatGPT Plus) bundle a fixed amount of usage with a flat monthly fee — useful for individual usage, not for production agent traffic where direct API pricing is correctly sized.

What are typical API pricing tiers?

Most frontier vendors offer three tiers: frontier (best reasoning — Claude Opus 4.8, GPT-5.5, Gemini 3 Pro), balanced (the sweet spot for most production work — Claude Sonnet 4.6, GPT-5.4, Gemini Flash), and value (cheap and fast for high-volume mechanical work — Claude Haiku 4.5, GPT-5.4 mini, Gemini Flash-Lite). Cross-tier price gaps are large: balanced is typically 5–10× cheaper than frontier; value is another 3–5× cheaper than balanced. Most teams over-spec by defaulting to frontier — picking the right tier is the single biggest cost lever in production agent workflows.

Is this an AI API pricing tracker?

Yes — this page is updated daily against vendor sources, with prices verified against an industry catalog before publication. Small drift (under 25% on input and output) is auto-applied the same day; larger changes go through human editorial review. The dataset is also published as an open JSON file you can fetch programmatically — see the 'Open dataset' section below. If you're building a comparison tool or calculator, the raw JSON is the path; if you're just checking prices for budget planning, this page is the path.

What's the difference between AI cloud pricing and direct API pricing?

Direct API pricing is what you pay the model vendor (Anthropic, OpenAI, Google) per token. Cloud pricing is what you pay a cloud provider (AWS Bedrock, Azure OpenAI, Vertex AI) for the same models, typically with a 10–30% markup in exchange for unified billing, identity management, and compliance certifications. For most builders, direct API is cheaper and simpler. For enterprises with strict procurement, identity, or compliance requirements, the cloud markup is the cost of doing business. The tier rates we publish on this page are direct API rates.

Related

Weekly digest

Get pricing changes the day they land.

Every price cut, every new model, every deprecation — surfaced in one short email. No dashboard to check, no discord to join.