Back to insights

ATI Lab insight

Anthropic Claude API Pricing: Cost per Business Task

Anthropic prices the Claude API per million tokens (MTok), billing input and output separately. On 2 October 2026 the list prices are $1 input and $5 output for...

Analysis for technology leaders and operators planning, buying, and governing AI systems.

Anthropic Claude API Pricing: Cost per Business Task

Anthropic prices the Claude API per million tokens (MTok), billing input and output separately. On 2 October 2026 the list prices are $1 input and $5 output for Claude Haiku 4.5, $2 and $10 for Sonnet 5.5, $4 and $20 for Opus 5.5, and $10 and $50 for Fable 5.1. There is no fee per request. The Batch API halves both rates, and cached input costs a tenth of the base input price or less.

For a business, that works out to fractions of a cent per unit of work. Classifying an email, extracting an invoice or drafting a support reply costs between $0.003 and $0.05 depending on the model, so 10,000 tasks a month usually lands in the tens or low hundreds of dollars in tokens. The worked examples below show the arithmetic so you can re-run it with your own volumes.

Our interest, stated up front: ATI Lab builds and runs Claude workflows for clients. Every price here comes from Anthropic's own pages, checked on 2 October 2026.

What does the Claude API cost per model?

These are the current models on Anthropic's API pricing page, in US dollars per million tokens. Anthropic's rule of thumb is about 0.75 English words per token, though newer models count more tokens for the same text (explained below).

ModelInputOutputCache write (5 min)Cache readBatch input / output
Claude Haiku 4.5$1$5$1.25$0.10$0.50 / $2.50
Claude Sonnet 5.5$2$10$2.50$0.20$1 / $5
Claude Opus 5.5$4$20$5$0.20$2 / $10
Claude Fable 5.1$10$50$12.50$0.25$5 / $25

Three details on the same page change real budgets:

  • Sonnet 5 kept its launch price. Anthropic introduced it at $2 / $10 as an offer running to 31 August 2026 and has since made that the standard price. The planned rise to $3 / $15 on 1 September will not happen.
  • Opus 5.5 is cheaper than the Opus models before it. Opus 4.5 through Opus 5 are listed at $5 / $25, so moving a workload from Opus 5 to Opus 5.5 cuts the list price by 20%.
  • Older models are retiring. Haiku 3.5, Sonnet 4, Opus 4 and Opus 4.1 are marked retired on the Claude API and remain only on some cloud platforms. If a cost estimate you inherited assumes Haiku 3.5 at $0.80 input, it no longer applies to the API.

How much does one Claude API request cost?

A request costs its input tokens times the input rate plus its output tokens times the output rate. Anthropic charges nothing per call on top. The formula is:

cost = (input tokens × input price + output tokens × output price) ÷ 1,000,000

Input includes more than the user's message. Every call carries your system prompt, any documents or retrieved passages, the conversation so far, and the definitions of any tools the model can use. Anthropic also adds a tool-use system prompt whenever tools are present (286 tokens on Sonnet 5.5 and Opus 5.5, per the pricing page). In the support example below, the fixed instructions and reference material are more than three times the size of the part that changes per ticket, which is why prompt caching usually saves more than shortening the message.

Some server-side tools add a usage charge on top of tokens. Web search costs $10 per 1,000 searches. Web fetch has no extra charge beyond the tokens of the fetched page. Claude Managed Agents, Anthropic's hosted agent runtime, adds $0.08 per session-hour while a session is running.

What do common business tasks cost on the Claude API?

The table prices three common operations tasks. The token counts are example assumptions, not measurements from a client system: about 2,300 input and 150 output tokens to classify and route one email, 4,000 and 600 to pull fields from a two-page invoice into JSON, and 6,500 and 500 to draft a support reply with policy and product context. They are stated in Haiku 4.5 tokens; for Sonnet 5.5 and Opus 5.5 we multiplied them by 1.3 to account for the newer tokenizer.

Task (monthly volume)Haiku 4.5Sonnet 5.5Opus 5.5
Classify and route an email (10,000)$0.0031 per email · $30.50 a month$0.0079 · $79.30$0.0159 · $158.60
Extract an invoice to JSON (2,000)$0.0070 per invoice · $14.00$0.0182 · $36.40$0.0364 · $72.80
Draft a support reply (3,000)$0.0090 per ticket · $27.00$0.0234 · $70.20$0.0468 · $140.40

To check one cell: an email on Haiku 4.5 costs (2,300 × $1 + 150 × $5) ÷ 1,000,000 = $0.00305.

These are list prices before caching or batching, covered below. For scanned invoices sent as images, the token count depends on image size; Anthropic's token counting endpoint gives the real number before you commit to a volume.

Why does the same text cost more tokens on newer Claude models?

Anthropic's pricing page states that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text," and that Sonnet 4.6 and earlier models use the previous one. Price comparisons that only read the per-MTok column miss this. Adjusted for a 30% token increase, the price for the same text looks like this:

ModelTokenizerList input / outputSame text, previous-tokenizer terms
Haiku 4.5Previous$1 / $5$1 / $5
Sonnet 4.6Previous$3 / $15$3 / $15
Sonnet 5.5Newer$2 / $10about $2.60 / $13
Opus 4.6Previous$5 / $25$5 / $25
Opus 5.5Newer$4 / $20about $5.20 / $26

Moving from Sonnet 4.6 to Sonnet 5.5 lowers the cost of the same work by about 13%, against the 33% the list price suggests. Moving from Opus 4.6 to Opus 5.5 leaves it roughly flat, a few percent higher. Moving from Opus 4.7, 4.8 or 5 to Opus 5.5 is a clean 20% cut, because those models already use the newer tokenizer. And Sonnet 5.5 costs about 2.6 times as much as Haiku 4.5 for the same text, against the 2× the list prices show.

Anthropic calls the 30% approximate and content-dependent, so treat the right-hand column as an estimate and count a sample of your real prompts on both models before switching a high-volume workflow.

How much do prompt caching and the Batch API save?

Prompt caching cuts the cost of the part of a prompt that repeats. The Batch API halves input and output for work that can wait. The two stack.

Caching, on the support-reply example. Suppose that on Sonnet 5.5 each ticket sends 6,500 tokens of fixed instructions, policies and product notes, 1,950 tokens of ticket text and retrieved passages, and gets back 650 tokens. Without caching, the input costs 8,450 × $2 and the output 650 × $10, per million: $0.0234 per ticket. With the fixed prefix cached and the cache warm, the prefix is read at $0.20 per MTok, so the ticket costs $0.0013 + $0.0039 + $0.0065 = $0.0117. That halves the per-ticket cost.

The 5-minute cache expires when traffic pauses. The first request after a gap pays the cache-write rate of $2.50 per MTok for the prefix, which makes that one ticket about $0.0267. If the cache goes cold ten times every working day (220 times a month), 3,000 tickets cost about $38.40 instead of $70.20. Per Anthropic, a 5-minute write pays for itself after one read and a 1-hour write after two.

Cost per support-reply draft, before and after prompt caching Sonnet 5.5, no caching $0.0234 per ticket Input $0.0169 Output $0.0065 Sonnet 5.5, cached prefix $0.0117 per ticket $0.0039 Output $0.0065 Haiku 4.5, cached prefix $0.0045 per ticket Cached input (read) Uncached input Output Example token counts, not measurements. Prices from Anthropic's pricing page, 2 Oct 2026.
Once the fixed prefix is cached, output tokens become the largest part of the bill. Shorter replies or a smaller model then save more than further prompt trimming.

Batching, on the invoice example. Invoices rarely need an answer within the minute. Sent through the Message Batches API, the Sonnet 5.5 invoice drops from $0.0182 to $0.0091, so 2,000 invoices cost $18.20 a month instead of $36.40. Anthropic's batch documentation says most batches finish within an hour and unfinished requests expire unbilled after 24 hours. It recommends the 1-hour cache for batches that share context, since a batch can outlast the 5-minute cache.

What else changes the Claude API bill?

  • US-only processing costs 10% more. On Claude 4.6 and later models, setting inference_geo to "us" applies a 1.1× multiplier to every token category. Global routing, the default, uses the list price. On Amazon Bedrock and Google Cloud, regional endpoints carry their own 10% premium over global ones, priced by the cloud provider.
  • Long prompts don't cost more per token. Claude 4.6 and later models include the full 1M-token context window at standard rates. Anthropic's example: a 900,000-token request is billed at the same per-token rate as a 9,000-token one.
  • Volume pricing is negotiated. Anthropic says volume discounts are agreed case by case through its sales team. Plan with list prices until you have a signed rate.

Is the Claude API cheaper than Claude seats for a business?

They pay for different things, so compare by the work you need done. A seat gives one person Claude's apps for varied daily work. The API runs a defined task inside your own systems, many times, without anyone typing into a chat.

For reference, Claude's pricing page lists Team standard seats at $25 per person per month, or $20 billed annually, and premium seats at $125, or $100 annually. Enterprise is $20 per seat per month billed annually plus usage at API rates. That last point matters if you are choosing Enterprise: your team's usage is priced on the same per-token table as an API build, so the arithmetic above applies to seat users as well. Our Claude Enterprise rollout guide covers the decisions to make before buying those seats.

Scale check: triaging 10,000 emails a month on Haiku 4.5 costs about $31 in tokens, a little more than one Team standard seat billed monthly, and a seat would need someone to paste in every email. Choose seats when people need help with varied, judgment-heavy work. Choose the API when the same task repeats hundreds or thousands of times and the result should land in a CRM, ERP or ticket queue. Many teams need both, and our Claude for business page explains how we set up each.

Why is the token bill only part of what an AI workflow costs?

When we review a running agent, token spend is one of three costs we measure. The other two are runtime (hosting, queues, orchestration and logging, mostly fixed) and the people who review outputs and handle exceptions. The provider invoice shows only the first, as one total.

Four of the six spend leaks listed on our AI agent cost page sit in the token line: whole documents pasted into every prompt instead of the passages needed, long sessions that keep re-sending old context, a premium model running steps a small one could do, and failed steps retrying in loops. The fixes are tighter retrieval, session summaries, routing easy steps to Haiku-class models, and a retry cap that hands the item to a person. In the support example, routing replies from Sonnet 5.5 to Haiku 4.5 would cut the token cost from $0.0117 to $0.0045 per ticket, but only if reply quality holds, so test it on a sample with human review before switching.

Whether to automate at all depends more on the hours the task takes your team today than on tokens. Our guide to calculating AI automation ROI shows the method, and the AI ROI calculator runs it on your own numbers.

Frequently asked questions

Is the Claude API free to try?

New accounts get a small amount of free credit to test the API, according to Anthropic's pricing FAQ. After that, usage is billed monthly by tokens used. Anthropic offers extended trials for enterprise evaluations through its sales team.

Does the Claude API charge per request or per token?

Per token. Each request is billed for its input tokens and output tokens at the model's rate, with no separate charge per call. Some server-side tools add their own fee, such as web search at $10 per 1,000 searches.

What is the cheapest Claude model on the API?

Claude Haiku 4.5, at $1 per million input tokens and $5 per million output tokens, or $0.50 and $2.50 through the Batch API. Haiku 3.5 was cheaper but is retired on the Claude API.

How do discounts stack on the Claude API?

Prompt caching and the Batch API combine, and both stack with the US data residency multiplier. Anthropic's pricing page states that the cache multipliers apply on top of the batch discount.

How can I estimate my own Claude API cost before building?

Take 20 to 50 real examples of the task, count their tokens with Anthropic's token counting endpoint on the model you plan to use, and multiply by monthly volume using the table above. Leave headroom for retries and tool definitions until you have logged real usage, then track cost per completed task rather than the monthly invoice total.

How this article was produced

Written by the ATI implementation team. All prices, multipliers and limits were checked on Anthropic's API pricing page, claude.com/pricing and the Message Batches documentation on 2 October 2026; recheck them before budgeting, because Anthropic has already revised Sonnet 5's scheduled pricing once this year. Token counts in the examples are illustrative assumptions, labelled as such. The cost layers and waste sources come from ATI's own agent cost reviews as published on our AI agent cost page. AI assistance was used in research and drafting; a human owns every claim and recommendation in it.

Next step

Turn the analysis into an implementation decision

Bring us the workflow, business constraint, or architecture question. We will help define the practical next step.