Updated for Claude Sonnet 5.5 · released 28 Sep 2026

Which effort level should you use — and what will it cost?

Effort is the single biggest lever on a Claude bill that most teams never touch. Pick your task below and get the recommended model and effort level, the cost per call, and the monthly number across Claude Sonnet 5.5, Opus 5.5 and GPT-6 Astra.

1. What are you doing?

This sets a starting recommendation. Everything stays editable.

2. Tune the workload

Defaults are typical for an API-backed app.


×8

Cost

Thinking tokens bill at the output rate, so they land on the output line.

$0 / call
Input
$0
Output
$0
Thinking tokens
0
Per day
$0
Per 30 days
$0
Ad slot · 728×90 · paste your AdSense unit here

Same workload, every model

Cheapest first. Sorted on your inputs, not on a benchmark score.

ModelPrice / 1MPer callPer 30 days

Monthly cost at a glance

Same numbers, drawn to scale.

What "effort" actually changes

Effort is a request parameter, not a model setting you buy. It caps how much thinking the model is allowed to do before it starts writing the answer. At low you get a single pass — fast, cheap, and fine for anything with one correct shape of answer. At max the model can chew on the problem for a long time before committing to a response.

The part that surprises people: thinking tokens are billed as output tokens. So the effort level does not touch your input cost at all — it moves the output line, and the output rate is typically 5× the input rate. That is why an unconsidered max on high-volume traffic is the fastest way to turn a $200/month bill into a $2,000/month bill.

The rule of thumb: choose the lowest effort level that still produces a correct answer, then raise it only for the specific task that failed. Effort is a per-request decision, not an account-wide one.

Where each level belongs

LevelThinkingReach for it when
low×1Classification, field extraction, formatting, translation, short replies
medium×3Chat bots, summarisation, routine analysis where latency matters
high×8Application code, refactors, debugging — the production default
xhigh×20Hard bugs, architecture, multi-file migrations, research
max×45Benchmark-grade problems. Rarely justified in production

Multipliers are planning assumptions you can edit with the slider above, not published vendor figures. Measure your own traffic, then set them to match.

Two mistakes that cost real money

  1. Running everything at max. Reasoning quality plateaus. Past a point you are buying latency and tokens, not correctness. Sample 50 real requests at high and 50 at xhigh and compare — most teams find the difference invisible outside genuinely hard tasks.
  2. Paying for reasoning on deterministic work. If the output is a label, a date or a JSON object with a fixed schema, effort cannot help you. That is a low job forever.

Frequently asked

Is a higher effort level always more accurate?

No. Accuracy improves steeply at first and then flattens, while cost keeps climbing linearly. For short, well-specified tasks the curve is essentially flat — you pay more and get the same answer.

Which model should I pair with which effort?

Prefer a stronger model at a lower effort over a weaker model at a higher effort for hard reasoning. Sonnet 5.5 at high handles most production work; reach for Opus 5.5 when the cost of a wrong answer exceeds the token bill.

Do these prices include prompt caching?

Cached input reads are billed far below the standard input rate, which is why long system prompts are worth caching. The calculator above uses standard input pricing, so treat it as a conservative upper bound.

How current are the numbers?

The data file behind this page records a verification date, shown in the footer. Pricing in this market moves monthly — always confirm against the vendor's own pricing page before you commit a budget.

Ad slot · 970×250 · in-content