Short answer
The AI model sets how capable the AI is and what each token costs; the effort level sets how much it thinks before answering, and that hidden thinking is billed as output tokens. To avoid overpaying, start with a strong model at its default effort and lower the effort until quality on your own cases starts to drop, before switching to a cheaper model. Judge by cost per correctly handled case, not price per token.
How do the model and the effort level differ?
The model decides how capable the AI is and what each token costs. The effort level decides how much the model thinks before it answers, which decides how many tokens you pay for. Both change quality, cost and speed.
A good default: start with a strong model at its default effort, measure quality on your own cases, then lower the effort until quality starts to drop. Try that before switching to a cheaper model. Judge the result by cost per correctly handled case, not by price per token.
What does the effort setting actually control?
Current models can think before they answer: they work through the problem in text you mostly never see, then write the reply. The effort setting tells the model how much of that work to do. At low effort it answers quickly and thinks only when a problem is clearly hard. At high effort it thinks on most requests, and at greater length.
Effort is a signal, not a hard limit. Anthropic's documentation says that at lower levels Claude still thinks on sufficiently difficult problems, just less than it would at a higher level. On Claude, effort also shapes everything else the model produces: at low effort it makes fewer, more consolidated tool calls and skips the preamble; at high effort it explains its plan, calls more tools and writes fuller summaries.
One thing effort does not reliably do is shorten the visible answer. Anthropic notes that on Claude Opus 5, changing effort changes the amount of thinking, not the length of the reply. If you want shorter answers, say so in the prompt.
The same dial under three names
Each provider names and scales the setting differently, and the default is not the same everywhere. That matters when you compare prices across providers: two models at their defaults may be doing very different amounts of work.
| Anthropic Claude | OpenAI GPT | Google Gemini | |
|---|---|---|---|
| Parameter | effort | reasoning.effort | thinking_level |
| Levels | low, medium, high, xhigh, max | none, minimal, low, medium, high, xhigh, max; each model supports a subset | low, medium, high on Gemini 3 models |
| Default | high | Varies by model; medium on GPT-5.5 | medium on current Flash models; high on 3.1 Pro |
| Hidden thinking billed as | Output tokens | Output tokens | Output tokens, in full even when only a summary is shown |
| Older control it replaced | A fixed thinking-token budget | Not applicable | thinking_budget in tokens, on Gemini 2.5 |
Where does the hidden cost sit?
All three providers bill thinking at the output-token rate, and output is the expensive side. On Claude Opus 5, input costs USD 5 per million tokens and output USD 25, five times as much. So a request that thinks for a few thousand tokens before a two-line answer can cost more than the answer suggests.
The thinking is also mostly invisible. OpenAI returns only a count of reasoning tokens; Claude returns thinking blocks that are empty by default on its newest models; Google bills the full thinking even when it shows you only a summary. The only way to see what you are paying for is to read the token counts in the API's usage data.
The output limit you set covers thinking and answer together. Set it too low and a high-effort request can run out of room before it writes the answer. OpenAI recommends reserving at least 25,000 tokens when you start experimenting, and Anthropic suggests starting at 64,000 for its top effort levels.
What do the measurements show?
In the test runs Anthropic has published, how much accuracy you lose by lowering effort depends mainly on the kind of work. The pattern depends heavily on the kind of work.
Research and knowledge work was nearly flat. Medium effort matched the default's accuracy at 70% to 85% of its cost, and low effort gave up only 1 to 3 points for a third to a half off. Long-running coding was a real trade-off: on Claude Opus 5, medium gave up about 2 points for half the cost, and low about 8 points for a quarter. Deep research across many subtopics gained about 2.4 rubric points at every step up, so there was no free saving there.
One pattern beat a single setting. Running every coding task at low effort and re-running only the failures at the default passed about 93% of tasks for about USD 0.70 each. Running everything at the default passed 91.7% for USD 1.39. The trick needs a way to detect failure automatically, such as tests or a validator.
Change the model or the effort first?
Try lowering effort before switching to a cheaper model. Anthropic reports that its newest models at lower effort often match or beat previous-generation models at high effort. Keeping one model also keeps your prompt cache working, because caches do not carry over between models.
Switching model still makes sense when the task is narrow and repetitive, when latency matters more than depth, or when a cheaper tier passes your tests at its best setting. For reference, Anthropic's current prices per million tokens, input then output, are USD 10 and 50 for Claude Fable 5.1, USD 5 and 25 for Opus 5, USD 2 and 10 for Sonnet 5, and USD 1 and 5 for Haiku 4.5.
If the job is only choosing an answer from a fixed list, such as routing a ticket or classifying a document, you may not need a thinking model at all. A decision model built for that job can be far cheaper; our guide to Jev and Laya covers that option.
Which effort level suits each type of work?
Start classification and extraction at low effort, and complex analysis and coding at high. These are starting points to test, not answers. The right level is a property of your workload, and it changes when the model changes.
- Classification, extraction, routing: low effort, or a decision model
- Customer chat and quick replies: low to medium, where speed matters most
- Summaries, document analysis, drafting reports: medium, stepping up if quality drops
- Complex analysis and hard reasoning: high, the default on Claude
- Coding and agents that run for a long time: high to xhigh, with a large output limit
- Max effort: only when tests show a measurable gain over the level below; on some structured tasks it overthinks
How do you choose from your own data?
Collect 30 to 100 real cases from the work, with the correct answer written down. Run all of them at two or three effort levels on the same model, changing nothing else. Record accuracy, total tokens and time for each level.
Choose the lowest level whose accuracy you can accept, then compare cost per correct case rather than cost per request. A cheaper request that fails more often, or needs a retry, is not cheaper.
Two practical details. Changing effort in the middle of a conversation usually resets the prompt cache, so keep one level per workflow; the newest Claude models can change it per message without losing the cache. And re-run the test whenever you change model or rewrite the prompt, because the curve moves with both.
Before you set a model and effort level
- Do you have real cases with known correct answers to test against?
- Have you measured at least two effort levels on the same model?
- Are you reading token counts from the API, including hidden thinking tokens?
- Is the output limit large enough for thinking plus the answer?
- Are you comparing cost per correct case, not price per token?
- Does the task need thinking at all, or only a choice from a fixed list?
- Is effort fixed per workflow, so the prompt cache keeps working?
- Will you re-test when the model or the prompt changes?
Frequently asked questions
What is the effort level in an AI model?
Effort is a setting that tells a reasoning model how much work to put into each request, mainly how much hidden thinking to do before answering. Low effort is faster and cheaper; high effort is more thorough. Anthropic calls it effort, OpenAI calls it reasoning effort, and Google calls it thinking level.
Is higher effort always better?
No. On research and knowledge work Anthropic measured almost no gain above medium effort, while the default cost roughly 20% to 40% more. Higher effort pays off on hard coding, long agent runs and deep multi-step research. The highest level can even overthink simple structured tasks.
Do I pay for the AI's thinking?
Yes. Claude, OpenAI and Gemini all bill thinking tokens at the output rate, which is usually several times the input rate. You mostly cannot see that thinking; it shows up only as a token count in the API's usage data, so check it when estimating cost.
Does lower effort make answers shorter?
Not reliably. Effort mainly controls how much the model thinks, not how long the visible reply is. Anthropic states this explicitly for Claude Opus 5. To get shorter answers, ask for a specific length or format in the prompt.
Should I lower the effort or use a cheaper model?
Try lowering effort first. Anthropic reports that its newest models at lower effort often match older models at high effort, and staying on one model keeps your prompt cache working. Move to a cheaper model when it passes your own tests at its best setting.
What effort level should a customer-service chatbot use?
Start at low and move to medium only if answers miss things your tests catch. Customer chat is usually latency-sensitive and rarely needs deep reasoning, and Anthropic recommends low for chat and other high-volume, speed-sensitive work.