Reasoning Models vs Standard LLMs: When Thinking Actually Helps

When to use a reasoning model vs a standard one: tasks where thinking helps, where it hurts, and how to set effort on OpenAI, Anthropic, Gemini, and DeepSeek.

A reasoning model is the same kind of language model as a standard one, with one added behaviour: it spends tokens “thinking” before it answers. That sounds free. It is not. Reasoning tokens are billed as output tokens, they eat latency, and they help a lot on some tasks and not at all on others. The right question is not “is the reasoning model smarter” - it is “does this task reward extra thinking, and am I willing to pay for it.” This guide walks through what each major provider means by “thinking,” where benchmarks show it helps, where it hurts, and how to control the effort on the four major APIs.

TL;DR

  • Reasoning is a behaviour, not a species. OpenAI, Anthropic, Google, and DeepSeek all expose it as a toggle or budget on top of the same kind of transformer. The OpenAI reasoning guide says reasoning models “introduce reasoning tokens in addition to input and output tokens,” and “reasoning tokens are not visible via the API [but] still occupy space in the model’s context window and are billed as output tokens.”
  • It helps when the task has a right answer to find. Math, formal code planning, multi-step agentic loops, scientific QA - the DeepSeek-R1 card reports AIME 2024 at 79.8 percent and MATH-500 at 97.3, with 32,768 tokens of generation headroom.
  • It hurts when the task is open-ended or latency-bound. Casual chat, rewriting, classification, voice, and short retrieval answers do not benefit. Anthropic’s guidance is to “match the starting point to the task” and to start “near the 1,024-token minimum” for simple prompts.
  • Every provider gives you a knob. OpenAI exposes reasoning.effort (none, minimal, low, medium, high, xhigh, max). Anthropic exposes manual budget_tokens on Claude 4.5 and earlier, and output_config.effort on adaptive-thinking models from Claude 4.6 onward. Google exposes thinking_level (minimal, low, medium, high) plus a dynamic-thinking default. DeepSeek exposes sampling parameters (temperature 0.5 to 0.7, top-p 0.95) and an optional reasoning prefix.
  • The single biggest mistake is leaving the default on. If you do nothing, you get medium effort on most providers, which adds tokens, latency, and dollars to tasks that did not need them.

What “reasoning” actually means across vendors

Every major provider now ships the same idea under different names. The mechanic is the same: the model generates internal tokens before its visible answer, and those internal tokens are charged as output.

OpenAI’s reasoning guide is explicit about the mechanic: reasoning models “use internal reasoning tokens before producing a response,” which helps them “plan, use tools effectively, inspect alternatives, recover from ambiguity, and solve harder multi-step tasks.” The current lineup in the guide includes gpt-6-astra, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5 (default medium effort), and older gpt-5 and gpt-5.4. Reasoning tokens “are not visible via the API” but “still occupy space in the model’s context window and are billed as output tokens.”

Anthropic calls it extended thinking. On Claude 4.5 and earlier you set thinking: {type: "enabled", budget_tokens: N} directly. The Anthropic page calls out a hard floor: “Minimum of 1,024 tokens. The API rejects smaller values.” The budget is a target, not a cap - “Claude may stop reasoning well before the budget is exhausted; max_tokens remains the hard ceiling on total output.” On Claude 4.6 and later the manual budget is deprecated in favour of thinking: {type: "adaptive"} plus output_config: {effort: ...}. Claude 4.7 and Claude Opus 5.5 reject manual thinking entirely: “requests using it…return a 400 error.” If you see that error today, your model uses adaptive thinking and you switch by removing budget_tokens.

Google’s Gemini thinking docs call the same behaviour “thinking” and expose it through thinking_level: "low" | "medium" | "high" (with "minimal" on some Flash models). Thinking is “on by default” for most Gemini 2.5 and 3.x models; gemini-2.5-flash-lite is the exception and ships with thinking off. The guide is firm about one trap: “max_output_tokens…sets the maximum number of tokens a response can generate, including thought tokens,” so a low cap silently cuts the thinking budget. “If the model hits this limit while reasoning, it stops generating with status incomplete…To reduce cost or latency without truncating responses, lower thinking_level (low or medium) instead of setting a small max_output_tokens.”

DeepSeek exposes the same mechanic through open weights. The DeepSeek-R1 card describes a 671B-parameter Mixture-of-Experts model with 37B activated, 128K context, MIT licence, and a recommended temperature of 0.6 (range 0.5 to 0.7). The card recommends prepending <think>\n to enforce the reasoning trace: “the model engages in thorough reasoning” when forced to begin with that prefix, and tends to “bypass thinking pattern” without it.

The open-weight world ships the same idea with a different dial. The gpt-oss-120b model card sets reasoning effort through the system prompt: “Reasoning: low” for “fast responses for general dialogue,” “Reasoning: medium” for “balanced speed and detail,” “Reasoning: high” for “deep and detailed analysis.” It is a 117B MoE with 5.1B active, MXFP4 quantisation so it runs on a single 80 GB GPU, and an Apache 2.0 licence.

The single sentence that ties this all together: reasoning is the same behaviour with five different control surfaces, and four of them let you turn it down.

Where reasoning helps

The DeepSeek-R1 card reports the strongest benchmark numbers for the open-weight reasoning models. The headline row, against OpenAI o1-1217 with a 32,768-token generation cap and 64 responses per query for pass@1:

BenchmarkDeepSeek-R1OpenAI o1-1217
MMLU (Pass@1)90.891.8
GPQA Diamond71.575.7
LiveCodeBench65.963.4
SWE-bench Verified49.248.9
AIME 202479.879.2
MATH-50097.396.4
Codeforces rating20292061

The DeepSeek-R1 distilled checkpoints drop a lot of that performance but keep the same shape: the 32B Qwen distill still scores 72.6 on AIME 2024 and 94.3 on MATH-500. The pattern in both tables is the same - reasoning helps the most where the task has a verifiable right answer, and the gap narrows on broad-knowledge benchmarks (MMLU 90.8 vs 91.8) where the standard model can lean on memorisation.

The Anthropic guidance lands at the same conclusion from the practitioner side. For complex tasks, Anthropic recommends a 16,000-token starting budget and tuning from there: “Larger budgets can improve response quality by enabling more thorough analysis for complex problems.” For simple tasks, the recommendation is the opposite: “For simple tasks, start near the 1,024-token minimum and increase incrementally.”

Gemini’s docs are even more pointed. The recommended mapping is:

  • “Simple: minimal/low - fact retrieval or classification (e.g., ‘Where was DeepMind founded?’).”
  • “Moderate: default - comparing concepts or creative reasoning (e.g., Compare electric and hybrid cars).”
  • “Complex: maximum - advanced coding, math, or multi-step planning (e.g., Solve AIME math problems).”

OpenAI’s effort table maps the same idea onto its six-step dial. The high-effort row is “hard reasoning, complex debugging, deep planning, and high-value tasks where quality and intelligence matters more than latency.” The xhigh row extends it: “deep research, asynchronous workflows and agentic tasks that require long runs.” The none row exists for “latency-critical tasks that do not benefit from any reasoning or multi-chained tool calls” - voice, fast info retrieval, classification. gpt-6-astra does not accept none and returns HTTP 400 if you try.

The rule that falls out of all four documents: reasoning helps when the task has a verifiable answer, a multi-step plan, or a long horizon. It does not help when the answer is “rewrite this paragraph” or “what is the capital of France.”

Where reasoning hurts

Three costs that always show up, regardless of vendor:

  1. Token cost. OpenAI is explicit: “Reasoning tokens…are billed as output tokens.” On Anthropic the budget is rendered into the prompt, so changing budget_tokens invalidates prompt caches. Anthropic’s example shows the cost: switching the budget from 4,000 to 8,000 tokens on the third request “re-creates the cache (cache_creation_input_tokens=1370, cache_read_input_tokens=0).” If you cache a long system prompt for cost reasons, leaving thinking effort on a moving dial can quietly defeat the cache.

  2. Latency. Gemini’s docs warn that lowering thinking_level is how you cut latency without losing the visible answer. Anthropic’s warning is stronger: “Pushing the model to think beyond 32k tokens produces long-running requests that can hit system timeouts and open-connection limits.” For a budget above 32K, “use batch processing to avoid networking issues.” OpenAI’s recommendation is “reserve at least 25,000 tokens for reasoning and outputs when you start experimenting” - and the failure mode is that the response can return status: incomplete with reason: max_output_tokens “before any visible output tokens are produced.”

  3. Quality regressions on the wrong task. Reasoning models verbose and over-think on simple prompts. The DeepSeek-R1 card flags the inverse failure: the model “tend[s] to bypass thinking pattern…when responding to certain queries, which can adversely affect the model’s performance.” The same model can both over-think and under-think, depending on how you prompt it. Anthropic’s tuning advice is to start small and grow: “start near the 1,024-token minimum” for simple prompts, then “increase incrementally.”

A practical pattern: keep reasoning off for classification, rewriting, retrieval, voice, and short chat. Keep it on for math, multi-step agentic planning, hard debugging, and long-horizon research. For mixed workloads, use the lowest effort that solves the hardest task in the set.

How each provider lets you control reasoning effort

The four vendors expose the same dial under different names.

OpenAI uses reasoning.effort on the Responses API (recommended for reasoning models). The six-step dial is none, minimal, low, medium, high, xhigh, max. gpt-5.5 defaults to medium; gpt-6-astra does not support none; gpt-5.6 and gpt-6 also expose reasoning.mode: "standard" | "pro", where pro “performs more model work than standard mode, increasing token usage and cost.” For long conversations, reasoning.context keeps prior reasoning in scope: current_turn (the default on most models) only carries the active turn; all_turns (default on gpt-5.6) renders earlier reasoning as input on later turns.

Anthropic uses budget_tokens on Claude 4.5 and earlier, and output_config.effort on adaptive-thinking models from Claude 4.6. The minimum is 1,024 tokens, and budget_tokens must be less than max_tokens (except in interleaved-thinking mode). On Claude 4.6 and later, manual mode is deprecated; on Claude 4.7 and Claude Opus 5.5, manual mode is rejected with HTTP 400. The migration is one-line: remove budget_tokens, set thinking: {type: "adaptive"}, and control depth with output_config: {effort: "low" | "medium" | "high" | "max"}. Anthropic notes one behaviour change to expect: “With a fixed budget, Claude thinks on every request. With adaptive thinking, Claude decides whether and how much to think on each request, and at lower effort settings it may skip thinking entirely on easy inputs.” That is a feature if your workload is mixed, and a footgun if you depend on reasoning happening.

Google Gemini uses thinking_level on generation_config. The default is dynamic thinking, where “Gemini models engage in dynamic thinking by default, automatically adjusting the amount of reasoning effort based on the complexity of the request.” You override the level explicitly with "minimal" | "low" | "medium" | "high". The trap is max_output_tokens: it is a hard ceiling that includes thought tokens, so a low cap cuts reasoning silently. Gemini’s fix is to lower thinking_level instead.

DeepSeek-R1 and the open-weight world do not have a vendor API parameter. You control reasoning through sampling: temperature 0.5 to 0.7 (the card recommends 0.6), top-p 0.95, and an optional <think>\n prefix on the user message to enforce the trace. Open-weight reasoning models like gpt-oss-120b take a different approach and put reasoning effort in the system prompt: “Reasoning: low | medium | high.” The card is clear that the chain-of-thought is “fully exposed for debugging” but “not intended to be shown to end users.”

A decision framework

Five questions, in order. Stop at the first “no” or “depends, but the default works” and you have your answer.

  1. Does the task have a verifiable answer, a multi-step plan, or a long horizon? If yes, a reasoning model is on the table. If no, a standard model is cheaper and faster.
  2. Is latency or cost the binding constraint? If yes, default to the lowest effort that solves the simplest task in your workload, and turn reasoning on only for the prompts that need it. OpenAI’s none and Anthropic’s lower effort settings are the levers.
  3. Do you cache long system prompts? If yes, hold the thinking configuration stable for the life of the cache. Anthropic’s example shows a single budget change re-creating a 1,370-token cache. OpenAI’s configuration_update between responses preserves the prefix - that is the safe pattern.
  4. Is the workload mixed (simple and hard in the same batch)? If yes, prefer adaptive thinking on Anthropic, dynamic thinking on Gemini, or route per-request on OpenAI. Fixed high effort on a mixed workload overpays.
  5. Are you running open weights? If yes, set temperature to 0.6, prepend <think>\n when you need the trace, and accept that effort is a prompt-time choice, not an API flag.

The Bottom Line

Reasoning is a behaviour, and it is the right tool for a narrow band of tasks. Math, formal code planning, long-horizon agentic workflows, and scientific QA show the strongest benchmark gains - DeepSeek-R1’s 79.8 on AIME 2024, 97.3 on MATH-500, 49.2 on SWE-bench Verified are the headline numbers for the open-weight world, and OpenAI’s effort table names the same tasks at the high end of the dial. For everything else, reasoning tokens are an output-token surcharge that does not improve the answer. Use the lowest effort that solves the hardest task in your workload, hold the configuration stable if you cache, and on open weights, prompt with <think>\n and temperature 0.6 rather than reaching for a non-existent API flag.