BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Gemini swaps its Extended Thinking toggle for Low, Medium and High reasoning levels
Google has replaced the Gemini app's on-off Extended Thinking toggle with Low, Medium and High reasoning efforts, NPowerUser reported. Choosing a tier is now a per-task decision about how much waiting to accept in exchange for deeper reasoning.
The Engineer · Build desk

What happened
- Medium is described as the new default baseline for many tasks and is aimed at most agentic loops and complex code segments.
- Low cuts internal thinking time for quick chat, simple instruction following and fast transcript searches where latency matters most.
- High applies the model's maximum internal compute to exhaustive multi-step planning, rigorous logic analysis and dense document work.
- API developers can set the levels programmatically, the site says, under Google thinking guidelines that list token and latency trade-offs per level.
Why it matters
- decision Tasks left on the Medium default run deeper than the report assigns chat and lookup work, so Low's latency savings arrive only where a team moves a task down on purpose.
- cost On the API each level carries its own token usage and latency, so the tier setting changes resource consumption per request and belongs with other configuration under review.
- capability Agentic loops and complex code that used to force a choice between full deliberation and a fast answer now get a middle setting built for them.
NPowerUser wrote that the old toggle was "a blunt instrument" and that "either you waited for deep thought or you stuck to rapid-fire responses" [5]. Power users switched it on to make the model deliberate longer on multi-step problems, mathematics and intricate coding scripts [5]. The new tiers adjust that same internal deliberation in three steps [1]. NPowerUser is the only source for the change, and the tier descriptions come from its account.
Those descriptions are claims about typical workloads. They hold for a team only as far as its prompts match them. A transcript search that is a simple lookup fits what the report files under Low [2]. A search that has to reconcile several long documents is closer to the dense analytical exploration it reserves for High [4]. The category name on a task says less than the number of reasoning steps the task actually needs.
The report does not give per-tier latencies, token counts, or the name of the API field that sets the level [7]. Without those numbers, the only way to place a task is to run it. We'd time a representative sample of prompts at each tier, score the answers against a set already known to be correct, and keep the lowest tier that passes. In our view a team that skips this step has put its latency budget in the hands of a label.
The same report covered the model roadmap. Logan Kilpatrick confirmed that Gemini 4 Argon is on track for public release, according to NPowerUser [8]. The site also relayed a Business Insider account of an internal variant codenamed Carbon, deployed in Google's Jetski development environment, where staff feedback compares its coding performance with Anthropic's Opus models [9]. That comparison comes from Google staff using an internal tool. We'd treat it as a claim about that workload until someone outside Google runs the same tasks on both models.
What to watch
- Google publishing per-tier latency and token figures for the API levels, which would let teams place tasks from numbers instead of labels.
- Whether API requests that set no level default to Medium, as the report describes for many tasks, or to something else.
- An outside evaluation of Gemini 4 Carbon against Anthropic's Opus models on coding tasks beyond Google's Jetski environment.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Google has replaced the all-or-nothing "Extended Thinking" toggle in the Gemini app and web interface with a tiered system of Low, Medium and High reasoning efforts.
- [2]
Low cuts down internal thinking time and is aimed at quick chat interactions, simple instruction following and fast transcript searches where latency matters more than deep contemplation.
- [3]
Medium serves as "the new default baseline for many tasks", balancing intermediate reasoning against turnaround time and handling most agentic loops and complex code segments.
- [4]
High is reserved for exhaustive multi-step planning, rigorous logic analysis and dense analytical document exploration, and uses the model's maximum internal compute.
- [5]
Power users relied on extended thinking to force models to deliberate longer on multi-step problems, mathematics and intricate coding scripts; NPowerUser called it "a blunt instrument" where "either you waited for deep thought or you stuck to rapid-fire responses".
- [6]
A task left at the Medium default runs at a deeper reasoning tier than the Low tier the report assigns to quick chat and transcript search.
- [7]
Developers working with the API can programmatically manage these parameters using updated Gemini thinking API guidelines, which outline token usage and latency trade-offs for each level.
- [8]
Logan Kilpatrick confirmed that Gemini 4 "Argon" is on track for public release.
- [9]
Business Insider reported that Google staff are testing an internal variant codenamed Carbon in Google's internal development environment, Jetski, and that internal feedback has engineering teams comparing its coding performance against Anthropic's Opus class models.
Sources
1 independent publisher whose own reporting we read for this story.
- Gemini: Low, Medium & High Reasoning & Gemini 4 Argon & Carbon leaks - NPowerUser
nokiapoweruser.com
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.