Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens

Alibaba's new 27B model defaults reasoning_effort to xhigh. Simon Willison measured 22,276 reasoning tokens and 21 minutes for one SVG that took two minutes with reasoning off.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens
Photo: thenextweb.com

What happened

  • Qwen 3.8 27B is an Apache 2 licensed, 27B parameter, vision-capable LLM from Alibaba's Qwen research lab, released on the Friday before 16 August 2026.
  • Simon Willison ran the model on a 128GB M5 Max MacBook Pro and an NVIDIA DGX Spark, using LM Studio and a 17GB Q4_K_M quantized build, and also tried llama-server directly on the Spark.
  • Qwen's documentation describes the model as defaulting to xhigh for reasoning effort.
  • Qwen 3.8 comes with official support for reasoning_effort with three levels: xhigh (default) for complex tasks demanding thorough analysis, medium balancing accuracy and speed, and low for efficient reasoning optimizing for speed and cost.
  • The LM Studio GGUF build Willison tested preserves the xhigh default.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Alibaba's Qwen lab released Qwen 3.8 27B on Friday, an Apache 2 licensed, vision-capable 27B parameter model [2]. According to Simon Willison, who ran it on a 128GB M5 Max MacBook Pro and an NVIDIA DGX Spark using LM Studio's 17GB Q4_K_M build [3], the decision that will govern your latency and your token spend is not the weights: the model's `reasoning_effort` ships set to `xhigh` [4].

Qwen's own documentation lists three levels: `xhigh` as the default, for complex tasks demanding thorough analysis; `medium`, balancing accuracy and speed; and `low`, for efficient reasoning optimising speed and cost [5]. The GGUF Willison tested in LM Studio preserves that default [6]. So the out-of-the-box configuration is the one Qwen describes as being for hard problems, applied indiscriminately to every prompt you send.

The numbers are the argument. Willison's pelican-on-a-bicycle SVG prompt took 21 minutes at the default setting, spending 22,276 reasoning tokens to produce 3,223 tokens of output [7]. The same prompt with reasoning turned off produced 3,715 tokens in 137 seconds [8]. That is roughly 6.9 reasoning tokens burned per token of output [18], about nine times the wall-clock cost [19], and the cheap run emitted about 15 percent more actual output than the expensive one [20].

The second-order damage is worse than the token bill. Willison first hit LM Studio's default context limit of 8,192 tokens, because the model was consuming the entire window thinking about mundane problems; loading it with the full 262,144 token context removed the symptom [9]. His reasoning trace alone was roughly 2.7 times the size of that default window [21]. An operator who installs both defaults together gets a model that appears broken rather than one that appears slow.

Willison is not dismissive about quality. He calls the result the best pelican SVG he has generated from a model running locally, from a 17GB file on disk [10], and he lists specifics: correct frame shape, legs on both sides of the bike, wings reaching the handlebars, motion lines behind rather than in front [11]. He is also blunt that 21 minutes was not worth it [12]. The failure mode shows up most clearly on trivial input. Asked to "draw an svg of a circle", the model spent its trace deliberating over concentric guide circles, tick marks, gradient fills and Bauhaus palette options, then several minutes later produced an elaborate animated circle that was not what was asked for [13]. His recommendation is to ignore the default and start at `low` or no reasoning at all [14].

Worth noting what is still unverified. The headline gains over Qwen 3.6 27B and the closed-weight Qwen 3.7-Plus are Qwen's self-reported benchmarks, and Willison says he wants to see independent numbers [1]. Qwen 3.7-Plus was one of the lab's strongest models of any size as recently as May [15], which is the scale of claim being made.

Watch three things: whether independent benchmarks reproduce the self-reported jump [1]; whether packagers such as LM Studio keep inheriting the `xhigh` default they currently pass through [6]; and how the 27B compares on real work against the much larger Qwen 3.8 2.4T-A95B released the week before [16]. Willison also reports the model is very good at bounding boxes [17], which is the kind of capability worth testing at `low` before you pay for `xhigh`.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories