Invest1 distinct publisher2 min readPublished
Hy4 preview carries 2.6 times Hy3's parameters and 3.9 times its context window. Tencent is giving the weights away. That means the build column in next quarter's model budget gets priced in accelerator memory rather than in tokens.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
Qwen 3.8 27B ships thinking at maximum, and one setting stands between you and 22,000 tokens1 distinct publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Qwen3.8's 27B dense checkpoint is the one operators can actually host1 distinct publisher
product
Ox Alpha was GLM-5.3-Flash, and the number that decides displacement is 18 billion1 distinct publisher
Sparsity is the first number I would put on the whiteboard: 49 billion active out of 770 billion total [1] is 6.4% of the weights doing work on any given token [5], against 7.1% for July's 295B/21B Hy3 [2][6], so the model got 2.6 times heavier to store [4] while the per-token compute path widened only 2.3 times [2].
The file size is where the buy-side arithmetic actually lives. Divide 1.56TB by 770 billion parameters and you get roughly 2.03 gigabytes per billion, which is two bytes a parameter, and Hy3's 598GB over 295 billion is the same 2.03 [7], so nobody has quantized anything on your behalf and the memory problem is yours to solve. At 80GB per accelerator, 1.56TB is twenty devices before the first token comes out [8]. That is the honest entry price of the build column.
Free weights still cap what a vendor can charge for text inference, though, because the buyer's alternative was never only self-hosting. Simon Willison ran his pelican-on-a-bicycle prompt through OpenRouter at the default high reasoning setting [4], which means somebody else had already bought the memory and was reselling access to it, and a reseller of downloadable weights competes on margin rather than on capability. I do not have those per-token prices in front of me. It is the first number I would go and get.
Tencent has put 1.56TB on Hugging Face [1], and that choice is its own statement: it cannot now sell access to this generation's text capability, so it is buying distribution and default-choice status with revenue it will not collect a token at a time.
The reasoning dial is the finding here, because there isn't one. The chat template accepts only "high" and "no_think" [3], so a team that wants a third of the thinking budget has no setting to ask for it, and hidden reasoning tokens are precisely the line item that breaks a monthly estimate. Willison noticed the trace writing in clipped, half-grammatical English, which he reads as perfect grammar not being worth the tokens [6]; that is a vendor compressing its own cost, and it implies the thinking runs long enough to be worth compressing.
This is probably wrong, but the direction I would bet on is that the marginal price of frontier-adjacent text tokens converges on hosting cost plus a thin margin. What would prove it wrong: hosted Hy4 prices arriving close to closed-model text pricing, which would say accelerator memory rather than model access is the genuinely scarce good.
Ranked by verification strength, evidence, and original report placement.
Tencent released Hy4 Preview, an open-weight text-input-only (no vision) LLM with 770B total parameters, 49B active parameters, a 1M-token context window, and 1.56TB of files on Hugging Face.
Tencent's previous model, Hy3, released in July, was 295B total parameters, 21B active, a 256,000-token context window, and 598GB.
Hy4's chat_template.jinja on Hugging Face defines only two reasoning_effort values, 'high' (the default) and 'no_think', raising an exception on any other value.
Willison tried his 'Generate an SVG of a pelican riding a bicycle' prompt against Hy4 with the default high reasoning via OpenRouter.
Willison observed that Hy4's reasoning trace uses slightly truncated English, which he attributes to perfect grammar not being useful or token efficient for hidden reasoning text.
Hy4's 770B total parameters are about 2.6 times Hy3's 295B.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 30, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary artifacts, single witness
Every number that matters here — 770B, 49B, 1M tokens, 1.56TB — is read off Tencent's own upload, and the reasoning-effort finding is quoted from the model's chat template rather than paraphrased, so a reader can verify it in a browser tab. The ceiling is set by who is doing the reading: one developer's notes, no second outlet, no benchmark, no independent check that the published files load and run as advertised.
Weights up, one prompt in
Public weights plus an OpenRouter endpoint good enough for a pelican test is the entire observed footprint. Nothing in this reporting shows a team serving Hy4, a throughput or latency figure, a quantized community build, or anyone who has actually budgeted the twenty-odd accelerators the file size implies — and 'Preview' in the name is itself a caution.
Specs solid, budget line inferred
On the specification side the reporting is, if anything, restrained — Willison logs a 2.6x jump as a size note and moves on. The stretch is ours: turning a file size into a claim about how next quarter's model spend gets priced requires a cost figure, a rental rate or a capacity constraint, and this story supplies none of the three.
No money in the frame
Nothing supplied touches commercial motive. Tencent uploaded weights and a developer wrote them up; we learn nothing about what the training run cost, what an open 770B release is meant to undercut, or whether anyone in the chain has a stake in how it lands. Any incentive score would be invention dressed as measurement.
Arithmetic holds, model is a preview
The multiples reproduce exactly and the two-bytes-per-parameter reading is division rather than judgement, so the quantitative spine is firm. Confidence stops in the middle band for two reasons: a preview's numbers can move before general release, and one writer's tag feed is the only witness we have to any of it.