Skip to content

LeadershipNot yet confirmed elsewhere1 publisher2 min readPublished

Tencent open-sources a 33-language translation model that fits on a device in 440 MB

Tencent has open-sourced its Hy-MT2 translation models for 33 languages, with the smallest compressed to 440 MB to run on devices. The claim that the small model beats Microsoft's API is Tencent's own, so teams paying per call have a candidate to test.

The Board Room · Leadership desk

How we use AISend a correction

Tencent says its models beat rivals; 1.8B fits in 440 MB How Tencent's translation release reaches teams weighing it, per Tencent's README. Performance claims are Tencent's own.

Per-call API users: Tencent says its 1.8B model surpasses Microsoft and Doubao APIs. Larger-model users: Tencent says 7B and 30B-A3B beat DeepSeek-V4-Pro and Kimi K2.6. On-device builders: 1.8B storage cut to 440 MB. Inference: max_context 8192, max_tokens 4096 recommended.

Tencent says its models beat rivals; 1.8B fits in 440 MB
WhoHowKindClaim
Teams paying per API callTencent says the 1.8B model surpasses mainstream commercial APIs from Microsoft and Doubao overalldecision1
Teams eyeing the larger modelsTencent says the 7B and 30B-A3B outperform DeepSeek-V4-Pro and Kimi K2.6 in fast-thinking modedecision6
On-device app buildersAngelSlim 1.25-bit quantization reduces the 1.8B model's storage requirement to 440 MB for on-device deploymentcapability4
Teams planning inferenceThe README lists max_context of 8192 and recommends max_tokens of 4096 for inferenceconstraint10

What happened

  • The family comes in three sizes, a 1.8B, a 7B and a 30B-A3B mixture-of-experts model, posted on HuggingFace and ModelScope on 21 May 2026.
  • Tencent says the 1.25-bit AngelSlim compression used for on-device deployment also makes the small model's inference 1.5 times faster.
  • Its 7B and 30B-A3B models outperform the open-source DeepSeek-V4-Pro and Kimi K2.6 in fast-thinking mode, according to Tencent.
  • Alongside the models, Tencent open-sourced IFMTBench, a benchmark for testing how well translation models follow instructions.

Why it matters

  • cost Self-hosting turns a per-call fee that scales with volume into fixed GPU and engineering costs that the team carries even in quiet months.
  • exposure The documented transformers loading path enables trust_remote_code, so code from Tencent's repository runs on the team's own machines and belongs in a security review.
  • constraint An 8,192-token context limit means long contracts and manuals must be split into chunks before they go through the model.
  • decision Picking a model size also picks which Tencent claim applies, since the Microsoft comparison covers the 1.8B and the DeepSeek comparison the two larger models.

Every quality figure in the Hy-MT2 release comes from Tencent's own testing. The README says "multi-dimensional evaluations" show strong results across general, business, domain-specific and instruction-following translation [8], and it sends readers to a technical report for the detail [12].

The size figure and the quality figure may describe different files. The 440 MB number applies to the 1.8B model after AngelSlim's 1.25-bit quantization [4]. A team that wants translation running on a device is choosing that compressed build. The README does not say whether the comparison with commercial APIs from Microsoft and Doubao was run on it or on the full model [1].

A skeptic would say a vendor's own table only shows which tests the vendor chose. For the commercial-API comparison, that holds until someone else measures. For instruction-following the objection is weaker, because Tencent published IFMTBench and outside teams can run it themselves [7]. Tencent has also put the models in front of a public competition. It is partnering with WMT26 on a video subtitle task and sponsoring awards for entrants who use Hy-MT models [9]. We'd expect the first comparisons Tencent did not run itself to come from those entries.

In our view a pilot is an engineering job before it is a procurement decision. On the transformers path, version 5.6.0 or later is required [13]. Both the vLLM and SGLang instructions build the serving software from source [11]. For llama.cpp, the quantized GGUF file depends on Tencent's STQ kernel, released at pull request #22836 [14]. Prompts need care as well: the models have no default system prompt, and language names must be written in full, in Chinese inside Chinese prompts [15].

Tencent's Hunyuan group has now shipped three translation releases in under nine months [19]. Hunyuan-MT-7B arrived on 1 September 2025 [16], HY-MT1.5 on 30 December 2025 [17] and Hy-MT2 on 21 May 2026 [3], gaps of 120 and 142 days [20]. This quarter's decision is whether to run Hy-MT2 against the API a team already pays for. If it goes into production, the team then owns a re-test with each new version, and the group has so far shipped one every four to five months [20]. We think the evidence supports a pilot on a team's own language pairs now, and supports moving production traffic only after someone other than Tencent publishes numbers.

What to watch

  • WMT26 results from entries built on Hy-MT models, which would be the first comparisons Tencent did not run itself.
  • Whether Tencent's technical report says which build of the 1.8B model was tested against the Microsoft and Doubao APIs.
  • Whether the STQ kernel that the 440 MB GGUF build depends on lands in mainline llama.cpp.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence30
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Tencent says the lightweight 1.8B model surpasses mainstream commercial APIs from providers such as Microsoft and Doubao overall.

    ReportedSupportedSource: Tencent Hy-MT2 README2 sources— create a free account to open themView cited source
  2. [2]

    Hy-MT2 is a family of "fast-thinking" multilingual translation models in three sizes, 1.8B, 7B and 30B-A3B (MoE), all supporting translation among 33 languages.

    ReportedSupportedSource: Tencent Hy-MT2 READMEView cited source
  3. [3]

    On 2026-05-21 Tencent open-sourced Hy-MT2-1.8B, Hy-MT2-7B, Hy-MT2-30B-A3B and IFMTBench on HuggingFace and ModelScope.

    ReportedSupportedSource: Tencent Hy-MT2 READMEView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. github.com

    1 article · October 11, 2026

    Tencent Open-Sources Hy-MT2

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories