Build1 publisher2 min readPublished
Moonshot publishes open weights for its trillion-parameter K2.7-Code model
Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- K2.7-Code is Moonshot's third major open-source model release in six months, following K2.5 in January and K2.6 in April.
- The model is built on K2.6 and keeps its architecture, including a 384-expert mixture-of-experts design, MLA attention and the vision encoder.
- Moonshot claims K2.7-Code uses 30% fewer reasoning tokens than K2.6.
- On MLS Bench Lite the model scored 31.5% higher than K2.6, according to a dev.to post covering the release.
- Weights and code are on Hugging Face under a Modified MIT License, with vLLM and SGLang as serving options.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Self-hosting has to be provisioned for all trillion parameters, so the 32B active count lowers per-token compute without lowering the memory a team must buy.
- constraint Because preserve-thinking is forced, operators cannot discard old reasoning to free context, and long agent sessions reach the 256K ceiling sooner than code and tool output alone would.
- decision Teams budgeting for proprietary coding agents now have a self-hostable comparison point, but they can only price it once the long-horizon gains reproduce on their own repositories.
- cost With point releases on one architecture arriving every few months, any team that adopts a K2.x model takes on re-evaluation as recurring work.
Each token runs through 32 billion of the trillion parameters, about 3.2 percent of the model [4][1]. That ratio sets compute per token. Memory is sized differently. A router chooses among 384 experts for every token [5]. Any expert can be picked for the next token, so a self-hosted server keeps the full trillion in memory.
The setting with the biggest effect on agent sessions is one the operator does not get to choose. According to the post, K2.7-Code forces "preserve thinking" mode, which keeps reasoning content across multi-turn interactions [14]. In a long agent run, that reasoning stays in the window next to the code and tool output, and all of it comes out of the same 256K tokens [4]. Moonshot calls its token reduction "less overthinking" [7]. With preservation forced on, a 30% cut [6] means each turn leaves less reasoning behind in the window, if the cut holds on prompts other than Moonshot's.
The post says most benchmarks rose even with the shorter reasoning [11]. It also puts K2.7 right next to GPT-5.5 on MLS Bench Lite [10]. That benchmark tests whether AI systems can invent generalizable ML methods [9]. For a score there to predict anything about a refactor queue, the queue would have to look like ML methods research. The long-horizon result is closer to what agent buyers pay for. The post reports that end-to-end task success went up [12]. "K2.7's improvements here feel like the real gain, even if the in-house benchmarks are harder to verify independently," its author wrote [13].
The post describes a release every two months [2]. By its own dates, K2.5 and K2.6 were about three months apart [2]. Each step kept the same architecture [3]. Moonshot versions a trillion-parameter model the way a library maintainer versions a minor release. With the architecture sheet unchanged [3], that is the correct version number.
The post does not say what the modification to the MIT license covers. A hosted API is available at platform.moonshot.ai [16]. I'd start there. I'd run a frozen set of in-house tasks, log reasoning tokens per turn, and price GPUs only after the 30% figure and the long-horizon gains show up on that set.
What to watch
- Third-party per-turn reasoning-token counts for K2.7-Code against K2.6, the direct test of Moonshot's 30% claim.
- Independent reproductions of the long-horizon coding results on benchmarks Moonshot did not build.
- The text of the Modified MIT License, and whether its changes restrict commercial self-hosting.