Build2 publishers3 min readPublished
Aleph Alpha's open Kolibri model routes each token through 3.46B of its 78.1B parameters
Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.
The Engineer · Build desk

What happened
- Aleph Alpha trained Kolibri on 24 trillion tokens using 768 NVIDIA B200 GPUs located in Germany and Finland.
- According to the dev.to write-up, Kolibri's UniBPE tokenizer needs 11.2% fewer tokens for German text than GPT-5's tokenizer does.
- The Hugging Face model card lists reasoning at four levels (none, low, medium and high), tool calling, and a knowledge cutoff of June 18, 2026.
- Aleph Alpha has signed the EU General-Purpose AI Code of Practice, the voluntary framework that accompanies the AI Act.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A European team self-hosting Kolibri for sovereignty still pays for hardware sized to the full model. The lower per-token compute shows up as throughput on that hardware.
- constraint A sovereignty requirement that covers data lineage as well as hosting and jurisdiction has to be checked against the third-party models the card lists in the data pipeline.
- decision The weights are Apache 2.0 and no provider can change or switch off the model, so adopting Kolibri depends on a team's own German evals and its memory budget.
Kolibri is German for hummingbird [17]. That suits the compute side of this design and flatters the memory side. In a mixture-of-experts model, the work per token follows the active parameters, and the memory a server must hold follows the total. The dev.to write-up describes Kolibri as light on compute but carrying considerable memory [14]. Each token runs through 3.46B parameters, about 4.4% of the 78.1B that stay loaded. The resident model is roughly 22.6 times the slice used on any one token [1]. Kolibri moves Aleph Alpha from the dense models of its earlier Luminous family to this design [15].
If the weights are held at 16 bits, two bytes each, the full set is about 156.2 GB before any KV cache [2]. Quantized to 8 bits, it is about 78.1 GB [3]. The write-up does not state the released precision or a recommended serving configuration.
On the tokenizer figures, a German passage that takes 1,000 tokens under GPT-5's tokenizer takes about 888 under UniBPE [4]. On a self-hosted deployment, that means fewer decode steps per German answer and more German text per context window [4]. The saving carries over only if your German resembles the text it was measured on. Running a sample of your own documents through both tokenizers is a cheap check.
Aleph Alpha's own evaluation puts Kolibri above every comparable model of its size in German and English [13]. That is the vendor grading its own work, and the dev.to piece tells readers to treat it with caution [13]. For the ranking to apply to you, your prompts need to resemble the evaluation set. "Its size" also needs to mean the size you pay for, which could be 3.46B active or 78.1B resident [2].
The sovereignty claim is about where the model was built and which law applied. Aleph Alpha wrote in its launch post, as quoted in an analysis on tej.as, that "teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control" [8]. The post uses the word "sovereign" several times [9]. The model card covers the data side. English training text was rewritten with Google's Gemma 4 and German text with Mistral-NeMo, and Qwen3-32B labeled data for the quality filters [11]. Aleph Alpha says it also filtered out political bias it measured in Chinese open models [12].
Listing that pipeline in the model card was good practice. The card shipped with the weights and a 189-page technical report on October 3, 2026 [1]. A reviewer can check a data-lineage requirement against the models it names. The dev.to author places Kolibri's sovereignty in the training infrastructure, the governing law and control of the final model, and not in every byte of data having come from Europe [16].
What to watch
- Independent German and English benchmark results that compare Kolibri against models matched on active and on total parameter counts.
- Guidance from Aleph Alpha or others on the released weight precision and the hardware needed to serve all 78.1B parameters.
- Whether EU public-sector buyers accept a sovereignty claim for a model whose training data was rewritten and labeled by Gemma 4, Mistral-NeMo and Qwen3-32B.