- What happened: On October 3, Aleph Alpha released the trained weights of Kolibri, an AI model specialized for German and English, under the Apache 2.0 license.
- Why it matters: Training data and efficiency techniques tuned to each language give companies and public agencies another option for handling internal documents on infrastructure they control themselves.
- What to watch next: How often the model answers incorrectly on real German and English documents, how long responses take, and how much GPU memory it requires will shape adoption decisions.
On October 3, 2026, German AI company Aleph Alpha released the trained weights of Kolibri, a large language model specialized for German and English. It has roughly 78.1 billion parameters in total, but only about 3.46 billion are actually used to process each token. Under the Apache 2.0 license, it can be deployed on infrastructure that organizations manage themselves.
For companies and public agencies, it offers a new way to process internal documents without sending them to an external AI inference service. Beyond a context length of up to about 1 million tokens, what stands out is that the model is designed with business use in mind, from the stage of collecting German training data to training the model to refrain from answering questions that lack supporting evidence. This is where the company's idea of "AI sovereignty" starts to take concrete shape.
78.1 billion parameters, but only about 3.46 billion used each time
Kolibri is a Mixture-of-Experts (MoE) model, which switches the parts it uses depending on the input. In each layer, it selects 6 of 384 expert networks and combines them with a shared network that is always active.
As a result, the parameters actually active per token come to about 3.46 billion, roughly 4.4% of the total. The design keeps the amount of computation down by not calculating everything every time, despite the large parameter count.
However, even if few parameters are used in computation, memory is still needed to store the weights of the entire model. According to the official model card, the weights alone take up about 78GB in FP8 format.
As examples of minimum configurations, the card lists two NVIDIA H100 SXM5 GPUs, or one H200 or B200. Processing long documents or handling requests from multiple users at once also requires memory to hold intermediate computation results.
Therefore, looking only at the figure of "about 3.46 billion parameters in actual use" and assuming the model runs easily on small GPUs would lead to underestimating the necessary equipment.
The model's structure also limits the computational load of long-text processing. Of its 50 layers, 40 look only at the nearest 512 tokens, while the remaining 10 look at the entire context.
Compared with a design in which every layer references the full text, the layers handling local processing can more easily limit growth in computation and memory consumption as documents get longer.
The maximum context length is 1,048,576 tokens, but the context length used in the final stage of training was 262,144 tokens.
The model card says quality and operational efficiency were verified with lengths extended beyond that, but recommends 262,144 tokens or fewer for uses that prioritize response speed or throughput, and for complex tasks.
In other words, about 1 million tokens is the technical upper limit, while about 260,000 tokens is the length to consider first in actual operation.
Users can also choose how much computation is spent before answering, from four levels: none, low, medium and high. This lets them adjust the amount of computation between simple tasks such as extracting items from a document and complex ones such as cross-checking multiple sources.
Japanese is not among the main target languages, so the model's suitability for Japanese-language work cannot be judged directly from its German and English performance.
Keeping German administrative documents in the training data
Aleph Alpha used 20 trillion tokens to pretrain Kolibri, allocating more than 20% of them to German.
The company says that simply translating English text and adding it was not enough to improve German performance. In its blog post on building German data, it points out that even if translated text is grammatically correct, biases in the geography, population and institutions it covers carry over from the original English web data.
To retain text originally written in German, post-collection filtering also has to be adapted to the language.
For example, applying an English-oriented criterion that excludes text containing extremely long words could mistakenly exclude German administrative documents, which are full of compound words.
The company adjusted such criteria for German and collected about 1.3 trillion unique German tokens from Common Crawl.
It also rewrote original German text into encyclopedia-style explanations and question-and-answer formats, generating about 1 trillion tokens of synthetic data.
Rather than adding new knowledge, this approach has the model learn the same information in different expressions. The aim is to supplement the limited volume of German data while preserving the original regional and institutional context.
The mechanism for splitting text into tokens is also designed with German in mind.
The tokenizer "UniBPE," with a vocabulary of 128,000, is designed to split German compound words along meaningful units. According to Figure 5 of the technical report published October 3, when processing the same German web data, FineWeb-2, the average number of UTF-8 bytes per token was 4.90 for Kolibri and 4.35 for the GPT-5 tokenizer.
Calculated from the averages Aleph Alpha published, Kolibri needs about 11.2% fewer tokens than the GPT-5 tokenizer to represent the same German web corpus.
For text of the same byte size, the number of tokens required is inversely proportional to the average bytes per token, so the reduction is calculated as (1 − 4.35 ÷ 4.90) × 100.
This is, however, a result calculated from published averages. It does not mean the token count falls by 11.2% for every individual document, nor that response time or operating costs fall by the same proportion.
It is also worth noting that standard BPE, using the same training data and vocabulary size, achieved an average of 4.89 bytes, nearly identical to UniBPE.
The technical report likewise explains that compression efficiency itself is on par with standard BPE. The gap with GPT-5 therefore cannot be attributed entirely to UniBPE's own algorithm.
The effect of building a tokenizer on data with enough German and the effect of splitting words with attention to their structure need to be evaluated separately.
Trained not to force an answer when there is no evidence
Kolibri is intended for use in RAG, where answers are based on documents retrieved by search, and in work that calls external tools.
When looking up administrative procedures or industrial product specifications, what matters more than fluent prose is whether the answer is based on the material provided.
Even if a question itself is plausible, if the answer is not in the material, the model is expected to refrain from answering rather than guess.
To teach this behavior, Aleph Alpha used a mechanism it calls "Merlin-Arthur."
The model being trained is Arthur, and Merlin creates documents containing the information needed to answer a question. Morgana, meanwhile, creates documents from which the evidence needed for that answer has been removed.
Arthur is not told which document it has been given. It is then trained to answer when evidence is present and to refrain when it is not.
For example, for a task of answering the conditions of use for a specific part from product documentation, one document would state those conditions and another would have just the relevant passage removed.
Even if the model happens to produce a correct answer for the latter using knowledge or guesswork alone, it is not treated as correct. What matters is not only the answer itself but the ability to judge whether it can be answered from this material.
In addition to conventional training that rewards correct answers, it can be seen as a mechanism for teaching the model to distinguish situations where it should answer from those where it should not.
In the company's evaluation, faithfulness to evidence and the ability to refrain from answering improved compared with Kolibri Origin, an earlier non-public model.
However, it cannot prevent wrong answers to every question, and its standing against competing models varies by benchmark.
For real deployments, companies and public agencies should measure both the rate at which the model answers without evidence and the rate at which it refrains even though it could have answered, using the documents they actually use.
"Low cost" claims depend on the measurement conditions
In the evaluations Aleph Alpha published, Kolibri outperformed Qwen3.6-35B-A3B, which has a similar number of active parameters, on math and expert-knowledge benchmarks.
On the other hand, Qwen scores higher on some items, such as tool calling and long-text processing.
| Benchmark | Kolibri | Qwen3.6-35B-A3B |
|---|---|---|
| AIME 2025 (English, math) | 96.9 | 84.6 |
| GPQA Diamond (English, expert knowledge) | 84.3 | 83.4 |
| BFCL v4 (tool calling, overall) | 61.4 | 67.2 |
| LongBench Pro (long-text processing) | 64.5 | 70.8 |
The figures come from the company's October 3 announcement; higher is better in every case.
They confirm that the model has areas of strength, but neither its German focus nor the release of its weights is grounds for judging it the highest-performing model for every use.
The company's claim of being "Pareto-optimal in quality and operating cost" also rests on specific measurement conditions.
In Appendix A of the technical report, generation speed was measured on a single node with eight NVIDIA B200 GPUs, using vLLM and handling many requests at once.
Multiple parallelization schemes were tried, the number of concurrent requests was raised until the memory holding intermediate results was nearly exhausted, and the highest generation throughput was adopted.
The metric used in place of cost was processing efficiency: how much output can be generated per second per GPU.
Because each model splits text into tokens differently, comparing only tokens per second would not be fair. The company therefore multiplies the token count by the average bytes per token, converting it into a measure closer to the amount of text actually generated.
This is useful for comparing generation efficiency in services that process large numbers of requests in parallel.
It does not, however, represent the overall cost of a business operation, including document-reading time, waiting for tool responses, GPU purchase costs, electricity and operating staff.
For the comparison models, quality was mainly measured in BF16, while generation speed was measured in FP8. This also carries the assumption that quality is sufficiently maintained after quantization to FP8.
What the company is claiming for Kolibri, then, is that among the models compared it achieves a high level of both quality and throughput under heavy generation loads.
For uses where a small number of users process long documents one at a time, the wait from inputting a document to receiving the first response may matter more than maximum throughput.
Developed in Europe, but the supply chain is not entirely European
Kolibri was developed by a team in Germany and trained on computing infrastructure in Germany and Finland. Pretraining used 768 NVIDIA B200 GPUs.
The technical report also names GLM-5.2, GLM-5.3 and Qwen3.8-27B as the main teacher models used to generate synthetic data for additional training.
In other words, having the ability to design and train a model in Europe is not the same as sourcing all hardware and the models used in training from European makers.
Aleph Alpha acknowledges that teacher models developed in China may carry biases on politically sensitive topics.
It then added training data aligned with European democratic values and inspected the answers the teacher models generated.
For data including free-form text, it also discloses a step in which entire conversations were classified with gpt-oss-120b and those containing answers aligned with the Chinese Communist Party's positions were excluded.
The idea is to use internationally developed models while retaining control over which outputs are used to train one's own model. However, it has not been demonstrated that this process completely removes biases inherited from the teacher models.
The company is similarly explicit about how it manages its training data.
According to the October 5 announcement, it checked training data against a blocklist containing more than 4.5 million URLs, and for third-party datasets verified licenses, provenance and handling of opt-out declarations.
It has also signed the EU's General-Purpose AI Code of Practice, but the European Commission describes the code as a voluntary tool for complying with AI Act obligations.
Signing the code or releasing trained weights does not guarantee that every way of using the model at a deploying organization is legally unproblematic.
Another feature is a setup that allows continuous management of the training process.
Using a system it calls the "Model Factory," the company managed everything from data processing to training and evaluation as a single pipeline, and continuously checked models mid-training.
It says it automatically recovered from 38 unexpected interruptions during the 21 days of pretraining.
The internal model Kolibri Origin finished pretraining on June 11, and Kolibri completed pretraining on September 11. That interval suggests a development setup able to improve a model in a short time and run large-scale training again.
However, the fact that Aleph Alpha can manage its internal training process itself is a separate matter from all training data and training code being public so that third parties can fully reproduce the same model.
The company itself is also changing.
As of October 5, Aleph Alpha has announced an agreement toward integration with Cohere. The deal awaits regulatory approval, and the two companies are to operate independently until it closes.
In assessing European "AI sovereignty," what matters is not only where models and servers are located but also who can decide future improvement policies and the terms of provision.
For companies and public agencies adopting Kolibri, the final basis for judgment will be the error rate on the German and English documents they actually use, the rate at which the model declines to answer, the GPU memory required, and the actual response time and throughput.
Only once they can measure these themselves, and choose where to run the model and how to update it, can the published weights become AI they control in their day-to-day work.
