Aleph Alpha releases German AI model with 78 GB of FP8 weights
Kolibri has 78.1 billion parameters, activates 3.46 billion per token and is released under Apache 2.0. Its model card calls for at least two A100 80GB GPUs.
By Ryan Merket · Published
Primary source: Tejas Kumar
Why it matters
Kolibri puts Aleph Alpha's sovereignty pitch into a downloadable model. Its roughly 78 GB FP8 weight footprint and two-A100 minimum show the hardware commitment behind running it locally.

Aleph Alpha, the German AI company co-founded by Jonas Andrulis, released Kolibri on October 3rd, according to its launch announcement, published at 9:36 a.m. UTC. The Hugging Face model card lists October 3rd as the release date and makes the weights available. In a same-day technical analysis, AI engineer Tejas Kumar also records the release date and links to the weights. Kolibri is a German-English mixture-of-experts model with 78.1 billion parameters and 3.46 billion active per token. Its model card estimates about 78 GB of FP8 weight memory and specifies a minimum of two A100 80GB GPUs.
For Andrulis, who posts as @jonasandrulis, the release puts a long-running company thesis into a model built for regulated institutions and industrial customers. Before starting Aleph Alpha with Samuel Weinbach in 2019, Andrulis founded AI company Pallas Ludens and worked on AI research and development at Apple. His professional biography describes Aleph Alpha's mission as building a sovereign stack for complex, critical environments. It records that he served as Aleph Alpha's CEO through 2025. The company's current leadership page lists Ilhan Scheer as CEO.
Aleph Alpha released Kolibri on German Reunification Day. The company pitches the model as AI built under German and European law that customers can deploy locally instead of sending sensitive documents to an external inference provider. Its launch post says Kolibri was developed in Germany and trained on infrastructure in Germany and Finland.

Sovereignty, with a German-language engineering bet
Kolibri emphasizes German data and text processing. Aleph Alpha says more than one-fifth of its 20-trillion-token pretraining run was German, or roughly 4.3 trillion tokens. Its technical report describes 24 trillion tokens across pretraining, mid-training and long-context extension. Aleph Alpha also developed a 128,000-token bilingual tokenizer intended to handle German compounds more efficiently than general-purpose tokenizers trained with less German material.
In a separate test, AI engineer Tejas Kumar ran several tokenizers over the German Basic Law. Kolibri used 35,190 tokens, compared with 41,482 for the tokenizer used by GPT-4o and GPT-5. The test covers one legal document and does not establish how Kolibri will perform across German-language work. Aleph Alpha's technical report says its tokenizer used 11.2% fewer tokens than GPT-5's on the German text it evaluated.

The model is a mixture of experts, a design that routes each token through only a subset of a larger network. Kolibri has 50 layers, each with 384 routed experts and one shared expert; six routed experts handle each token. That reduces computation per token, while the full model still has to be held in memory. Kumar's technical breakdown estimates about 78 GB for the weights in FP8, before accounting for other serving needs. The model's small active-parameter count measures compute per token; it does not mean the model will run on ordinary office hardware. Aleph Alpha's model card sets a minimum of two A100 80GB GPUs, with larger configurations recommended for serving.
The hardware needs point to likely buyers. A government department, bank or manufacturer may value keeping internal data on premises enough to provision data-center GPUs. A small team seeking a cheap local model may find the memory requirement harder to justify. Aleph Alpha says the native context window is 262,144 tokens and that it validated use up to 1,048,576 tokens, a scale aimed at large document collections rather than casual chat.
Open weights, not an entirely open pipeline
Aleph Alpha calls Kolibri sovereign because it says it controlled the development and training pipeline and customers can deploy the weights on their own infrastructure. The weights and configuration files use Apache 2.0 terms. Aleph Alpha retains rights to its training code and methods. Open weights allow customers to deploy the model with some freedom, but they do not make every part of model creation open or independently reproducible.
The training process also drew on models built elsewhere. The model card and Kumar's account describe Google Gemma 4 and Mistral-Nemo being used to rephrase web text, with Qwen3-32B used to label data for quality filtering. The use of those models is compatible with the deployment control Aleph Alpha is selling. Here, "sovereign" refers to legal jurisdiction, operational control and the ability to run the model locally, not a supply chain untouched by foreign AI.
Aleph Alpha says Kolibri outperformed every model of similar size in its English and German evaluations. Those are company-run comparisons, not independent benchmark results. Buyers will have to weigh its German-language specialization against the hardware and operational work required to serve it.
Andrulis's founding bet was that European institutions could build advanced AI without handing control of sensitive systems to a small set of foreign providers. Kolibri carries that bet forward with open weights, extensive German training data and support for local deployment. Customers will have to decide whether those features justify the hardware and operational work required to use the model.