Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
Oct 4, 2026, 12:01 AM · MarkTechPost

Aleph Alpha’s Kolibri is a German-built MoE that keeps most of its 78.1B parameters asleep — and makes the EU sovereignty pitch concrete with Apache 2.0 weights that fit one H200.
Why it matters
MarkTechPost’s Asif Razzaq reports that Aleph Alpha released Kolibri, a bilingual English-German Mixture-of-Experts model with 78.1 billion total parameters that activates only 3.46 billion — about 4.4% — per token. Context stretches to 1,048,576 tokens, reasoning effort is set per request, and the FP8 checkpoint is about 78GB on a single B200, B300 or H200 (or two H100 SXM5s) via vLLM.
The company framed the release for sovereign deployment in regulated sectors: public administration, industry, aerospace. Training ran in Germany and Finland, the pipeline redacts personal data, and Aleph Alpha says the design targets the EU General-Purpose AI Code of Practice, the AI Act and GDPR.
That combination — open weights under Apache 2.0, EU-local training, and a deployable single-GPU footprint — is rarer than another English-first frontier chat model. For buyers who cannot send German government or industrial text to a US API, this is a product argument, not a press slogan.
From the desk
We’re watching European open-weight releases for whether they stay demos or become the default stack inside regulated walls. Kolibri looks closer to the latter. Sparse MoE with hybrid attention (sliding-window on most layers, full attention every fifth block) is how you sell a million-token window without pretending every layer grows the KV cache forever. Aleph Alpha also claims a German-tuned UniBPE tokenizer that uses fewer tokens on German web text than GPT-5’s — the kind of boring efficiency that compounds in production.
Benchmarks in the MarkTechPost write-up put Kolibri ahead on GPQA Diamond and AIME 2025/2026 in English, tied with Qwen3.5 35B-A3B on an English agentic average, and trailing on BFCL tool calling. Dense Qwen3.8 27B still scores higher overall while activating far more parameters per token. So the honest read is not “Europe beats everyone.” It’s “Europe ships a competitive sparse bilingual model you can actually host.”
We’re for that. Useful AI in public administration should not require routing every document through a hyperscaler chat endpoint. The downside if this scales badly is familiar: open weights travel, and a strong German-English MoE will get fine-tuned for spam, fraud and influence ops the same way every other strong open model does. Sovereignty cuts both ways — local control and local misuse.
I’m also watching the Merlin-Arthur abstention training mentioned in the report. Teaching a model to refuse when retrieved context doesn’t support an answer is the right instinct for government use. Whether that holds once buyers bolt on tools and agent loops is the real test.
Context
Aleph Alpha has long positioned itself as the European alternative to US labs. Kolibri-1 sits against peers like Qwen MoEs, NVIDIA Nemotron 3 Super and Mistral Small 4 in the MarkTechPost comparison table — different licenses, contexts and active-parameter budgets, same fight for deployable open weights.
Who feels it
- Enterprises & public sector in the EU
- A path to host bilingual models under GDPR-shaped pipelines without sending core text to a US SaaS endpoint — if ops teams can actually run vLLM on H200-class iron.
- Open-weight developers
- Apache 2.0 FP8 weights plus dedicated reasoning/tool parsers make Kolibri a forkable base, not a gated API teaser.
- US hyperscaler APIs
- Not an existential threat tomorrow, but another data point that regulated buyers will keep shopping for local alternatives.
What to watch
- Whether BFCL and real tool-calling evals close the gap with Qwen-class MoEs
- First production deployments in German public administration or aerospace
- How quickly community fine-tunes and agent harnesses appear on Hugging Face
Companies: NVIDIA