The newest open-weight model from Germany outperformed a variety of international rivals in European language evaluations. In tests conducted by developer Aleph Alpha, Kolibri scored 96.9% on the AIME 2025 math benchmark and 84.3% on GPQA Diamond, with a German overall score that narrowly surpassed Qwen3.5 35B-A3B and GPT-OSS 120B. Its performance in coding was less dominant, reaching 66.4% on SWE-Bench Verified, behind the 73.8% marked by Qwen3.6.

Kolibri uses a mixture-of-experts architecture with 78B total parameters and 3.46B active parameters, supporting a 1M token context window. It was built to comply with the EU AI Act and GDPR, utilizing a bilingual tokenizer where 21.3% of pre-training tokens are German. While released under an Apache 2.0 license for local hosting, researcher Ethan Mollick noted the system was fine-tuned on synthetic data generated by Chinese models GLM and Qwen.

Sign in to suggest edits

Key sources

  1. SOURCE@aleph__alpha“78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.”x.com
  2. SUPPORT@testingcatalog“The Aleph Alpha team developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German.”x.com
  3. SUPPORT@emollick“Europe's new “sovereign model” is fine-tuned on data generated by… GLM and Qwen”x.com
  4. SUPPORT@teortaxestex“Hybrid SWA, 3.46B active. Europe on the Pareto frontier indeed.”x.com
  5. SUPPORT@kimmonismus“In Aleph Alpha’s own evaluations, it scores 96.9% on AIME 2025 and 84.3% on GPQA Diamond. Its German overall score narrowly beats Qwen3.5 35B-A3B and GPT-OSS 120B.”x.com
  6. SOURCEmarketbrief.now
Markdown