---
format: "aidr-story-markdown/v1"
id: "f3cde25b3582210e89fb19f8bf8ea0944492c1bf3813c47f90fdc68533f22216"
canonical_url: "https://aidr.today/f3cde25b?lang=en"
title: "Aleph Alpha Kolibri Beats Qwen With 96.9% AIME Score"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-04T09:10:03.000Z"
category: "Models"
topics: ["llm"]
source_urls: ["https://huggingnews.com/ai/aleph-alpha-kolibri-beats-qwen-with-969percent-aime-score-50cae607","https://x.com/Aleph__Alpha/status/2106306840657297814","https://x.com/testingcatalog/status/2106453084020920491","https://x.com/emollick/status/2106487816460910912","https://x.com/teortaxesTex/status/2106403475999412602","https://x.com/kimmonismus/status/2106669429320761788","https://marketbrief.now/ai/aleph-alpha-kolibri-beats-qwen-with-969percent-aime-score-50cae607"]
summary: "The newest open-weight model from Germany outperformed a variety of international rivals in European language evaluations. In tests conducted by developer Aleph Alpha, Kolibri scored 96.9% on the AIME 2025 math benchmark and 84.3% on GPQA Diamond, with a German overall score that narrowly surpassed Qwen3.5 35B-A3B and GPT-OSS 120B. Its performance in coding was less dominant, reaching 66.4% on SWE-Bench Verified, behind the 73.8% marked by Qwen3.6. Kolibri uses a mixture-of-experts architecture with 78B total parameters and 3.46B active parameters, supporting a 1M token context window. It was built to comply with the EU AI Act and GDPR, utilizing a bilingual tokenizer where 21.3% of pre-training tokens are German. While released under an Apache 2.0 license for local hosting, researcher Ethan Mollick noted the system was fine-tuned on synthetic data generated by Chinese models GLM and Qwen."
---

# Aleph Alpha Kolibri Beats Qwen With 96\.9% AIME Score

> [Open the canonical story](<https://aidr.today/f3cde25b?lang=en>)

**Published:** 2026-10-04T09:10:03.000Z
**Category:** Models
**Topics:** llm

## Summary

The newest open\-weight model from Germany outperformed a variety of international rivals in European language evaluations\. In tests conducted by developer Aleph Alpha, Kolibri scored 96\.9% on the AIME 2025 math benchmark and 84\.3% on GPQA Diamond, with a German overall score that narrowly surpassed Qwen3\.5 35B\-A3B and GPT\-OSS 120B\. Its performance in coding was less dominant, reaching 66\.4% on SWE\-Bench Verified, behind the 73\.8% marked by Qwen3\.6\. Kolibri uses a mixture\-of\-experts architecture with 78B total parameters and 3\.46B active parameters, supporting a 1M token context window\. It was built to comply with the EU AI Act and GDPR, utilizing a bilingual tokenizer where 21\.3% of pre\-training tokens are German\. While released under an Apache 2\.0 license for local hosting, researcher Ethan Mollick noted the system was fine\-tuned on synthetic data generated by Chinese models GLM and Qwen\.

## Sources

- [Story source](<https://huggingnews.com/ai/aleph-alpha-kolibri-beats-qwen-with-969percent-aime-score-50cae607>)
- [Story source](<https://x.com/Aleph__Alpha/status/2106306840657297814>)
- [Supporting source](<https://x.com/testingcatalog/status/2106453084020920491>)
- [Supporting source](<https://x.com/emollick/status/2106487816460910912>)
- [Supporting source](<https://x.com/teortaxesTex/status/2106403475999412602>)
- [Supporting source](<https://x.com/kimmonismus/status/2106669429320761788>)
- [Story source](<https://marketbrief.now/ai/aleph-alpha-kolibri-beats-qwen-with-969percent-aime-score-50cae607>)

