---
format: "aidr-story-markdown/v1"
id: "ae1e175bda7e0d6c70af75e8e68e74fb60814a755e5f872b2ae5b3b866a9dccd"
canonical_url: "https://aidr.today/ae1e175b?lang=en"
title: "Aleph Alpha's Kolibri AI Uses 21% German Data for EU Compliance"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-03T21:09:51.000Z"
category: "Models"
topics: ["regulation"]
source_urls: ["https://marketbrief.now/ai/aleph-alphas-kolibri-ai-uses-21percent-german-data-for-eu-compliance-9910e6ea","https://x.com/Aleph__Alpha/status/2106306840657297814","https://x.com/testingcatalog/status/2106453084020920491","https://x.com/emollick/status/2106487816460910912","https://x.com/ClementDelangue/status/2106348150684303422","https://x.com/aidangomez/status/2106401716417568841","https://x.com/cohere/status/2106454131023995094","https://huggingnews.com/ai/aleph-alphas-kolibri-ai-uses-21percent-german-data-for-eu-compliance-9910e6ea"]
summary: "The German AI research lab designed its latest model to adhere to the General-Purpose AI Code of Practice and the GDPR from its inception. This regulatory alignment includes a specific focus on copyright law to develop what the company calls trustworthy technology. To enhance regional utility, Aleph Alpha utilized a bilingual German-English tokenizer and ensured organic German content comprised 21.3% of the system's pre-training tokens. While the release of Kolibri has been hailed as a win for sovereign European AI, researcher Ethan Mollick noted the model was fine-tuned on synthetic data from Chinese models GLM and Qwen. The system uses a mixture-of-experts architecture with 78B total parameters, 3.46B of which are active, and supports a 1M token context window. It is distributed under an Apache 2.0 license, allowing for local hosting and private infrastructure use."
---

# Aleph Alpha's Kolibri AI Uses 21% German Data for EU Compliance

> [Open the canonical story](<https://aidr.today/ae1e175b?lang=en>)

**Published:** 2026-10-03T21:09:51.000Z
**Category:** Models
**Topics:** regulation

## Summary

The German AI research lab designed its latest model to adhere to the General\-Purpose AI Code of Practice and the GDPR from its inception\. This regulatory alignment includes a specific focus on copyright law to develop what the company calls trustworthy technology\. To enhance regional utility, Aleph Alpha utilized a bilingual German\-English tokenizer and ensured organic German content comprised 21\.3% of the system's pre\-training tokens\. While the release of Kolibri has been hailed as a win for sovereign European AI, researcher Ethan Mollick noted the model was fine\-tuned on synthetic data from Chinese models GLM and Qwen\. The system uses a mixture\-of\-experts architecture with 78B total parameters, 3\.46B of which are active, and supports a 1M token context window\. It is distributed under an Apache 2\.0 license, allowing for local hosting and private infrastructure use\.

## Sources

- [Story source](<https://marketbrief.now/ai/aleph-alphas-kolibri-ai-uses-21percent-german-data-for-eu-compliance-9910e6ea>)
- [Story source](<https://x.com/Aleph__Alpha/status/2106306840657297814>)
- [Supporting source](<https://x.com/testingcatalog/status/2106453084020920491>)
- [Supporting source](<https://x.com/emollick/status/2106487816460910912>)
- [Supporting source](<https://x.com/ClementDelangue/status/2106348150684303422>)
- [Supporting source](<https://x.com/aidangomez/status/2106401716417568841>)
- [Supporting source](<https://x.com/cohere/status/2106454131023995094>)
- [Story source](<https://huggingnews.com/ai/aleph-alphas-kolibri-ai-uses-21percent-german-data-for-eu-compliance-9910e6ea>)

