The German AI research lab designed its latest model to adhere to the General-Purpose AI Code of Practice and the GDPR from its inception. This regulatory alignment includes a specific focus on copyright law to develop what the company calls trustworthy technology. To enhance regional utility, Aleph Alpha utilized a bilingual German-English tokenizer and ensured organic German content comprised 21.3% of the system's pre-training tokens.

While the release of Kolibri has been hailed as a win for sovereign European AI, researcher Ethan Mollick noted the model was fine-tuned on synthetic data from Chinese models GLM and Qwen. The system uses a mixture-of-experts architecture with 78B total parameters, 3.46B of which are active, and supports a 1M token context window. It is distributed under an Apache 2.0 license, allowing for local hosting and private infrastructure use.

Content history (1)
  • 2026-10-04 · Summary · vi · Wording fix
    Phòng thí nghiệm nghiên cứu AI Đức đã thiết kế mô hình mới nhất tuân thủ Bộ quy tắc Thực …
    Phòng thí nghiệm nghiên cứu AI Đức đã thiết kế mô hình mới nhất tuân thủ Bộ quy tắc Thực …
Sign in to suggest edits

Key sources

  1. SOURCE@aleph__alpha“78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.”x.com
  2. SUPPORT@testingcatalog“focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German”x.com
  3. SUPPORT@emollick“Europe's new “sovereign model” is fine-tuned on data generated by… GLM and Qwen”x.com
  4. SUPPORT@clementdelangue“Small bird, fast wings, Kolibri is here.”x.com
  5. SUPPORT@aidangomez“Congrats to the AlephAlpha team on a very cool German-English model”x.com
  6. SUPPORT@cohere“New Apache 2.0 model from @Aleph__Alpha. It's really good.”x.com
  7. SOURCEhuggingnewshuggingnews.com
Markdown