Ornith has released the Ornith-1.5 family of open source large language models featuring 9B Dense, 35B Mixture of Experts (MoE), and 397B MoE variants. The 397B model outperformed Claude Opus 4.8 on four reasoning and coding benchmarks, including Terminal-Bench 2.1 and SWE-bench Verified, by using a reinforcement learning loop that generates its own tasks and grading scaffolds. All variants are distributed under the MIT License for unrestricted commercial and research use.
The launch includes a 9B-Mobile version quantized to 1.5 GB for deployment on iPhone and Android devices. The 35B MoE model activates only 3B parameters per token while exceeding the performance of Google's Gemma 4-31B and Meta's Muse Glimmer-30B. The series is available in FP8, NVFP4, GGUF, and MLX formats for integration with platforms such as Ollama, LM Studio, and Unsloth.
Key sources
- SOURCE@ornith_“The model proposes new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning”x.com
- SOURCE@ornith_“despite activating only 3B parameters per token, it also outperforms dense models Gemma 4-31B and Meta's Muse Glimmer-30B”x.com
- SOURCE@ornith_“compress the model's size to 1.5 GB”x.com
- SOURCE@ornith_“You can easily run them on @ollama Ollama, @atomic_chat_hq AtomicChat, and @lmstudio LM Studio”x.com
- SUPPORT@rohanpaul_ai“Only 397B weights, but 4 wins over Claude Opus 4.8: Terminal-Bench 2.1 (86.1 v 85), SWE-bench Verified (86 v 85.8), WideSearch (80.8 v 72.9), BrowseComp (86.6 v 84.3)”x.com
- SUPPORT@testingcatalog“quantized 9B-Mobile build that targets iPhone and Android, putting a coding agent on device”x.com
- SOURCE@ornith_“developed on top of Qwen3.5 @Alibaba_Qwen with additional continued pretraining, mid-training, and post-training”x.com
- SUPPORT@vllm_project“MIT-licensed family from 9B to 397B, SOTA among open models on coding and agentic tasks”x.com