---
format: "aidr-story-markdown/v1"
id: "1ff6c60284be5bd3a77565900cc8ea7987b60aeb6c4d6d587f7385173db3ca04"
canonical_url: "https://aidr.today/1ff6c602?lang=en"
title: "Fireworks Research Cuts Kimi K3 Reasoning Tokens by 40% in Ember-1"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-27T21:21:06.000Z"
category: "Research"
topics: ["reasoning","inference"]
source_urls: ["https://huggingnews.com/ai/fireworks-research-cuts-kimi-k3-reasoning-tokens-by-40percent-in-ember-1-bca1f42a","https://marketbrief.now/ai/fireworks-research-cuts-kimi-k3-reasoning-tokens-by-40percent-in-ember-1-bca1f42a","https://huggingnews.com/ai/update-fireworks-ai-cuts-reasoning-tokens-by-71percent-in-first-research-a958e166","https://marketbrief.now/ai/update-fireworks-ai-cuts-reasoning-tokens-by-71percent-in-first-research-a958e166","https://www.marktechpost.com/2026/09/28/fireworks-ai-releases-ember-1-a-post-trained-kimi-k3-that-uses-about-40-fewer-tokens/"]
summary: "A post-trained version of the Kimi K3 model generates shorter responses during its thinking process to lower compute requirements. Fireworks Research released the model as Ember-1, which uses roughly 40% fewer tokens in reasoning than the base model while maintaining top-tier quality. The specialized tool is 40% faster and cheaper to operate by optimizing how it consumes tokens to increase overall efficiency. Following its release, the model has trended on the Hacker News developer forum."
---

# Fireworks Research Cuts Kimi K3 Reasoning Tokens by 40% in Ember\-1

> [Open the canonical story](<https://aidr.today/1ff6c602?lang=en>)

**Published:** 2026-09-27T21:21:06.000Z
**Category:** Research
**Topics:** reasoning, inference

## Summary

A post\-trained version of the Kimi K3 model generates shorter responses during its thinking process to lower compute requirements\. Fireworks Research released the model as Ember\-1, which uses roughly 40% fewer tokens in reasoning than the base model while maintaining top\-tier quality\. The specialized tool is 40% faster and cheaper to operate by optimizing how it consumes tokens to increase overall efficiency\. Following its release, the model has trended on the Hacker News developer forum\.

## Sources

- [Story source](<https://huggingnews.com/ai/fireworks-research-cuts-kimi-k3-reasoning-tokens-by-40percent-in-ember-1-bca1f42a>)
- [Story source](<https://marketbrief.now/ai/fireworks-research-cuts-kimi-k3-reasoning-tokens-by-40percent-in-ember-1-bca1f42a>)
- [Story source](<https://huggingnews.com/ai/update-fireworks-ai-cuts-reasoning-tokens-by-71percent-in-first-research-a958e166>)
- [Story source](<https://marketbrief.now/ai/update-fireworks-ai-cuts-reasoning-tokens-by-71percent-in-first-research-a958e166>)
- [Story source](<https://www.marktechpost.com/2026/09/28/fireworks-ai-releases-ember-1-a-post-trained-kimi-k3-that-uses-about-40-fewer-tokens/>)

