---
format: "aidr-story-markdown/v1"
id: "174c63be3e50aa638d7e59ef26647249086f392b498c81bcb83d048261800a00"
canonical_url: "https://aidr.today/174c63be?lang=en"
title: "Inco AI Launches DFlash 2 With 4.6x Speedup for Local AI Inference"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-19T14:58:33.000Z"
category: "Infra"
topics: ["inference","llm","open-source","vllm","dflash"]
source_urls: ["https://huggingnews.com/ai/inco-ai-launches-dflash-2-with-46x-speedup-for-local-ai-inference-edee71c5","https://x.com/LLiebenwein/status/2089854076070883426","https://x.com/Xianbao_QIAN/status/2090039971365433811","https://x.com/NielsRogge/status/2090029334455058906","https://x.com/zhijianliu_/status/2089836737132650504","https://x.com/jun_song/status/2089849870651892183","https://x.com/elliotarledge/status/2089995979462148275","https://huggingnews.com/ai/dflash-2-boosts-local-ai-inference-46x-to-hit-70-tokens-per-second-on-m5-92a312b5"]
summary: "Performance gains for frontier-scale models and personal computers are now available via a new speculative decoding method. Inco AI released the tool, known as DFlash 2, to power its inference stack and deliver a processing boost of up to 4.6 times over standard decoding. The system targets both massive frontier-scale models and smaller versions running on local hardware. The software is currently available on the vLLM project, an open source inference engine. DFlash 2 evolves a previous block diffusion method for flash speculative decoding, and a technical write-up has been made available on Papers with Code."
---

# Inco AI Launches DFlash 2 With 4\.6x Speedup for Local AI Inference

> [Open the canonical story](<https://aidr.today/174c63be?lang=en>)

**Published:** 2026-08-19T14:58:33.000Z
**Category:** Infra
**Topics:** inference, llm, open\-source, vllm, dflash

## Summary

Performance gains for frontier\-scale models and personal computers are now available via a new speculative decoding method\. Inco AI released the tool, known as DFlash 2, to power its inference stack and deliver a processing boost of up to 4\.6 times over standard decoding\. The system targets both massive frontier\-scale models and smaller versions running on local hardware\. The software is currently available on the vLLM project, an open source inference engine\. DFlash 2 evolves a previous block diffusion method for flash speculative decoding, and a technical write\-up has been made available on Papers with Code\.

## Sources

- [Story source](<https://huggingnews.com/ai/inco-ai-launches-dflash-2-with-46x-speedup-for-local-ai-inference-edee71c5>)
- [Story source](<https://x.com/LLiebenwein/status/2089854076070883426>)
- [Supporting source](<https://x.com/Xianbao_QIAN/status/2090039971365433811>)
- [Supporting source](<https://x.com/NielsRogge/status/2090029334455058906>)
- [Story source](<https://x.com/zhijianliu_/status/2089836737132650504>)
- [Supporting source](<https://x.com/jun_song/status/2089849870651892183>)
- [Story source](<https://x.com/elliotarledge/status/2089995979462148275>)
- [Story source](<https://huggingnews.com/ai/dflash-2-boosts-local-ai-inference-46x-to-hit-70-tokens-per-second-on-m5-92a312b5>)

