Performance gains for frontier-scale models and personal computers are now available via a new speculative decoding method. Inco AI released the tool, known as DFlash 2, to power its inference stack and deliver a processing boost of up to 4.6 times over standard decoding. The system targets both massive frontier-scale models and smaller versions running on local hardware.

The software is currently available on the vLLM project, an open source inference engine. DFlash 2 evolves a previous block diffusion method for flash speculative decoding, and a technical write-up has been made available on Papers with Code.

Sign in to suggest edits

Key sources

  1. SOURCE@lliebenwein“one of many innovations powering the inference stack at @inco_ai”x.com
  2. SUPPORT@xianbao_qian“DFlash 2 is available on @vllm_project”x.com
  3. SUPPORT@nielsrogge“I've made the write-up available on Papers with Code for anyone to learn more”x.com
  4. SOURCE@zhijianliu_“seeded at Z Lab and upgraded at Inco AI”x.com
  5. SUPPORT@jun_song“faster than Fable or Sol”x.com
  6. SOURCE@elliotarledge“Up to 4.6× the speed of autoregressive decoding, with the same output.”x.com
  7. SOURCEhuggingnewshuggingnews.com
Markdown