← Back to live feed1 story
Researchers developed a technique to discard intermediate reasoning tokens that lose importance as a model continues to process, allowing existing systems to run 3x faster without training. The approach, called Prefix Sliding, avoids the memory failures associated with vanilla full attention and the loss of detail common in compaction methods.
Using reinforcement learning with Prefix Sliding enables models to scale to reasoning traces exceeding 100,000 tokens. The system maintains performance by retaining only the initial prefix and a sliding window of the most recent few thousand tokens.
Sign in to suggest edits
Key sources
- SOURCE@muennighoff“vanilla full attention OOMs on long tasks & compaction loses important details”x.com
- SUPPORT@iscienceluvr“Training with Prefix Sliding using reinforcement learning can achieve better performance by enabling scaling to reasoning traces beyond a hundred thousand tokens”x.com
- SUPPORT@percyliang“Prefix Sliding for efficient test-time scaling”x.com