---
format: "aidr-story-markdown/v1"
id: "4e74c3c42aee8b32e32f08d4da646c92ee4b825b55ba7dd9be36ac81fc078233"
canonical_url: "https://aidr.today/4e74c3c4?lang=en"
title: "Prefix Sliding Speeds AI Models 3x for Reasoning Traces Over 100,000 Tokens"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-27T17:02:04.000Z"
category: "Research"
topics: ["llm","inference","reasoning","efficiency"]
source_urls: ["https://huggingnews.com/ai/prefix-sliding-speeds-ai-models-3x-for-reasoning-traces-over-100000-toke-593254ff","https://x.com/Muennighoff/status/2092960068685692974","https://x.com/iScienceLuvr/status/2092902633468022979","https://x.com/percyliang/status/2092986067230269655"]
summary: "Researchers developed a technique to discard intermediate reasoning tokens that lose importance as a model continues to process, allowing existing systems to run 3x faster without training. The approach, called Prefix Sliding, avoids the memory failures associated with vanilla full attention and the loss of detail common in compaction methods. Using reinforcement learning with Prefix Sliding enables models to scale to reasoning traces exceeding 100,000 tokens. The system maintains performance by retaining only the initial prefix and a sliding window of the most recent few thousand tokens."
---

# Prefix Sliding Speeds AI Models 3x for Reasoning Traces Over 100,000 Tokens

> [Open the canonical story](<https://aidr.today/4e74c3c4?lang=en>)

**Published:** 2026-08-27T17:02:04.000Z
**Category:** Research
**Topics:** llm, inference, reasoning, efficiency

## Summary

Researchers developed a technique to discard intermediate reasoning tokens that lose importance as a model continues to process, allowing existing systems to run 3x faster without training\. The approach, called Prefix Sliding, avoids the memory failures associated with vanilla full attention and the loss of detail common in compaction methods\. Using reinforcement learning with Prefix Sliding enables models to scale to reasoning traces exceeding 100,000 tokens\. The system maintains performance by retaining only the initial prefix and a sliding window of the most recent few thousand tokens\.

## Sources

- [Story source](<https://huggingnews.com/ai/prefix-sliding-speeds-ai-models-3x-for-reasoning-traces-over-100000-toke-593254ff>)
- [Story source](<https://x.com/Muennighoff/status/2092960068685692974>)
- [Supporting source](<https://x.com/iScienceLuvr/status/2092902633468022979>)
- [Supporting source](<https://x.com/percyliang/status/2092986067230269655>)

