Researchers at Stanford University developed a method to reduce the computational cost of long-chain-of-thought inference by discarding unimportant intermediate tokens from memory. The technique, called Prefix Sliding, maintains the original prompt and a sliding window of recent reasoning, allowing existing models to run up to 3x faster while matching full-attention performance without requiring retraining.

The approach addresses the tendency for inference costs and memory usage to grow as reasoning chains lengthen. When combined with reinforcement learning fine-tuning, Prefix Sliding allows AI agents to scale reasoning beyond 100,000 tokens by capping total memory regardless of the chain length.

Sign in to suggest edits

Key sources

  1. SOURCE@shipfrontierai“Cost per new token stays constant instead of growing with chain length”x.com
  2. SUPPORT@omarsar0“Intermediate tokens steadily lose importance as the model keeps going”x.com
  3. SUPPORT@askalphaxiv“Without retraining, Prefix Sliding is up to 3x faster while matching performance”x.com
Markdown