---
format: "aidr-story-markdown/v1"
id: "182622bf1f20238f6f2bfff70324a1e402fd10a7dd123aa3db0ab5bb17a413fd"
canonical_url: "https://aidr.today/182622bf?lang=en"
title: "Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-02T04:00:00.000Z"
category: "Research"
topics: ["reasoning"]
source_urls: ["https://arxiv.org/abs/2610.00047"]
summary: "arXiv:2610.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firin"
---

# Characterizing a Configuration Where Inference\-Time PRM\-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

> [Open the canonical story](<https://aidr.today/182622bf?lang=en>)

**Published:** 2026-10-02T04:00:00.000Z
**Category:** Research
**Topics:** reasoning

## Summary

arXiv:2610\.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain\-of\-thought has motivated inference\-time interventions built on a natural design: when a process reward model \(PRM\) prunes a chain, its high\-PRM prefix is extracted and grafted verbatim as an in\-context demonstration into a still\-decoding sibling\. We isolate this mechanism, PRM\-Pruned Fragment Grafting \(PPFG\), as the most cost\-minimal operationalization of cross\-trajectory step\-level transfer, and test it at the operating point where prior fragment\-grafting work reports gains only under additional compensating ingredients\. On Qwen2\.5\-7B\-Instruct with Math\-Shepherd on full MATH500 \(n\=500, three seeds\), PPFG in both stagnation\- and random\-targeting variants is statistically indistinguishable from an independent parallel\-CoT baseline on every measured axis\. We characterize why: a four\-bucket classification of 322 stagnation\-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change\. No compound\-gate refinement jointly achieves well\-targeted firin

## Sources

- [Story source](<https://arxiv.org/abs/2610.00047>)

