---
format: "aidr-story-markdown/v1"
id: "54c50ce8b3586d28068de0903599b860a57bddbaf67558eeceec3519d33156d6"
canonical_url: "https://aidr.today/54c50ce8?lang=en"
title: "LFM2.5-2.6B Model Hits 54% Solve Rate With New Multi Harness RL"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-01T15:55:57.000Z"
category: "Research"
topics: ["agent","llm"]
source_urls: ["https://huggingnews.com/ai/lfm25-26b-model-hits-54percent-solve-rate-with-new-multi-harness-rl-e6d96d0e","https://marketbrief.now/ai/lfm25-26b-model-hits-54percent-solve-rate-with-new-multi-harness-rl-e6d96d0e"]
summary: "← Back to live feed · 1 stories across 1 day A new capture proxy for OpenEnv permits the direct training of AI models via reinforcement learning within tool-using environments like Claude Code, Codex, and OpenCode. Developed by Adithya S.K. and Ben Burtenshaw, the tool sits between the harness and the model to record exact token IDs and logprobs of every call, allowing reinforcement learning on any task set without requiring modifications to the harness or training code. The LFM2.5-2.6B model demonstrated the system's utility, raising its average solve rate across four harnesses from 42% to 54% while reducing tool calls by 31%. In specific tests within Claude Code, the model's solve rate increased from 33% to 49%. The workflow utilizes Harbor for task sandboxing and the TRL library for asynchronous GRPO training."
---

# LFM2\.5\-2\.6B Model Hits 54% Solve Rate With New Multi Harness RL

> [Open the canonical story](<https://aidr.today/54c50ce8?lang=en>)

**Published:** 2026-10-01T15:55:57.000Z
**Category:** Research
**Topics:** agent, llm

## Summary

← Back to live feed · 1 stories across 1 day A new capture proxy for OpenEnv permits the direct training of AI models via reinforcement learning within tool\-using environments like Claude Code, Codex, and OpenCode\. Developed by Adithya S\.K\. and Ben Burtenshaw, the tool sits between the harness and the model to record exact token IDs and logprobs of every call, allowing reinforcement learning on any task set without requiring modifications to the harness or training code\. The LFM2\.5\-2\.6B model demonstrated the system's utility, raising its average solve rate across four harnesses from 42% to 54% while reducing tool calls by 31%\. In specific tests within Claude Code, the model's solve rate increased from 33% to 49%\. The workflow utilizes Harbor for task sandboxing and the TRL library for asynchronous GRPO training\.

## Sources

- [Story source](<https://huggingnews.com/ai/lfm25-26b-model-hits-54percent-solve-rate-with-new-multi-harness-rl-e6d96d0e>)
- [Story source](<https://marketbrief.now/ai/lfm25-26b-model-hits-54percent-solve-rate-with-new-multi-harness-rl-e6d96d0e>)

