---
format: "aidr-story-markdown/v1"
id: "ba29a6eddf3cbc7473f2fc2115c9560ef14261abb035316d9401ff9705ae378e"
canonical_url: "https://aidr.today/ba29a6ed?lang=en"
title: "CLM-8B Hits 87.6% on Terminal-Bench to Set New Agentic Coding Record"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-24T16:40:30.000Z"
category: "Research"
topics: ["clm-8b","terminal-bench","agentic-coding","inference","benchmark","llm","coding","reasoning"]
source_urls: ["https://marketbrief.now/ai/clm-8b-hits-876percent-on-terminal-bench-to-set-new-agentic-coding-recor-932dacd2","https://huggingnews.com/ai/clm-8b-hits-876percent-on-terminal-bench-to-set-new-agentic-coding-recor-932dacd2"]
summary: "The new Contrastive Language Model delivers up to 9x faster inference than the Jev model while achieving an 81.6% score on the DeepSWE agentic coding benchmark. Known as CLM-8B, this \"System One\" model uses a contrastive learning objective to connect states and actions, allowing it to function as a more effective verifier for long-horizon tasks than previous models like Jev. The developers reduced inference latency by disaggregating states and actions so their embeddings can be cached and reused independently when action sets remain fixed. CLM-8B was trained on internet-scale data and follows scaling laws where test contrastive loss decreases as a power law relative to model size, training compute, and dataset size."
---

# CLM\-8B Hits 87\.6% on Terminal\-Bench to Set New Agentic Coding Record

> [Open the canonical story](<https://aidr.today/ba29a6ed?lang=en>)

**Published:** 2026-09-24T16:40:30.000Z
**Category:** Research
**Topics:** clm\-8b, terminal\-bench, agentic\-coding, inference, benchmark, llm, coding, reasoning

## Summary

The new Contrastive Language Model delivers up to 9x faster inference than the Jev model while achieving an 81\.6% score on the DeepSWE agentic coding benchmark\. Known as CLM\-8B, this "System One" model uses a contrastive learning objective to connect states and actions, allowing it to function as a more effective verifier for long\-horizon tasks than previous models like Jev\. The developers reduced inference latency by disaggregating states and actions so their embeddings can be cached and reused independently when action sets remain fixed\. CLM\-8B was trained on internet\-scale data and follows scaling laws where test contrastive loss decreases as a power law relative to model size, training compute, and dataset size\.

## Sources

- [Story source](<https://marketbrief.now/ai/clm-8b-hits-876percent-on-terminal-bench-to-set-new-agentic-coding-recor-932dacd2>)
- [Story source](<https://huggingnews.com/ai/clm-8b-hits-876percent-on-terminal-bench-to-set-new-agentic-coding-recor-932dacd2>)

