---
format: "aidr-story-markdown/v1"
id: "d66945ef830bb2f2ad0617705b74c2613816d03b3e615a5c7d93f809901ea95b"
canonical_url: "https://aidr.today/d66945ef?lang=en"
title: "Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-02T04:00:00.000Z"
category: "Research"
topics: ["agent"]
source_urls: ["https://arxiv.org/abs/2610.00025"]
summary: "arXiv:2610.00025v1 Announce Type: new Abstract: Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics, each with a pre-specified threshold $\\tau$ anchored to a cheap non-LLM baseline and a CI-aware eligibility rule (a configuration passes only if its confidence bound clears $\\tau$). Sweeping Qwen3 0.6/1.7/4/8B at their best (FP16, greedy, one frozen prompt, no tuning), we find an eligibility gap: 0 of 16 (4 tasks $\\times$ 4 models) configurations pass (verified by checking the raw outputs and parser behavior). A logprob decision-threshold diagnostic (T1/T3/T4; T2 via a context-length/cascade probe) separates the failures into capability deficits and failures that can be addressed by changing the decoding threshold (4 regimes). Quantization to 4-bit (RTN/GPTQ/AWQ) does damage that depends on m"
---

# Measuring the Microtask Eligibility Gap: When Is an Off\-the\-Shelf SLM Enough for an Agent Harness?

> [Open the canonical story](<https://aidr.today/d66945ef?lang=en>)

**Published:** 2026-10-02T04:00:00.000Z
**Category:** Research
**Topics:** agent

## Summary

arXiv:2610\.00025v1 Announce Type: new Abstract: Agent harnesses increasingly want to run small language models \(SLMs\) on the microtasks around a frontier large language model \(LLM\) planner: auto\-approving shell commands, writing memory, selecting tools, ranking past turns\. We ask whether off\-the\-shelf SLMs meet practitioner\-defined thresholds and, when they fail, why, and whether quantization changes the answer\. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics, each with a pre\-specified threshold $\\tau$ anchored to a cheap non\-LLM baseline and a CI\-aware eligibility rule \(a configuration passes only if its confidence bound clears $\\tau$\)\. Sweeping Qwen3 0\.6/1\.7/4/8B at their best \(FP16, greedy, one frozen prompt, no tuning\), we find an eligibility gap: 0 of 16 \(4 tasks $\\times$ 4 models\) configurations pass \(verified by checking the raw outputs and parser behavior\)\. A logprob decision\-threshold diagnostic \(T1/T3/T4; T2 via a context\-length/cascade probe\) separates the failures into capability deficits and failures that can be addressed by changing the decoding threshold \(4 regimes\)\. Quantization to 4\-bit \(RTN/GPTQ/AWQ\) does damage that depends on m

## Sources

- [Story source](<https://arxiv.org/abs/2610.00025>)

