---
format: "aidr-story-markdown/v1"
id: "b6cab15d8a527082365458d2baccc2b9ffeed91cae005e59d9972e020b9dc353"
canonical_url: "https://aidr.today/b6cab15d?lang=en"
title: "The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-02T04:00:00.000Z"
category: "Research"
topics: ["llm"]
source_urls: ["https://arxiv.org/abs/2610.00054"]
summary: "arXiv:2610.00054v1 Announce Type: new Abstract: Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood-scoring evaluation harnesses produce. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama-3.1-8B and Phi-3.5-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89.7% of them when the responses are swapped, against 47.5% read after generation (paired difference +0.422, 95% CI [+0.365, +0.467]). The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one. A second, smaller failure occurs even when the judge does lead with a verdi"
---

# The First Token Is Not the Verdict: Hidden Costs of Reading LLM Judges Without Generating

> [Open the canonical story](<https://aidr.today/b6cab15d?lang=en>)

**Published:** 2026-10-02T04:00:00.000Z
**Category:** Research
**Topics:** llm

## Summary

arXiv:2610\.00054v1 Announce Type: new Abstract: Reading an LLM judge's verdict from the logits of its first generated token is cheap, requires no generation, and is exactly what constrained decoding and likelihood\-scoring evaluation harnesses produce\. We show that this readout distorts position bias in one direction: it overstates it in every condition we test, so figures obtained this way behave as upper bounds\. The mechanism is that judges do not always lead with a verdict token, on 12% to 49% of pairs for three Qwen3 judges and under 3% for Llama\-3\.1\-8B and Phi\-3\.5\-mini, and forcing a read on those pairs returns whichever response was shown first rather than a judgment\. Pooled over the 924 pairs where a judge did not commit, the forced read flips on 89\.7% of them when the responses are swapped, against 47\.5% read after generation \(paired difference \+0\.422, 95% CI \[\+0\.365, \+0\.467\]\)\. The distortion is specific to what is measured: it moves position bias by 42 points while moving judge accuracy by under one point in seven of ten conditions, so it misleads whoever audits a judge rather than whoever uses one\. A second, smaller failure occurs even when the judge does lead with a verdi

## Sources

- [Story source](<https://arxiv.org/abs/2610.00054>)

