Paper Detail

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

Qiyao Yan, Chenpeng Wang, Liangming Pan

arxiv Score 14.3

Published 2026-08-31 · First seen 2026-09-01

General AI

Abstract

When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct answers even when native sequence scoring completely collapses due to structural biases. To test whether instance-specific logic survives this collapse, we introduce a diagnostic protocol using a minimal, target-label-free additive correction. Fitting just two parameters on as few as 25 unlabeled examples recovers 9--34 accuracy points for Qwen3.5 models, transferring successfully to OLMo-2-1B and Llama-3.1-8B. Crucially, these recovered decisions persist on hard instances unresolved by simple lexical overlap and significantly exceed count-preserving permutation baselines. Our results show that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{yan2026wrong,
  title = {Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores},
  author = {Qiyao Yan and Chenpeng Wang and Liangming Pan},
  year = {2026},
  abstract = {When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct answers even when native sequence scoring completely collapses due to structural biases. To test whether instance-specific logic survives this collapse, we introduce a diagnostic p},
  url = {https://arxiv.org/abs/2608.31068},
  keywords = {cs.AI},
  eprint = {2608.31068},
  archiveprefix = {arXiv},
}

Metadata

{}