Paper Detail

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Baha Rababah, Cuneyt Gurcan Akcora, Carson K. Leung

arxiv Score 7.3

Published 2026-07-09 · First seen 2026-07-10

General AI

Abstract

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections. These findings expose an illusion of equivalence between base and quantized models and motivate behavioral evaluation beyond conventional performance metrics.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{rababah2026illusion,
  title = {The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs},
  author = {Baha Rababah and Cuneyt Gurcan Akcora and Carson K. Leung},
  year = {2026},
  abstract = {Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization sche},
  url = {https://arxiv.org/abs/2607.08734},
  keywords = {cs.AI},
  eprint = {2607.08734},
  archiveprefix = {arXiv},
}

Metadata

{}