Paper Detail

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Orăsan, Chrysoula Zerva, Ricardo Rei, Frédéric Blain, André F. T. Martins, Marco Turchi, Matteo Negri, Rajen Chatterjee, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya

arxiv Score 9.3

Published 2026-08-17 · First seen 2026-08-18

General AI

Abstract

Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes. On it, we benchmark six prompted LLMs and three COMET metrics on segment-level QE, and three systems on APE. Two of the axes are defined partly on the direct assessment and select a compressed slice of it, so each axis is compared against a control drawn from the same language pair with the same score distribution. Only one survives that control: segments whose holistic and token-level quality signals conflict are ranked worse than equally-scored segments of the same language, for all nine systems and all seven pairs that carry the axis. Annotator disagreement, which looks second-hardest without the control, has no effect with it. Few-shot prompting costs every model $\leq$ $3.4$B both correlation and output-format compliance. Within-language accuracy does not make scores comparable across pairs: of the three trained metrics, the one with the best within-language correlation loses most when the pairs are pooled. The benchmark and code will be released.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{kanojia2026indicqe,
  title = {IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages},
  author = {Diptesh Kanojia and Archchana Sindhujan and Sourabh Deoghare and Daria Sokova and Shenbin Qian and Girish Koushik and Tharindu Ranasinghe and Constantin Orăsan and Chrysoula Zerva and Ricardo Rei and Frédéric Blain and André F. T. Martins and Marco Turchi and Matteo Negri and Rajen Chatterjee and Anoop Kunchukuttan and Mitesh M. Khapra and Pushpak Bhattacharyya},
  year = {2026},
  abstract = {Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \textbackslash{}indicqe: \$126\{,\}754\$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an e},
  url = {https://arxiv.org/abs/2608.16344},
  keywords = {cs.CL},
  eprint = {2608.16344},
  archiveprefix = {arXiv},
}

Metadata

{}