Paper Detail

A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

Toqeer Ehsan, Thamar Solorio

arxiv Score 9.8

Published 2026-09-16 · First seen 2026-09-17

General AI

Abstract

Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper, we present a multi-step framework to enhance the annotation quality of NER datasets by employing automated techniques. We propose a frequency-based iterative approach that leverages self-training and a dual-threshold mechanism to enhance inference confidence. Experimental evaluations on different NER datasets demonstrate significant improvements in NER performance with respect to the original datasets. This work further explores the potential of generative Large Language Models (LLMs) to perform NER for low-resource languages.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{ehsan2026scalable,
  title = {A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages},
  author = {Toqeer Ehsan and Thamar Solorio},
  year = {2026},
  abstract = {Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper, we present a multi-step framework to enhance the annotation quality of NER datasets by employing automated techniques. We propose a frequency-based iterative approach that leverages self-training and a dual-threshold mechanism to enhance inference confidence. Experimental evaluations on different NER datasets demonstrate signif},
  url = {https://arxiv.org/abs/2609.18739},
  keywords = {cs.CL, cs.AI},
  eprint = {2609.18739},
  archiveprefix = {arXiv},
}

Metadata

{}