Paper Detail

A Living Benchmark for Information Retrieval from Electronic Health Records

Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer

arxiv Score 18.2

Published 2026-09-24 · First seen 2026-09-26

General AI

Abstract

Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{cahoon2026living,
  title = {A Living Benchmark for Information Retrieval from Electronic Health Records},
  author = {Jordan L. Cahoon and Chloe O. Stanwyck and Sulaiman Somani and Philip Chung and Kevin R Keet and Kameron C. Black and Andrea T. Fisher and Sarita Khemani and Jerry Liu and Stephen Ma and Saloni K. Maharaj and Rita M. Pandya and Eduardo Perez-Guerrero and Priyanka Pillai and Lisa Shieh and David J. H. Wu and James Xie and James C. McAvoy and Teresa Nguyen and Jessica Tran and Lucy Yin and Bridget Lin and Alison Callahan and Jason A. Fries and Nigam H. Shah and Emily Alsentzer},
  year = {2026},
  abstract = {Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from lon},
  url = {https://arxiv.org/abs/2609.30205},
  keywords = {cs.AI},
  eprint = {2609.30205},
  archiveprefix = {arXiv},
}

Metadata

{}