Paper Detail

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu

arxiv Score 17.3

Published 2026-08-17 · First seen 2026-08-18

General AI

Abstract

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{chen2026harnesseval,
  title = {HarnessEval-W: Agentifying the Evaluation of Visual Worlds},
  author = {Weiliang Chen and Haowen Sun and Jun Gao and Jiawei Chi and Hanyang Wang and Qiyu Dai and Yihao Li and Hao Li and Jingnan Gao and Yi-Hsin Hung and Xingzhuo Guo and Shangchen Miao and Zhiyuan Shi and Xiang Li and Fengrui Tian and Weihua Du and Ziqi Huang and Shenyuan Gao and Siqiao Huang and Mingyu Liu and Yifei Li and Shizun Wang and Xi Wang and Tianqi Zhang and Xue Luo and Xiyin Ren and Jinshan Ren and Xiaoyang Shen and Xiaobo Hu and Zhiyang Dou and Mingyu Ding and Yichao Yan and Xinchao Wang and Yizhou Wang and Shilong Liu and Wenzhao Zheng and Yueqi Duan and Yuan Gong and Ziwei Liu and Ming-Yu Liu and Jialong Wu and Jiangran Lyu and Fangfu Liu},
  year = {2026},
  abstract = {A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-},
  url = {https://arxiv.org/abs/2608.16859},
  keywords = {cs.CV, HarnessEval-W, agentified evaluation pipeline, harness paradigm, world model benchmarking, sub-agents, evidence tree, reasoning chain, code available, huggingface daily},
  eprint = {2608.16859},
  archiveprefix = {arXiv},
}

Metadata

{}