Research Paper Cockpit

Daily Digest - 2026-08-07

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-08-18.

Papers

45 visible entries

huggingface Score 22.4

ChronoVision: Temporal Reasoning via Latent State Reconstruction

2026-08-06 · Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao

General AI

Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To …

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.2

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

2026-08-06 · Zelong Sun, Jun Wang, Kaicheng Yang, Tiancheng Gu, Ziyong Feng, Zhiwu Lu

General AI

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to con…

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.4

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

2026-08-06 · Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu

General AI

Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that rep…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.4

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

2026-08-06 · Jiaming Wei, Zekun Wu, Adriano Koshiyama, Maria Perez-Ortiz

Research Track B · General AI

Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask what choosing per task would buy. The modes are complementary: each solves tasks the others…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

2026-08-06 · Jonas Gann, Michael Gertz

General AI

Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning steps are difficult to verify and cannot be reliably attributed to specific evidence. Moreo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.2

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

2026-08-06 · Abdulkadir Külçe, Alihan Esen, Cağla Fikir, Berke Kurt, Kuzey Arar, Gökhan Ercan, Faik Boray Tek

General AI

This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a unified system. The core module is an agentic chatbot built on a ReAct loo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.2

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

2026-08-06 · Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer, Christopher Lee, Sajeev Singh, Piyum Zonooz, Navin Kumar, Zeeshan Ahmed, Priyadarshini Kachroo

General AI

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

2026-08-06 · ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng

General AI

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

2026-08-06 · Fardin Afdideh, Fernando Seoane, Farhad Abtahi

Research Track A · General AI

Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibration, and Multimodal Instruction Tuning. However, the literature remains fragmente…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

2026-08-06 · Tao Wang, Qihao Yang, Rongjiao Liang, Lianghong Lin, Haitao Wang, Xinyu Cao, Tianyong Hao

General AI

Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

2026-08-06 · Qifeng Zhang, Kaixiang Huang, Heng Dong, Huang Fang, Junting Chen, Junjie Zhu, Yonghang Chen, Zhiyu Zhang, Wei Li

General AI

Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.4

Continual Learning in Transition

2026-08-06 · Zhiyan Hou, Dan Zhang, Tao Feng, Liyuan Wang, Wei Li, Xiangzhao Hao, Hongyan An, Junfeng Fang, Haokai Ma, Zhaohui Xu, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua, Xinyu Tang

Research Track A · General AI

Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. However, emerging paradigms are reshaping the scope of CL beyond this traditional model adaptation view…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

2026-08-06 · Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan Xue

General AI

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.2

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

2026-08-06 · Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li

General AI

Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLL…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.2

The Bitter Lesson of Tool Calling

2026-08-06 · Ishan Patel, Sahil Sen, Elias Lumer, Vamse Kumar Subbiah

General AI

Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that chain and parallelize naturally. However, a systematic evaluation of tools as code on an established benchmark across curr…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

2026-08-06 · Xinye Wang, Junxiao Liu, Shujian Huang

General AI

Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising paradigm, providing dense token-level supervision on student-generated rollouts, yet their objec…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

2026-08-06 · Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang, Tongran Liu, Jingbo Zhu

General AI

Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation ari…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.4

DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval

2026-08-06 · Xi Chen, Xu Chen, Xiangyang Jia, Wei Wang, Xu Zhang, Zhenyuan Sun

Research Track A · General AI

With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challenging because scale variation and distribution shifts in RS aggravate cross-modal alignment space di…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.4

MASS: Multiplayer World Models with Authoritative Shared State

2026-08-06 · Ziqi Cai, Siqi Yang, Yimu Wang, Zixian Gao, Yunheng Liu, Shuchen Weng, Erwin Wu, Kaipeng Zhang, Boxin Shi

General AI

Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MAS (Multiplayer world models with Authoritative Shared State) to resolve this limitation. Inspired b…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

2026-08-06 · Noam Koren, Roy Bar-Haim, Abigail Goldsteen

General AI

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework th…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

2026-08-06 · Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo

General AI

Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our objective is to assess the impact of SPT on the performance and scalabil…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

Learning Globally Reusable Skills for Coding Agents

2026-08-06 · Chen Yang, Jiashuo Tian, Ziqi Wang, Xinyin Liu, Meiru Ye, Junjie Chen

Research Track A · General AI

Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches typically treat skill evolution as a sequence of local updates, overlooking relationships among skills and often producing overfitted skill updates that fail to generali…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

2026-08-06 · Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu

General AI

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on ques…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.4

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

2026-08-05 · Xuying Ning, Dongqi Fu, Tianxin Wei, Hanqing Zeng, Yuanchen Bei, Bingxuan Li, Zihao Li, Qifan Wang, Xiang Shen, Yifan Wu, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He

Research Track A · General AI

Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled challenges: state formation from noisy interaction traces and runtime control over externa…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.4

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

2026-08-06 · Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang

General AI

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser super…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.4

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

2026-08-06 · Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong

General AI

Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes. This paper develops an implementation roadmap f…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Learning When to Trust via Selective Context Preference Optimization

2026-08-06 · Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong

General AI

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We rec…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

2026-08-06 · Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, Furong Huang

General AI

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executab…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.9

STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models

2026-08-06 · Songpan Gao, Yajie Zhang, Guanxing Chen, Jiayu Qian, Zhenzhen Liu, Shijun Li, Xiaowei Zhu, Yao Hu, Kay Chen Tan, Yu-An Huang, Shiqi Wang, Zhi-An Huang

Research Track A · General AI

Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs signi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.2

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

2026-08-06 · Boning Li, Yu Chen, Longbo Huang

General AI

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while na…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.2

Automatic Translation of Unstructured Requirements into Linear Temporal Logic through Large Language Models

2026-08-06 · Alexandra Newcomb, Omar Ochoa

General AI

Automatically translating unstructured natural language requirements into formal specifications remains a challenge in requirements engineering and formal methods, particularly for safety- and mission-critical systems whose verification depends on mathematically precise specifications. This paper evaluates whether cont…

Review
pending
Role
unreviewed
Read
now
huggingface Score 9.4

WorldClaw: Agentic 3D Open-World Generation at Scale

2026-08-05 · Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang

General AI

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.4

On-Policy Delta Distillation for Multilingual Math Reasoning

2026-08-06 · Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

General AI

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japan…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

2026-08-06 · Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju

General AI

Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with \texttt{**kern} score encodings for 9,468 clips originating from 6,066 unique songs, the fi…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

On-Policy Self-Distillation without Any Supervision

2026-08-06 · Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos

General AI

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-d…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

2026-08-06 · Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li

General AI

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenge…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.7

QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction

2026-08-06 · Mutasim Fuad Sarker, Adiba Rahman Namira, Wafa Binte Alam, Md Adnan Arefeen, Mahzabeen Emu, Sumaiya Tabassum Nimi

General AI

Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admission. Such approaches ignore the temporal p…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.4

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

2026-08-06 · Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu

General AI

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their gen…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

2026-08-06 · Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia

General AI

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous termin…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Sensor-Level Fault Diagnosis for Automotive Software Validation Using Large Language Models

2026-08-06 · Mohammad Abboush, Hamza Ouarrad, Andreas Rausch

General AI

The pre-series validation of automotive software on hardware-in-the-loop (HIL) platforms produces large volumes of multivariate sensor recordings whose assessment against functional safety requirements exceeds what manual review can sustain at campaign scale. Threshold-based tooling reports that a deviation has occurre…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

A Master-Salve Robot Manipulator for Needle-Based Teleoperation in MRI Chamber

2026-08-06 · Omar Curiel, Jing-Yuan Huang, Po-Chih Chen, Ji Ma, Qing Dai, Wenqi Zhou, David Lu, Holden H. Wu, Tsu-Chin Tsao

General AI

We present a MR safe, master-slave robot manipulator for abdominal interventions in the MRI chamber. A human operated 2+1-DoF master controller manipulator transmits motion and force to a 2+1-DoF slave manipulator via fluid transmission. Jointly, a digital master controller provides multimodal control capability beyond…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

2026-08-06 · Arya Labroo, Mengjie Qian, Kate Knill

General AI

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based found…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

2026-08-06 · Praphul Chandra, Sujit Gujar, Ganesh Ghalme

General AI

We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism seeks to establish the …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.2

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

2026-08-06 · Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang, Xichen Yuan, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Yang, Shanghang Zhang

General AI

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-ce…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.2

Fair and Efficient Balanced Allocations for Additive Valuations

2026-08-06 · Benjamin Cookson, Nisarg Shah, Paritosh Verma

General AI

We study the existence of fair and efficient allocations of indivisible goods under the balancedness constraint, which requires that any two agents' bundles differ in size by at most one. Our main result establishes the existence of balanced allocations that satisfy envy-freeness up to one good (EF1) and fractional Par…

Review
pending
Role
unreviewed
Read
later