Research Paper Cockpit

Daily Digest - 2026-07-14

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-07-25.

Papers

54 visible entries

huggingface Score 24.4

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

2026-07-11 · Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Mingyang Yin, Zedong Chu, Mu Xu

General AI

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-l…

Review
pending
Role
unreviewed
Read
now
arxiv Score 23.2

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

2026-07-13 · Kerui Chen, Jinglu Wang, Xiaoyi Zhang, Yan Lu

General AI

Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple ca…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.4

Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation

2026-07-13 · Paul A. Bereuter, Mark D. Plumbley, Alois Sontacchi

Research Track A

State-of-the-art speech enhancement models benefit from large-scale labeled datasets, whereas singing voice separation models suffer from limited available training data. To address this limitation, we formulate singing voice separation as domain adaptation from speech enhancement to singing voice separation. We invest…

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.0

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

2026-07-08 · Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen

General AI

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored i…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.2

The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory

2026-07-12 · Yixiong Chen, Xinyi Bai, Alan Yuille

Research Track B · General AI

Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve f…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

2026-07-11 · Chengguang Gan, Hanjun Wei, Yunhao Liang, Zhixi Cai, Qinghao Zhang, Shiwen Ni

Research Track B · General AI

Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a few buttons on a single page: it takes a sequence of actions that unfo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

Evidence-Backed Video Question Answering

2026-07-13 · Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu, Caiming Xiong, Silvio Savarese, Chen Sun, Juan Carlos Niebles

General AI

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such …

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

2026-07-13 · Mikhail Komarov, Ivan Bondarenko, Stanislav Shtuka, Oleg Sedukhin, Roman Shuvalov, Yana Dementyeva, Matvey Solovyov, Nikolay O. Nikitin

Research Track A · General AI

Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an open-source modular GraphRAG engine, addresses this by separating extraction fro…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.9

CA-DGCL: Dynamic Graph Continual Learning via Condensation and Attachment

2026-07-13 · Tingxu Yan Ye Yuan

Research Track A

Dynamic graph continual learning (DGCL) is an effective manner for handling catastrophic forgetting in dynamic graphs. However, existing DGCL methods underutilize temporal information across graph snapshots. To address this critical issue, we propose a novel framework for Dynamic Graph Continual Learning via Condensati…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.9

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR

2026-07-13 · Ziang Ren, Guodong Lin, Yuchen Ai, Kaize Tan, Wei-Qiang Zhang

Research Track A

Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue, existing methods struggle to regulate cross-task interference in multilingual settings, where…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.5

Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models

2026-07-08 · Sergi Masip, Alicja Dobrzeniecka, Jonathan Swinnen, Joachim Collin, Bartłomiej Twardowski, Szymon Łukasik, Tinne Tuytelaars

Research Track A · General AI

Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- require models to adapt continuously from unlabeled streams. This has led to the development of continual self-supervised learning (CSSL), a rapidly growing area that lacks a dedicated,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

2026-07-13 · Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen

General AI

Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and ofte…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes

2026-07-13 · Iman Johary, Guillaume Bied, Alexandru C. Mara, Tijl De Bie

General AI

Large-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from pre-standardized occupational codes with LLM-synthesized rather than authentic free text. We prese…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

2026-07-13 · Kaixin Ma, Di Feng, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan

General AI

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs i…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

2026-07-13 · Runhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao

General AI

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the origin…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

LightMem-Ego: Your AI Memory for Everyday Life

2026-07-13 · Yijun Chen, Boyi Xiao, Yixian Zhao, Haoting Xia, Buqiang Xu, Jizhan Fang, Yanya Li, Yaqi Zheng, Xuehai Wang, Zirui Xue, Liuxin Zhang, Hui Li, Ningyu Zhang

General AI

Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory that can continuously accumulate, organize, and retrieve long-term experiences, which remains challeng…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.2

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

2026-07-13 · Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su, Yiming Wang, Xiaojian Zhang, Jingyu Zhu, Junhao Zhu, Zhuowen Liang, Jiazhen Peng, Lianggui Weng, Zhihao Ding, Kerui Yi, Qifeng Wang, Rong Zhu, Bolin Ding, Liyu Mou, Jingren Zhou

General AI

Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, ex…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

2026-07-13 · Yuanzhi Liang, Xufeng Zhan, Haibin Huang, Chi Zhang, Xuelong Li

General AI

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted cons…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis

2026-07-13 · Yongqian Sun, Rongchen Gao, Yu Luo, Wenwei Gu, Shenglin Zhang, Qingyi Guo, Qiuai Fu, Yaoliang Wu, Dan Pei

General AI

Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experience. Existing LLM-based methods improve diagnosis through agentic reasoning or knowledge augmentation, but they often lack a mechanism to coordinate the evolving diagnostic state wi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

2026-07-13 · Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Cheewei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, Jianxin Lin

General AI

The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from th…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.4

Can a Language Model Learn Facts Continually in Its Weights?

2026-07-13 · Charles O'Neill

Research Track A

Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided. We follow invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later wri…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

2026-07-13 · Esteban U. Vega Barajas

General AI

Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a h…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

From Global to Factor-Wise Expert Composition in Discrete Diffusion Models

2026-07-13 · Haozhe Huang, Yudong Xu, Abhijoy Mandal, Alán Aspuru-Guzik

General AI

Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.4

ABot-N1: Toward a General Visual Language Navigation Foundation Model

2026-07-11 · Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu

General AI

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?

2026-07-13 · Nishant Aggarwal, Ayushi Dubal, Sreeraj Kannakarankodi, Ian McDougall, Adarsh Mittal, Vishnu Ramadas, Noah Scott, Ranganath Selagamsetty, Weichu Yang, Karthikeyan Sankaralingam

General AI

Can large language models perform deep technical comprehension of computer architecture papers -- not summarization, but structured critique that names the core mechanism, surfaces buried assumptions, and connects a contribution beyond its own scope? We study Gauntlet, an open-source pipeline that analyzes a paper thro…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

How Temperature Shapes Ideological Discourse in Retrieval-Augmented Generation?

2026-07-13 · Elmira Salari, Hazem Amamou, José Victor de Souza, Shruti Kshirsagar, Maria Nunes Delfino, Anderson Avila

General AI

Retrieval-Augmented Generation (RAG) has been increasingly adopted to reduce hallucinations and strengthen the factual grounding of large language models (LLMs). While robustness to errors in the retrieval process has been explored, the impact of ideological bias on LLM outputs has been overlooked. For instance, if the…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

2026-07-13 · Ayoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu, Peter Railton, Lu Wang

General AI

Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reason…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.9

UNIT: Unleash Large Language Models Potential for Graph Continual Learning

2026-07-11 · Tairan Huang, Yili Wang, Beibei Hu, Yiting Shi, Qiutong Li, Changlong He, Jianliang Gao

Research Track A · General AI

In real-world multimodal web scenarios, graph-structured data often arrives in a streaming manner, making graph continual learning a crucial paradigm for continuously modeling such evolving structures. However, existing graph continual learning methods still face two fundamental challenges. 1) semantic-structural separ…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.2

Multi-Agent Reinforcement Learning for C-V2X RAT Selection

2026-07-13 · Moritz Schaffenroth, Uwe Kölbel, Heike Lepke, Alexander Prinz, Alfred Höß

General AI

Vehicles are increasingly equipped with advanced V2X communication capabilities. While early V2X apps utilized services such as Cooperative Awareness Messages, recent developments have allowed more advanced applications including cooperative driving, shared perception, and sensor-sharing services. The broader mix of ap…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.9

Stateful Worlds, Stateless Elasticity: Exact-State Serving for Interactive World Models

2026-07-11 · Jin Li, Jiawei Chen

Research Track A

A persistent interactive world model keeps its running state resident on the GPU that serves it: a multi-gigabyte attention cache, almost all of it rewritten at every generation step. That state cannot be recomputed in interactive time or approximated without changing the world, so a live session pins its device. The p…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.9

Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game

2026-07-13 · Landon Liu, Mary Kelly, Alan Tsang

Research Track A · General AI

Shared meaning in language requires people to learn and agree on categories. We ask how characteristics of agents' memories change the emergence and evolution of shared meaning. Without a coordination game, models of conceptual semantics cannot explain how shared meaning emerges and changes in groups of people; however…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.2

LSTrans: Efficient Knowledge Transfer for Lightweight and Automated ECG Classification

2026-07-12 · Yi Zhao, Jiajun Gao, Chenyang Xu, Yuxi Zhou, Hao Wang

General AI

Deploying deep learning models for automated electrocardiogram classification on resource-constrained wearable devices remains challenging due to high computational costs. To address this, we propose LSTrans, a lightweight hybrid model designed for efficient and sensitive ECG analysis. LSTrans introduces a specialized …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

AutoPath: Learning Transferable Goal-Conditioned Stochastic Path Prior for Safe Navigation Without Human Demonstrations

2026-07-13 · Ziyang Zhang, Boyang Zhou, Zesong Yang, Haocheng Peng, Zeming Gai, Xiao Liang, Yujun Shen, Danping Zou, Ruizhen Hu, Hujun Bao, Zhaopeng Cui

General AI

Real-time navigation in cluttered and dynamic environments requires collision-free and dynamically feasible motion under limited perception. However, feasible navigation behaviors are inherently multimodal because multiple paths may exist around obstacles. In this paper, we formulate navigation as learning a transferab…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments

2026-07-13 · Divya Mereddy, Jeevan Beedareddy

General AI

This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated i…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems

2026-07-13 · Yibo Hu, Ren Wang

General AI

As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the att…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.0

Weak-to-Strong Generalization via Direct On-Policy Distillation

2026-07-08 · Shiyuan Feng, Huan-ang Gao, Haohan Chi, Hanlin Wu, Zhilong Zhang, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

General AI

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.4

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

2026-07-13 · Daocheng Fu, Rong Wu, Yu Yang, Xuemeng Yang, Jianbiao Mei, Licheng Wen, Pinlong Cai, Yong Liu, Botian Shi, Yu Qiao

General AI

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation

2026-07-11 · Finn Ferchau, Daniel Pommer, Cristian Axenie

General AI

Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We present a systematic study of Low-Rank Adaptation (LoRA) for $π_0$, a flow-matching VLA, ev…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

TabLoRA: Parameter-Efficient Low-Rank Ensemble Learning for Large-Scale Tabular Data

2026-07-11 · Jiaqi Luo, Shixin Xu

General AI

Tabular learning is still dominated by gradient-boosted decision trees (GBDTs), while recent deep learning approaches have become increasingly competitive. However, applying deep tabular models to large-scale datasets remains challenging, as large sample sizes, high feature dimensionality, or many target classes can in…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems

2026-07-13 · Sheng Li, Jing Li, Felix Schijve, Jun Hu, Emilia Barakova

General AI

Automatic speech recognition (ASR) has become a critical component of modern robotic systems because it is one of the most natural and intuitive ways for humans to interact with robots. A commonly used method is to directly use API services online. But is that all we can do? This article provides an overview of how ASR…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts

2026-07-13 · Christelle Schneuwly Diaz, Narmina Baghirova, Duy-Thanh Vu, Duy-Cat Can, Gilles Allali, Philippe Ryvlin, Oliver Y. Chén

General AI

Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data. Left unaddressed, these barriers prevent reliable disease modelling and hinder effective clinical evaluation. Conventional imputation strategies in…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

Metacognition in LLMs: Foundations, Progress, and Opportunities

2026-07-13 · Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, Jordan Thomas, Mark Steyvers, Arman Cohan

General AI

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse re…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

Paradoxes of Game Theoretic Equilibria and Price of Anarchy

2026-07-13 · Georgios Piliouras, Ian Gemp, Siqi Liu, Luke Marris

General AI

For decades, static solution concepts (Nash, Correlated, and Coarse Correlated Equilibria) and the Price of Anarchy (PoA) have formed the bedrock of algorithmic game theory, with no-regret learning proving fast convergence to such game-theoretic equilibria. We show that reducing multi-agent learning to static equilibri…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

2026-07-13 · Yunhai Feng, Natalie Leung, Jiaxuan Wang, Lujie Yang, Haozhi Qi, Preston Culbertson

General AI

Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves comp…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP

2026-07-13 · Michael Rizvi-Martel, Satwik Bhattamishra, Guillaume Rabusseau, Michael Hahn

General AI

A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theore…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

2026-07-13 · Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song, Xiuying Chen

General AI

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operatio…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

2026-07-13 · Tiberiu Musat, Tiago Pimentel, Nicholas Zucchet, Thomas Hofmann

General AI

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Mixture of Frames Policy: Multi-Frame Action Denoising for Bimanual Mobile Manipulation

2026-07-13 · Dian Wang, Jisang Park, Xiaomeng Xu, Han Zhang, Shuran Song, Jeannette Bohg

General AI

Robotic manipulation is inherently multi-frame: local actions may be simple in an end-effector frame, while transport, upright-object handling, and whole-body coordination are better represented in a base-aligned frame. However, modern diffusion-based visuomotor policies typically commit to a single predefined action f…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

2026-07-11 · Mustafa Ozan Duman, Ahmet Emir Dirik

General AI

Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. This study presents a comprehensive benchmarking of three architectural frameworks: Frozen Self-Supervised Learning (SSL-FRZ), Fine-Tune…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space

2026-07-13 · Mingchao Sun, Luyang Tang, Yu Liu, Xu Yan, Zhan Li, Yunwei Zhang, Fei Yu, Zengye Ge, Yumin Liu, Jiacheng Zhang, Yongchang Zhang, Jiawei Zhang, Zhicheng Liu, Zhongxu Sun, Tianjian Ouyang, Wenzheng Chen, Shixing Yang, Nianfei Fan, Guodong Sun, Huan Li, Zheng Zhou, Yongze Li, Yingliang Peng, Mengmeng Du, Yuan Liu, Haozhe Shi, Chunnuo Gong, Chengzhen Yu, Chunxue Jia, Yang Liu, Shiying Zeng, Junnan Lai, Hang Zhang, Ning Guo, Baoquan Chen, Mu Xu, Hongyu Pan

General AI

We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds. At the heart of our framework is a unified Spatial Generative Primitive (SGP), a compact tuple of a high-quality panorama and a spatial point cloud that delivers an efficie…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.0

NeuroCogMap Reveals Cognitive Organization of Large Language Models

2026-07-01 · Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen

General AI

Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible …

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.2

A Fokker-Planck approach to a stochastic multiplicative wealth model with taxation and redistribution

2026-07-13 · Iago Nascimento Barros, Marcelo Lobato Martins, Celia Anteneodo

General AI

We develop a Fokker-Planck description of the dynamics of wealth distribution in a stochastic multiplicative economic growth model with taxation and redistribution, as introduced by P.M.C. de Oliveira. Extending the original formulation, our theoretical framework includes general redistribution protocols, encompassing …

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.2

Supporting Reflection in LLM-based Exploratory Search

2026-07-13 · Giulia Di Fede, Salvatore Andolina

General AI

Large Language Models (LLMs) can make exploratory search more efficient but may undermine the reflection and iterative sensemaking needed in unfamiliar domains. Existing LLM tools often prioritize rapid answers over supporting users in tracking how their understanding evolves and how well their strategies align with th…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.0

4D Human-Scene Reconstruction from Low-Overlap Captures

2026-07-10 · Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park

General AI

Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overl…

Review
pending
Role
unreviewed
Read
later