Research Paper Cockpit

Daily Digest - 2026-09-26

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-10-04.

Papers

60 visible entries

arxiv Score 21.4

Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

2026-09-23 · Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney

Research Track A · General AI

Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific instruction following (M…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.2

GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI

2026-09-24 · Arunabh Srivastava, Mohammad A., Khojastepour, Srimat Chakradhar, Sennur Ulukus

General AI

Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing $\textbf{GRASP}$, a strategy-aware, multi-stage planning framework. GRASP d…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.4

Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

2026-09-23 · Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni

Research Track B · General AI

Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 steps derived from succ…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

A Living Benchmark for Information Retrieval from Electronic Health Records

2026-09-24 · Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer

General AI

Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, …

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

Multimodal Thinking with Renderable Programs

2026-09-24 · Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan

General AI

Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domai…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

2026-09-24 · Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li

General AI

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-a…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.4

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

2026-09-24 · Xingyu Wu, Yuchen Yan, Zhengxi Lu, Siqi Chen, Xin ZHANG, Aiting Liu, Chao Deng, Jie Liu, Jin Ma, Jian Shao, Jun Xiao, Yongliang Shen

General AI

Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories int…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.2

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

2026-09-22 · Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu, Xitong Gao

Research Track B · General AI

Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for ta…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.2

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

2026-09-24 · Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Conceição Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath

General AI

Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools rel…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.9

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

2026-09-24 · Linghua Zhang

Research Track B · General AI

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

2026-09-24 · Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri

General AI

Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can verify claims directly from speech. We introduce VeriSpeak, a probe benchmark for studying speech-based fact verification…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.4

OmniEcho: Spatial Audio Understanding for Embodied Agents

2026-09-20 · Ruixun Liu, Yuxuan Wang, Jiacheng Xie, Yuhuan You, Donghua Cai, Junming Lin, Xiong-Hui Chen, Zhifang Guo, Yunfei Chu, Qize Yang, Xize Cheng, Jin Xu, Yiwu Zhong

General AI

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks

2026-09-24 · Xueshu Chen, Yan Wang, Zihao Xue, Jiefu Li, Zhenfang Liu, Jayden Chen, Zhen Bi, Jungang Lou

General AI

Long-horizon tasks require preserving and later recovering cross-session evidence under a bounded, query-blind memory budget. Existing compression can discard fine-grained visual cues or conflate semantically similar but incompatible observations. We present C3M, a cross-session multimodal memory organization that main…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Coding Agents for Generalized Task and Motion Planning Problems

2026-09-24 · Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, Tom Silver

General AI

Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce plann…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

2026-09-24 · David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Anselm Paulus, Ameya Prabhu, Maksym Andriushchenko

General AI

A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce EvasionBench, a benchmark of 50 diverse task-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.4

Brain-Inspired Hierarchical Modularity for General Continual Learning

2026-09-21 · Hongwei Yan, Kanglei Zhou, Qi Cheng, Weiyi Dong, Chunyan Lan, Guanglong Sun, Jun Zhou, Qian Li, Yi Zhong, Liyuan Wang

Research Track A · General AI

Continual learning, the ability to learn from sequential experience while retaining and adapting prior knowledge, is central to intelligent systems operating in changing environments. However, conventional continual learning is typically studied with offline task-wise training and clear task boundaries, leaving a subst…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

2026-09-24 · Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao

General AI

Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.4

Exemplar-Free Analytic Learning for Multi-Label Audio Class-Incremental Learning

2026-09-24 · Siyuan Luo, Yang Xiao, Ting Dang

Research Track A

Audio classification is inherently a multi-label task, as real-world acoustic environments contain multiple simultaneous sound events. When new sound classes emerge, models must incorporate them without forgetting previously learned ones: a challenge known as class-incremental learning. Existing methods rely on storing…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

2026-09-24 · Beomsoo Kim, Byeongju Kim, Dohyun Kim, Dongwon Kim, Eunchong Kim, Hongmin Kim, Hyeojung Im, Hyeonbin Hwang, Hyeonghwan Kim, Hyoseok Seol, Insub Im, Irene Chen, Jaeseung Jeon, Jimin Hong, Kiyoon Yoo, Minkyoung Park, Seohyeon Jung, Seungjun Chung, Sue Hyun Park, Sungwoo Kim, Youngin Cho, Yujeong Son, Kangwook Lee, Hyunseung Kim

General AI

We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency const…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.2

Faster Visuomotor Policy Learning on Action Manifolds via Riemannian MeanFlow

2026-09-24 · S. Talha Bukhari, Austin Garrett, Yi Wei, Ruiqi Ni, Zachary Kingston, Aniket Bera

General AI

Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector f…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.2

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

2026-09-24 · Nayoung Choi, Shengjian Chen, Xiaokai Wei, Wenzheng Zhang, Daiyao Yi, Rachit Pareek, Vincent Su, Michelle Gong, Jinho D. Choi

General AI

Query understanding (QU) plays a critical role in production search systems, translating raw user queries into search execution plans that drive downstream retrieval and ranking. While large language models (LLMs) have enabled QU to be framed as a structured multi-task generation problem (e.g., intent classification, q…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.0

AgentKernel: The Trust-Native Agentic Operating System

2026-08-29 · Zhenhua Zou, Sheng Guo, Qiuyang Zhan, Lepeng Zhao, Shuo Li, Zhuotao Liu

General AI

Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actio…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

Agentic Detection of Online Conspiracies

2026-09-24 · Lior Biton, Oren Tsur

General AI

Conspiratorial discourse on social media is not always expressed through explicit claims or stable lexical markers. The same surface content may express endorsement, legitimate concerns, criticism, satire, or mockery. The main challenge is therefore not only recognizing conspiracy-related claims, but inferring the spea…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

RAPID: Robot Agentic Programming from Demonstrations

2026-09-24 · Yuyao Liu, Jiayuan Mao, David Hsu, Leslie Pack Kaelbling, Tomás Lozano-Pérez

General AI

Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstrat…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

2026-09-24 · Xinyue Zeng, Jiawei Zhang, Yujun Yan, Dawei Zhou

General AI

Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.4

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

2026-09-22 · Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers

Research Track A · General AI

Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Ter…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

Hyperbolic Multimodal Continual Learning: A Closest-Admissible Solution

2026-09-24 · Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King

Research Track A · General AI

Existing continual-learning methods protect parameters, replayed examples, or Euclidean feature subspaces. When applied to hyperbolic multimodal models, they do not explicitly preserve the Lorentz geometry that jointly encodes within-modality similarity, cross-modal correspondence, and semantic hierarchy; sequential up…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.4

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

2026-09-24 · Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan

General AI

In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-gra…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints

2026-09-24 · Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi, Leslie Barrett, Madhavan Seshadri, Enrico Santus

General AI

U.S. employment-discrimination complaints describe complex event sequences that are not explicitly captured by lexical or embedding-based representations alone. We present ARGUS, a source-grounded pipeline that combines a 5W1H-inspired schema, legal-domain models, and LLM-based structured generation to construct docume…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

IronViT: Toward Efficient Generalist Visual Representation Learning

2026-09-24 · Jiaxi Huang, Yueqi Hu, Xin Zhu, Xiaopeng Zhang, Huiting Qiao, Yanglin Zhang, Zefeng Ji, Rongxue Li, Yifei Xu, Huiying Yu, Wei Liu, Jiayin Zheng, Yinggan Xu, Peipeng Chen, Yin Zhang, Jian Yao

General AI

A generalist vision encoder must capture semantic, spatial, language-aligned, and action-relevant cues within a unified representation, yet softmax attention underlying today's most capable visual backbones becomes prohibitively expensive at high resolution. A natural attempt to address both challenges is to distill mu…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority

2026-09-24 · Mehmet Iscan

General AI

An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

2026-09-24 · Sudip Bhujel, Shanghao Shi, Ruiquan Huang, Ning Zhang, Yang Xiao

General AI

Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive E…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

The Alignment Illusion in Multimodal Large Language Models

2026-09-24 · Hong-Han Wang, Yuntao Wang, Hu Ding

General AI

Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interact…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.4

MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference

2026-09-23 · Romain Facq, Sami Ben Ali, Olivier Sentieys

Research Track A · General AI

Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. However, applying these methods efficiently in convolutional layers is not straightforward. A naive approach transfers full-precision weights and activati…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.4

Rufus-Air: An Open LLM Post-Training Recipe

2026-09-24 · Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Tuo Zhao

General AI

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and s…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.4

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

2026-09-24 · Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou, Huaizu Jiang

General AI

Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained teacher and an online critic to estimate distributional discrepancies through diffusion scores. In t…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.4

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

2026-09-24 · Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

General AI

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfol…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.2

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

2026-09-24 · Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno

General AI

Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independ…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.2

NEUROTESTGEN: Neuro-Symbolic Guided Test Generation with Large Language Models

2026-09-24 · Ruixin Zhang, Jiho Shin, Hung Viet Pham, Song Wang

General AI

Ensuring high structural coverage remains a fundamental challenge in automated test generation, particularly for complex software systems where reaching specific lines or branches requires satisfying intricate control- and data-flow constraints. Large Language Models (LLMs) have recently demonstrated strong capabilitie…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.2

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

2026-09-24 · Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang

General AI

Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.9

Automatic Harness Evolution for Hardware Design Verification: Can LLMs Consolidate Gains Across Discovered Harnesses?

2026-09-24 · Kidus Seyoum, Ajay Mittur

Research Track A · General AI

Agent behavior depends on the harness surrounding a language model, but it remains unclear whether language models can reliably improve such harnesses for hardware-design tasks. We study automatic harness evolution around a fixed subject model on 12 proprietary design-verification root-cause localization tasks. Across …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.4

An Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics

2026-09-23 · Teeratham Vitchutripop, Alyssa Quarles, Wenhe Zhang, Richard Xue, Daniel Rakita

Research Track A · General AI

Over the course of a lifetime, robots may encounter novel scenarios unaccounted for in its original training that result in performance degradation. One common approach to mitigating this issue is to further grow the offline training dataset in hopes of producing a policy robust to these changes. In contrast, biologica…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

Rolling-WAM: World Action Models with Rolling Imagination

2026-09-24 · Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

General AI

World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formula…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.4

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

2026-09-24 · Zhiyang Zhou, Yingxin Shang, Zhou Wang, Hongwei Cai, Weixu Wang, Shuran Zhou, Shuofeng Zhao, Wenke Fan, Qingxiang Guo, Dawei Yang, Lin Yang, Yang Song

Research Track A · General AI

Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which u…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.4

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

2026-09-24 · Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang

General AI

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement lear…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.7

Learning and interpreting policies for simultaneous entanglement requests in quantum networks

2026-09-24 · Leon Rode, Sumeet Khatri, Supartha Podder

General AI

Future quantum networks will make use of entanglement to perform numerous tasks, such as sending quantum information over long distances, distributed quantum computing, and quantum sensing. In general, these tasks will need to be performed simultaneously in various regions of a network, while minimizing resources and l…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.4

DeltaWAM: Delta World Action Models for Bimanual Manipulation

2026-09-23 · Han Yan, Zishang Xiang, Haokai Jiang, Zeyu Zhang, Qilin Wang, Weiyu Guo, Yandong Guo, Boxin Shi, Hao Tang

General AI

World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to n…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.4

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

2026-09-24 · Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan

General AI

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a syste…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Adapting a Large Language Model Crash-Severity Pipeline to Tennessee: Performance Across Sampling Strategies

2026-09-24 · Abhilasha Saroj, Pranav Govindu, Bharat Sharma, Usman Ahmed

General AI

State crash databases differ in structure, coding, and injury-severity distributions, limiting direct reuse of predictive workflows across jurisdictions. This study adapts the SafeTraffic Copilot large language model (LLM) crash-severity workflow to a three-year Tennessee inventory of 624,392 crashes. Tennessee crash, …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Jev in the Wild: A Data-Driven Analysis of the Jev Model's Functionality, Applications and Ecosystem

2026-09-24 · Guoming Ling, Muen Xue, Zijian Ye

General AI

Jev is a fast, low-cost decision model that answers natural-language questions with choices, binary judgments, and scores. As its public ecosystem grows rapidly, it remains unclear how Jev is used across applications and how public attention relates to project distribution. To answer these questions, we conduct a large…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

2026-09-24 · Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond

General AI

Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a pa…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.4

Learning to Discover Interesting Mathematics

2026-09-23 · Niket Patel, Ahmad Rammal, Amaury Hayat, Remi Munos, Julia Kempe

General AI

Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it re…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.4

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

2026-09-24 · Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina

General AI

While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the Superposition Lineari…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

OneTrans-V2: Unifying Retrieval, Pre-rank, and Fine-rank with One Transformer in Industrial Recommender

2026-09-23 · Hannan Cao, Jun Guo, Haolei Pei, Zhaoqi Zhang, Tianyu Wang, Ziyang Wang, Youchen Sun, Yue Xue, Yucheng Mao, Lintao Yan, Yufei Feng, Shaowei Liu, Rongkun Xing, Feiling Gong, Xinyu Chenli, Cong Xu, Mingge Zhang, Yunjia Zhu, Yajing Zhang, Pengfei Ren, Yue Lin

General AI

Industrial recommendation systems typically operate as a \emph{cascade} of retrieval, pre-rank, and fine-rank, but these stages are usually trained and served as separate models, causing repeated user-sequence encoding, isolated optimization, and duplicated engineering effort. Building on OneTrans' model-level unificat…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

2026-09-24 · Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao

General AI

Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-e…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization

2026-09-24 · Zebang Xie, Chuanyang Zheng, Yik-Chung Wu, Yihang Gao

General AI

Low-rank adaptation (LoRA) has become a popular parameter-efficient fine-tuning method for large language models. A key challenge in LoRA is how to determine the rank of each adaptation matrix, as rank directly controls its capacity and efficiency. Existing adaptive-rank methods typically allocate ranks according to ma…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection

2026-09-24 · Ben Liang, Chao Sui, Junqi Bai, Yuan Liu, Chunlai Li, Xiubao Sui, Qian Chen

General AI

In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enhancement, while the …

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

LLM Agents Can Easily Tamper With Their Own Traces

2026-09-24 · Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko

General AI

Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail t…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.2

Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS

2026-09-24 · Se Un Park, Hakjun Kim, Taehoon Roh, Junyoung Park

General AI

We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer initialized from English-trained weights attains 9.95 - 12.19% character error rate (CER) un…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.0

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

2026-09-19 · Chenyu Zhu, Ruoyu Zhao, Zhichao Lu

General AI

Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations rece…

Review
pending
Role
unreviewed
Read
later