Research Paper Cockpit

Daily Digest - 2026-08-04

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-09-22.

Papers

61 visible entries

huggingface Score 28.0

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

2026-08-03 · Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang

General AI

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for ope…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.5

Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning

2026-08-02 · MD Shaikh Rahman, Syed Maudud E Rabbi, Muhammad Mahbubur Rashid

Research Track A

Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali sentiment classification framework combining domain-adaptive continual pretraining and paramete…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.8

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

2026-08-03 · Michael Farmer

General AI

Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis generation requires an agent continuously coupled to the physical world. We defend a narrower claim: online embodiment is not necessary for every abductive scientific …

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.0

DiffusionGemma Technical Report

2026-07-31 · DiffusionGemma Team, Adrien Ali Taïga, James Assiene, Daniele Calandriello, Rahma Chaabouni, João Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, João Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, Çağlar Ünlü, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Clément Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor

General AI

We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional au…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience

2026-08-03 · Sitong Gong, Caixin Kang, Tianyu Yan, Guo Chen, Bo Zheng, Kaipeng Zhang, Yunzhi Zhuge, Xiang Ruan, Huchuan Lu, Yifei Huang

General AI

A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primarily support question-conditioned recall, whereas proactive assistants typically use separate memory and control mechanisms. We introduce GROV…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

2026-08-03 · Zhaoxin Yu, Qi Shen, Hengli Li, Zhaowei Zhang, Song-Chun Zhu, Chi Zhang, Zilong Zheng

General AI

Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assig…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

Self-Improving Large Language Models via Progressive Experience Evolution

2026-08-03 · Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng

General AI

Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract e…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.0

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

2026-08-03 · Ziyu Ma, Hailang Huang, Shun Zou, Yong Wang, Shidong Yang, Yiming Hu, Fei Wei, XiangXiang Chu

Research Track B · General AI

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

UEmbed: Unified Sparse and Dense Multimodal Embeddings

2026-08-03 · Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Zhijie Nie, Yilun Zhao, Shu Wu

General AI

Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

ACEM: A Cost Estimation Model for Agentic Software Engineering

2026-08-03 · Mohammad El-Ramly

General AI

Traditional software cost estimation models, such as COCOMO II, Function Points, and Story Points, assume that development effort is primarily driven by human labor in design, coding, and testing. Agentic software engineering, where autonomous AI agents perform substantial implementation work and humans focus on planni…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning

2026-08-03 · Yiran Gao, Tao Li, Kim Hammar

General AI

Incident response is currently managed by security operators using predefined playbooks, resulting in slow, labor-intensive security decision-making processes. Consequently, there is a growing need for automated incident response planning. Decision-theoretic approaches based on control, optimization, and reinforcement …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

2026-08-03 · Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Hu Wei, Bing Zhao

General AI

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invali…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

2026-08-03 · Kang Liu, Zijing Wang, Yongkang Liu, Mengjie Zhao, Xiaocui Yang, Shi Feng, Yifei Zhang, Daling Wang

General AI

Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as reasoning trajectories grow, models may become less effective at using information established earlier in the context, increasing the risk of reasoning errors. Existin…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.5

Plasticity of Growing and Elastic Neural Networks in Online Continual Learning

2026-08-02 · Jeong Min Kong, Richard S. Sutton

Research Track A

Neural networks that can grow or both grow and shrink during learning, referred to as growing neural networks and elastic neural networks, respectively, have recently been explored in offline continual learning with a particular focus on catastrophic forgetting. Driven by the observations that 1) online continual learn…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.5

Learning What to Remember: Test-Time Training via Context Distillation

2026-08-03 · Zixuan Wang, Xingyu Dang, Rui-Jie Zhu, Zixin Wen, Hengyu Fu, Wenhao Chai, Jason D. Lee

Research Track A · General AI

Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruc…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems

2026-07-31 · Yashaswi Malla, Sandra Siby

Research Track B · General AI

Large Language Model (LLM)-based web agents are increasingly evolving from single-agent systems (SAS) to multi-agent systems (MAS). While MAS can lead to improved task performance by decomposing complex tasks across specialized sub-agents, such role decomposition introduces new structural attack surfaces that are absen…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

2026-08-03 · Taye Akinrele, Sindhuja Penchala, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi

Research Track A · General AI

Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive capabilities across a wide range of tasks, many fundamental cognitive…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

2026-08-03 · Zhichen Liu, Ruihan Sun, Hengjie Yang, Zipeng Wu, Zhaohan Chen, Xiaofan Zhang, Yang Xu

Research Track A · General AI

Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the full lifecycle when working context changes. We formulate this missing inferenc…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.0

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection

2026-08-02 · Fei Li, Yue Yu, Yuran Wang, Xinghan Li, Jingjing Chen, Yu-Gang Jiang

Research Track A · General AI

AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail to generalize to newly emerging generative models, making this task challenging. Therefore, continual learning (CL) is essential for improving the adaptability. However,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

2026-08-03 · Brandon Wang, Andrei S. Tyrin, Daniil A. Boiko

General AI

Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific int…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.5

UCBound-Net: Uncertainty-Guided Boundary-Aware Continual Learning for Domain-Incremental Ultrasound Segmentation

2026-08-02 · Mohammad Amanour Rahman

Research Track A · General AI

Continual learning in clinical imaging faces a dual challenge: a model must assimilate knowledge from new anatomical domains while retaining representations learned from prior tasks, a problem known as catastrophic forgetting. Existing mitigation strategies, including regularization and knowledge distillation, treat al…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.0

Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

2026-08-03 · Anusha Madan Gopal, Aras Pirbadian, Kristofor D. Carlson, M Anthony Lewis, Jonathan Tapson

Research Track A · General AI

Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. State-Space Models (SSMs) avoid the second cost by construction; we eliminate the first, collapsing prefill from $O(L_{context})$ to…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

2026-07-31 · Chengbo Liu, Lifang Zhou, Ruijie Yan, Pei Tan, Ao Sun, Haojun Huang, Guichun Hua, Sining Wei, Yining Chen, Yingying He, Yutao Xie

Research Track B · General AI

Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Successful trajectories are expensive to collect and often contain inefficient detours. After supervised fine-tuning (SFT), full trajectory corpora are dominated by routine…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

2026-08-03 · Vishwajeet Shivaji Hogale, Anjali Pai, Nitya Ravi

General AI

Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remains unreliable for fine-grained, visually ambiguous targets. We study this gap in the context of automated vehicle damage assessment, where fine-grained defects such as …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

2026-08-03 · Ambarish Govindarajulu Kaliamurthi, Kaikai Liu

General AI

Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact model size and reliable metric grounding. We present MoRAL (Multimodal Reasoning for Autonomous Language Models), a two-stage fine-tuning pipeline that teaches Cosmos-…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.0

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

2026-08-02 · Changwoo Baek, Kyeongbo Kong

General AI

Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial reasoning for tasks such as 3D question answering. However, this design generates thousands of tokens per scene, resulting in substantial computational and memory overhead…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+

2026-08-03 · Christoph Weinhuber, Maximilian Prokop, Giuseppe De Giacomo, Moshe Y. Vardi

General AI

Linear Temporal Logic (LTL) is one of the most widely adopted languages for specifying temporal extended objectives in AI, with applications ranging from reactive synthesis to stochastic planning in Markov decision processes and reinforcement learning. Traditionally, solving any of these problems requires translating t…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents

2026-08-03 · YuFei Luo, Xiucheng Xu, Zhen Yang

General AI

Long-term memory is critical for LLM agents operating over long-horizon interactions. However, several persistent limitations of existing memory systems can be traced to two recurring misalignment patterns in long-term interaction settings: Temporal-Structural Misalignment (TSM) and Delayed Utility Manifestation (DUM).…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.5

KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement

2026-08-03 · Gusseppe Bravo-Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain

Research Track A · General AI

Data drift poses significant challenges for machine learning systems in production, requiring continuous model updates to maintain performance. We present KC-Agent, a dual-process cognitive architecture for automated ML model improvement that combines fast pattern recognition (System 1) with deliberate incremental upda…

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.0

Progressive Agent Skill Generation via Reinforcement Learning

2026-08-03 · Junhao Shen, Zhanqiu Zhang, Yiwen Guo, Hong Cheng

Research Track A · General AI

Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. In contrast, learning-based approaches offer a more unified way to model skill generation across heterogeneous sources. However, learning-based skill generation …

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.0

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

2026-08-03 · Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, Lei Bai

General AI

Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthe…

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.0

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

2026-08-03 · Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, Kang Liu

General AI

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents u…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

2026-07-31 · Chaimae Abouzahir, Musa Khan, Hala Ali-Hassan, Congbo Ma, Khaled Saleh, Yousra Sadqi, Jihad Mallat, Walid Al-Eisawi, Nizar Habash, Farah E. Shamout

General AI

Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermed…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

2026-08-03 · Qiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic

General AI

The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes months. Agentic AI promises to automate this search. Off the shelf, however, it fa…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

2026-08-03 · Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang

General AI

Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image su…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

2026-08-03 · Luan Zhang, Ruochen Zhou, Dandan Song, Zhengyu Chen, Yuhang Tian, Jun Yang, Huipeng Ma, Chenhao Li, Guangyuan Feng, Xudong Li, Yizhou Jin, Yan Xu

General AI

Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing method…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

2026-08-03 · Saman Sarker Joy, Niloy Farhan

General AI

Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than pressured patient-facing conversations. We introduce MedPRESS, a multi-turn benchmark for measuring patient-pressure-induced sycophancy in LLMs. MedPRESS contains 600 m…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes

2026-08-03 · Nan Chen, Zhouhao Yang, Soufiane Hayou

General AI

Intent classification in Large Language Models (LLMs) involves categorizing user prompts into predefined classes. For instance, given a user prompt, the system must determine whether it primarily concerns mathematics, coding, or general text processing. Such classification enables routing prompts to specialized models …

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.5

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning

2026-08-01 · Malavika Suresh, Ikechukwu Nkisi-Orji, Nirmalie Wiratunga

Research Track A

Achieving continual learning (CL) with deep neural networks requires balancing stability and plasticity while enabling knowledge transfer. In this work, we focus on offline learning algorithms under the constraints: (I) no access to training data from prior tasks (II) no access to task-id at inference time. We introduc…

Review
pending
Role
unreviewed
Read
now
huggingface Score 9.0

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

2026-07-31 · Haoran Ling, Yuecheng Li, Zeyu Song, Jing Yao, Shuwen Kang, Chi Lu, Wenjin Wu, Peng Jiang

General AI

Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to both select modification directions and generate concrete hypotheses often leads …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.0

AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning

2026-08-02 · Pengxiang Wang, Hongbo Bo, Jun Hong, Weiru Liu, Kedian Mu

Research Track A

Catastrophic forgetting is a major problem in task-incremental learning, where neural networks tend to overwrite previously learned knowledge when trained on new tasks. A number of architecture-based approaches have been proposed to address this problem. However, the architecture-based approaches suffer from another pr…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework

2026-08-03 · Junjie Yin, Buxin She, Xinyu Feng, Fangxing, Li

General AI

Artificial intelligence (AI) is increasingly central to power and energy systems, supporting modeling, forecasting, optimization, and control. Yet most existing works emphasize specialized applications and offer little reusable material for newcomers or interdisciplinary learners, who increasingly rely on large languag…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.0

A manifold-aware Neural ODE surrogate model for stochastic induction heating with anisotropic electrical conductivity

2026-08-03 · Wouter J. Schuttert, Mohammed Iqbal Abdul Rasheed, Bojana Rosić

Research Track A

Induction welding plays a central role in enabling lightweight, integrated structures made from fibre-reinforced thermoplastic composites. From a modelling perspective, the induction welding process can be approximated by one-way coupled electromagnetic and heat-transfer equations. In practice, material parameters such…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Implicit Neural Representations for Multimodal Longitudinal Image Imputation and Interpolation

2026-08-03 · Sina Wendrich, Lukas Förner, Zoe Reinke, Kartikay Tehlan, Ansgar Berlis, Michael Frühwald, Matthias Wagner, Thomas Wendler

General AI

Longitudinal multiparametric MRI is central to follow-up imaging in oncology, yet real-world clinical data are characterised by missing sequences, heterogeneous acquisition protocols, and varying spatial resolutions across time points. We propose a patient-specific conditional implicit neural representation (INR) that …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment

2026-08-03 · Nan Bi, Taoyue Wang, Lijun Yin, Vandana Sharma

General AI

Automated pain assessment in real clinics is limited by scarce clinically grounded facial video data with weak labels (often sequence-level self-report) and by the fact that pain cues can be subtle or near-neutral in RGB, while thermal and depth signals are informative yet impractical to deploy routinely. To address th…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Unpaired Modality-Agnostic Generative Recommendation

2026-08-03 · Weihao Shen, Wei Chen, Fuwei Zhang, Meng Yuan, Yuqin Lan, Guojun Liu, Qingsong Hua, Wei Lin, Fuzhen Zhuang

General AI

Generative Recommendation (GR) formulates recommendation as autoregressive generation over discrete semantic identifiers (IDs). Although recent multimodal GR methods improve semantic ID construction with visual and textual information, they typically require item-level paired observations, restricting tokenization to t…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.0

CADENA: Stepwise CAD Reverse Engineering

2026-08-01 · Soslan Kabisov, Gennadiy Savrasov, Maksim Elistratov, Antonio Rodriguez, Daniil Ignatiev, Nikita Gavrilov, Rustam Uzdenov, Alexey I. Boyko, Igor Pasechnik, Anton Konushin, Andrey Kuznetsov, Dmitrii Zhemchuzhnikov

General AI

Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still demands substantial expert effort. Most AI systems emit the entire CAD program in a single pass, never inspecting the intermediate geometry. In contrast, human engineers build a part feature by feature, c…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.8

EulerLoRA: Rank-Driven Jump Dynamics for Calibrated Parameter-Efficient Fine-Tuning

2026-08-02 · Srinivas Anumasa, Dianbo Liu

General AI

Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning, but standard LoRA produces a single deterministic model and does not directly support predictive uncertainty estimation. We introduce EulerLoRA, a stochastic extension of LoRA that generates multiple predictive trajectories by sampling structured varia…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.8

ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation

2026-08-02 · Mohamed Farag, Genc Hoxha, Yahia Maleki, Chris McCool, Ribana Roscher

General AI

Reliable decision-support in digital agriculture requires accurate predictions and well-calibrated uncertainty estimates, particularly for dense prediction tasks such as semantic segmentation. Ensemble methods provide strong uncertainty quantification, but their computational and memory demands limit practical use, whi…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.8

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

2026-08-03 · Jiajun Liang, Yucheng Liao, Yukang Cao, Jiazhe Wei, Ken Li, Wende Tan, Jiankun Zhang, ZY Cui, Jingkang Yang, Liucheng Guo, Shiqi Yang, B. Yang, Caifeng Shan, Ziwei Liu, Chenyang Si

General AI

Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces not designed for joint generation and decoding, or c…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.8

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

2026-08-03 · Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang

General AI

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to ge…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 4.8

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

2026-08-03 · Natalie Isak, Matthew Dressman

General AI

The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-o…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 4.8

The Condition-Number Barrier in Sparse Least Squares

2026-08-03 · Honghao Lin, Vahab Mirrokni, David P. Woodruff

General AI

In [AS21], Axiotis and Sviridenko conjectured that the linear dependence on the restricted condition number in sparse convex optimization cannot be improved by a polynomial-time algorithm. We establish their conjectured lower bound for least-squares objectives, conditional on the randomized exact-volume Small-Set Expan…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 4.0

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

2026-07-29 · Rongxiang Zhang, Songhua Liu

General AI

Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel f…

Review
pending
Role
unreviewed
Read
later
huggingface Score 4.0

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

2026-08-01 · Tongsheng Ding, Zhen Luo, Yixuan Yang, Boyu Wang, Luyang Xie, Jinyu Yang, Feng Zheng

General AI

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predictors either rely on privileged inputs such as video, depth, or CAD models, or rec…

Review
pending
Role
unreviewed
Read
later
huggingface Score 4.0

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

2026-08-01 · Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan

General AI

Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image and text load errors can cancel at one …

Review
pending
Role
unreviewed
Read
later
arxiv Score 3.8

On Defining Chart Types Boundaries

2026-08-03 · Chang Han, Andrew Mcnutt, Katherine E. Isaacs

General AI

What makes a Gantt chart? This question proved unexpectedly difficult to answer when we set out to build a design space for Gantt charts. Existing definitions, each shaped by their respective research goals, made different scope choices that we could not directly reconcile. We reasoned about what should and should not …

Review
pending
Role
unreviewed
Read
later
arxiv Score 3.8

Pairwise-Independent Dithering for Single-Stage Hadamard Quantization

2026-08-03 · Honghao Lin, Vahab Mirrokni, David P. Woodruff

General AI

Quantizing high-dimensional vectors is fundamental to similarity search, distributed learning, and model compression. Feng, Indyk, Kapralov, Krachun, and Prokhorov established sharp guarantees for an unbiased dithered quantizer based on a randomized Hadamard transform [FIK+26]. Their $1/d$-scale inner-product estimator…

Review
pending
Role
unreviewed
Read
later
arxiv Score 3.8

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

2026-08-03 · Haijie Yang, Jindi Bao, Yixuan Dong, Hongliang Zhang, Jian Bi, Hao Tang, Zhenyu Zhang, Jianjun Qian, Jian Yang

General AI

Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion control remains challenging. Existing methods often suffer from insufficient emotion-aware motion priors and expensive appearance computation during multi-step denoising…

Review
pending
Role
unreviewed
Read
later
arxiv Score 3.8

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification

2026-08-03 · Chao Ji, Shiyu Xuan, Zechao Li

General AI

Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric deformation. Existing methods attempt to learn view-invariant representations exclusively within the 2D image space, where drastic viewpoint variations cause the learned fe…

Review
pending
Role
unreviewed
Read
later
arxiv Score 3.8

WIP: Chat-Debugging: Large Language Model as a Hardware Debugging Assistant

2026-08-03 · Andrew Ash, John Hu

General AI

This work-in-progress research paper explores Chat-Debugging, a novel use case for large language models as an assistant for hardware debugging tasks to improve students' debugging skills. Hardware debugging can be a time-consuming and stressful skill to develop, leading to frustration and other negative emotions. Whil…

Review
pending
Role
unreviewed
Read
later