Research Paper Cockpit

Daily Digest - 2026-07-21

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-07-25.

Papers

55 visible entries

arxiv Score 29.4

Rethinking Quantum Continual Learning with Quantum Fisher Information

2026-07-17 · Yu-Chao Hsu, Yu-Cheng Lin, Tai-Yue Li, Nan-Yow Chen, En-Jui Kuo

Research Track A

Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variational quantum classifiers (VQCs) are prone to catastrophic forgetting under nonstationary task distributions. We propose quantum elastic weight consolidation (QEWC), a quantum Fisher i…

Review
pending
Role
unreviewed
Read
now
arxiv Score 23.8

Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

2026-07-20 · Haochen Zhao, Yongxiu Xu, Xinkui Lin, Dong Xie, Jiarui Lu, Yuqi Qian, Yubin Wang, Hongbo Xu, Gaopeng Gou

General AI

Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may d…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.0

A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive Manufacturing

2026-07-20 · Yi-Ping Chen, Ying-Kuan Tsai, Vispi Karkaria, Seul Lee, Daniel Apley, Wei Chen

Research Track A

Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as operating conditions evolve, a phenomenon known as concept drift. Maintaining surrogate fidelity under drift, particularly when models must also capture aleatoric uncertainty, remains an open challenge. Exist…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.8

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

2026-07-20 · Keuntae Kim, Beomseok Lee, Hyunwoo Kim, Yong Suk Choi

General AI

Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correction. Diffusion Multimodal Large Language Models (dMLLMs) unmask tokens in an order-agnostic process, improving efficiency and enabling itera…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization

2026-07-20 · Rui Chu, Yingjie Lao

General AI

Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Multimodal Large Language Model(M-LLM) has shown its capability in video understanding. Video summarization, a specific domain of video understanding, has proven its importance for effi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory

2026-07-20 · Qingcan Kang, Mingyang Liu, Shixiong Kai, Kaichao Liang, Zhentao Tang, Yuqi Cui, Tao Zhong, Mingxuan Yuan

Research Track A · General AI

Language agents depend on memory across interactions. However, the limited context windows of large language models (LLMs) and their inference costs constrain how much memory can be used at once. Existing systems mainly follow two strategies: memory retention and memory consolidation. Retention keeps raw records and pr…

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.0

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

2026-07-19 · Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang

General AI

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoi…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.4

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

2026-07-14 · Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu

General AI

Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their m…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.9

Rethinking Transfer in Continual Learning: A Replay-Based Realisation

2026-07-17 · Yang Meng, Zhenya Liu, Zhuokai Zhao, Yuxin Chen

Research Track A

Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch. Existing methods, whether rehearsal-based (replaying stored past data) or rehearsal-free (regularising or isolating parameters), overwhelmingly target one objective: preventing catastroph…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

2026-07-18 · Yang Liu, Weixing Chen, Xinshuai Song, Tao Pu, Siwen Mo, Yongjie Bai, Zihao Chen, Qianran Sun, Liruo Zhong, Ying Shen, Liang Lin

General AI

Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, semantic verification, and persistent experience across heterogeneous embodiments. We present PhyAgentOS, a runtime foundation delivering schedu…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions

2026-07-20 · Jingzhe Fang, Guozhi Xu, Yunfan Cui, Xiaochen Yang, Zhangyu Hua

General AI

AI companions are judged not only by single-turn fluency but by whether they sustain emotional continuity: remembering who the companion is, what the user prefers, and how the relationship has felt. We present ZifaMem, a structured memory system that organizes dialogue into session summaries, episodic memories, and a c…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.0

TypiCore: A Hybrid Active Query Strategy for Class-Incremental Learning on Time Series

2026-07-20 · Gabor Szucs, Samuel Jacsev, Marcell Nemeth, Davide Dalle Pezze, Gian Antonio Susto

Research Track A · General AI

Time series data play a pivotal role across numerous domains, including healthcare and manufacturing. In real-world environments, models must cope with distribution shifts over time, a challenge commonly addressed through Continual Learning (CL) techniques. However, existing CL methods face a critical limitation: real-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

SGA: Plug&Play Geometric Verification for Educational Video Synthesis

2026-07-20 · Lopez Jhon, Hinojosa Carlos, Ghanem Bernard

General AI

Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. However, ensuring spatial correctness and visual legibility remains challenging, as existing frameworks emphasize pedagogical content while overlooking geometric occlusions. We propos…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.0

Distilled Reinforcement Learning for LLM Post-training

2026-07-19 · Chen Wang, Zhaochun Li, Jionghao Bai, Yining Zhang, Hexuan Deng, Ge Lan, Yue Wang

General AI

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and lim…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.0

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

2026-07-20 · Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang

General AI

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluatio…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

2026-07-20 · Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen

General AI

Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving systems and auto-parallelism compilers comm…

Review
pending
Role
unreviewed
Read
now
huggingface Score 13.0

Group Entropy-Controlled Policy Optimization

2026-07-18 · Guangran Cheng, Chengqi Lyu, Songyang Gao, Wenwei Zhang, Kai Chen

General AI

Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce distinct entropy regimes under the same policy, m…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

2026-07-20 · Thomas MacDougall, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Maxim Malkov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov

General AI

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molec…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications

2026-07-20 · Daniela Rojas, Abdulwahab Albassam, Aidan G. Leung, Jett Ngo, Ryan Luo, Peter R. Quawas, Junpyung Kim, Kangkai Liang, Mansi Nanavati, Jonathan Mai, Meng-Chi Tsai, Yun-Tong Tsai, Yize Chen, Yuanyuan Shi

General AI

Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and act in technical domains. In smart grids, recent work applies agentic schemes to forecasting, optimization, and control, wrapping trusted solvers behind language interfaces and orc…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

2026-07-20 · Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu

General AI

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performan…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs

2026-07-20 · Huzaifa Shaaban Kabakibo, Eric Schniedermeyer, Artem Burchanow, Lin Wang

General AI

Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of Natural Language Processing (NLP) tasks, but their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Existing approaches to model compression and optimization oft…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.4

Recursive Harness Self-Improvement

2026-07-17 · Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang

Research Track A · General AI

Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future …

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

2026-07-19 · Qing Zong, Yue Guo, Mengxin Yang, Yiwen Guo, Yangqiu Song

General AI

This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over ti…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

2026-07-20 · Kehan Li, Bohan Hou, Minghao Zhu, Tianyi Zhang, Zesen Cheng, Zhikai Wang, Sicong Leng, Xin Li, Xiao Lin, Biying Yao, Minghua Zeng, Jiangpin Liu, Ronghao Dang, Jiayan Guo, Siteng Huang, Haoyu Zhao, Heng Ping, Yaxi Zhao, Kexiang Wang, Tong Lu, Shengke Xue, Jiahao Tang, Yulei Wang, Zejing Wang, Jianwei Gao, Shijian Lu, Chengju Liu, Jianfei Yang, Mingxiu Chen, Deli Zhao

General AI

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces …

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?

2026-07-20 · Yimeng Chen, Nathanaël Denis, Roberto Di Pietro, Jürgen Schmidhuber

General AI

Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. We refer to this class of threats as self-state attacks. In this paper, we investigate the OS resilie…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

2026-07-20 · Ruicheng Li, Qixiu Li, Ruichun Ma, Yu Deng, Lin Luo, Zhiying Du, Jianfeng Xiang, Huizhi Liang, Ruicheng Wang, Jiaolong Yang, Baining Guo

General AI

Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

2026-07-20 · Kevin Du, Clara Kümpel, Michelle Wastl, Alex Warstadt

General AI

Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true. In others, they should stick to their prior knowledge. Notably, users' expressions of belief (EoBs) can take linguistically diverse forms - using presuppositions, evidentia…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

2026-07-20 · Hang Zhang, Warren J. Gross

General AI

Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark

2026-07-20 · Sania Bano, Shahzad Ahmad, Santosh Kumar Vipparthi, Sukalpa Chanda, Subrahmanyam Murala

General AI

Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can in…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

2026-07-20 · Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros, Richard Zhang

General AI

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a l…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval

2026-07-20 · Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng, An-Zi Yen

General AI

Large language models are increasingly used in practical systems, making efficient model selection important for reducing deployment cost. LLM routing has emerged as a practical solution for allocating each input query to an appropriate model under a desired cost-performance trade-off. Existing routing methods often es…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.4

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

2026-07-17 · Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu

General AI

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We in…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

How to Build Marcus's Algebraic Mind: From Minsky's Emotion-Machine Viewpoint

2026-07-18 · Hiroyuki Chuma, Kanji Otsuka, Yoichi Sato

Research Track A · General AI

In The Algebraic Mind, Marcus identified three cognitive components: operations over variables, recursively structured representations, and an individual/kind distinction. He left the neural substrate open. A companion paper solves this with VaCoAl, an architecture built on GF(2) XOR-and-shift. It offers exact reversib…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.8

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

2026-07-20 · Sheldon Yu, Tong Yu, Xunyi Jiang, Rohan Surana, Gagan Mundada, Sungchul Kim, Lina Yao, Julian McAuley, Junda Wu

General AI

Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches and operate at the input level, offering no fine-grained control over the reasoning proc…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

2026-07-20 · Brian K Chen

General AI

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical fo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation

2026-07-20 · Christina Nasika, Feng Wang, Antonis Krasakis, Han Xiao

General AI

Listwise rerankers are the discriminative core of agentic retrieval pipelines, yet production deployment demands efficiency, domain robustness, and fluency on semi-structured data at the same time. We present jina-reranker-v3.5, a 0.6B-parameter listwise reranker that meets these demands together without sacrificing th…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.5

Hierarchy-Aware and Anatomy-Guided Learning for Lung Ultrasound Video Classification

2026-07-20 · Alya Almsouti, Lotfi Mecharbat, Noha Aboukhater, Yousef Alabrach, Siddiq Anwar, Andre Kumar, Ibrahim Almakky, Mohammad Yaqub

Research Track A · General AI

Lung ultrasound (LUS) is a bedside tool for assessing pulmonary edema in patients at risk due to heart failure or impaired kidney function. However, automated LUS analysis remains challenging because of speckle noise, imaging artifacts, and operator-dependent acquisition variability. In this work, we present a deep lea…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

Beyond Storage: State as a Runtime Control Problem in Parallel and Distributed Systems

2026-07-18 · Shuhao Zhang, Haoran Peng, Chun Ho Ma, Yancan Mao, Yufeng Du, Shifeng Liu, Ruijie Qiu, Xiaofei Liao, Hai Jin

General AI

Shared state increasingly shapes both performance and failure behavior in streaming, serving, retrieval, and continual-learning systems. Existing studies, however, often isolate access control, hardware-aware execution, memory management, and long-horizon updates. The review organizes this literature around three coupl…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning

2026-07-18 · Sana Tonekaboni, Viktoria Schuster, Caroline Uhler

General AI

Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities. However, training multimodal models faces two main obstacles. First, collecting large-scale, well-aligned paired multimodal datasets is often impractical, making end-to-end multimodal training diffi…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer

2026-07-20 · Nikhil Ghosh, Tetiana Parshakova, Robert M. Gower

General AI

Low-rank adaptation (LoRA) makes finetuning large language models cheaper by adding to each weight matrix a trainable low-rank update parameterized as the product of two matrices. These matrices are usually trained with Adam, which treats them as a single flat vector of parameters and ignores both the matrix and produc…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

2026-07-20 · Yuhang Wang, Yuling Shi, Shaoqiu Zhang, Jialiang Liang, Shilin He, Siyu Ye, Yuting Chen, Kai Cai, Xiaodong Gu

General AI

Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context when rea…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.9

ImprovedVBGS: Real-time Continual Variational Bayes Gaussian Splatting

2026-07-17 · Damani Mguni-Coker

Research Track A · General AI

On-the-fly reconstruction is a key requirement for many applications in robotics and autonomous navigation. Variational Bayes Gaussian Splatting (VBGS) enables continual learning without replay buffers using Coordinate Ascent Variational Inference (CAVI), but its per-frame iterations over all observed points make it to…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

OR Else: A Differentiable Trust Region for Policy Optimization

2026-07-20 · Chinmay Rane, Kanishka Tyagi, Michael Manry

General AI

PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in the scalar objective's derivative. We ask whether Output Reset (OR), a smooth one-sided saturation rule, offers a useful alternative for large language model post-training. PPO-OR …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference

2026-07-20 · Masahiro Kato, Taka Kato

General AI

We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedding space, the generator estimates condi…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach

2026-07-17 · Lichao Mou, Shilan Zhang, Chunlei Li, Bingcong Yan, Jingliang Hu, Yilei Shi, Shengwu Xiong, Xiao Xiang Zhu, Lei Li, Yaxiong Chen

General AI

Large vision-language models (LVLMs) can be adapted to specialized medical imaging tasks via parameter-efficient fine-tuning approaches such as low-rank adaptation (LoRA), leading to a growing ecosystem of expert models tailored to specific imaging modalities and clinical scenarios. However, deploying multiple expert L…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.8

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

2026-07-20 · Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo

General AI

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Robust Multimodal Dynamic Object Segmentation

2026-07-20 · Zhe Xin, Hanzhi Chang, Penghui Huang, Yinian Mao, Guoquan Huang

General AI

Dynamic object segmentation plays a critical role in many visual applications such as static scene reconstruction from dynamic videos. However, existing optical flow-based methods fail to ensure consistent static/dynamic segmentation along object boundaries, while 3D reconstruction-based approaches are highly sensitive…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation

2026-07-20 · Hanchen Yang, Kaiwen Yang, Junpeng Zhuang, Yang He, Keting Cen, Bochao Liu, Zhongbo Sun, An Liu, Zhongteng Han, Chenyi Lei

General AI

User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment ev…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization

2026-07-20 · Alex Mathai, Shobini Iyer, Aleksandr Nogikh, Petros Maniatis, Franjo Ivancic, Junfeng Yang, Baishakhi Ray

General AI

Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, building applications, and prototyping. However, despite their value as coding assistants, agent-generated code tends to be larger and more verbose than the corresponding human-written implementation. In thi…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.0

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

2026-07-19 · Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu

General AI

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Cross-Coordinate Correspondence Pruning for Image-to-Point Cloud Registration

2026-07-19 · Xin Liu, Rong Qin, Huipeng Lin, Leizhi Shu, Jin Wu, Chi-Man Vong, Liang Lin, Jufeng Yang

General AI

Recent detection-free approaches have shown significant efficacy in image-to-point cloud (I2P) registration by employing a coarse-to-fine matching pipeline. In the coarse stage, down-sampled image features and voxelized point cloud features are typically fused to establish initial coarse correspondences for subsequent …

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

ThRIve: Thermally Robust CNN Inference via Low-Rank Adaptation in Heterogeneous PIM Architectures

2026-07-19 · Vibhanshu Sharma, Pratyush Dhingra, Janardhan Rao Doppa, Partha Pratim Pande

General AI

Processing-In-Memory (PIM) has emerged as a promising technology for accelerating machine learning (ML) workloads. Specifically, non-volatile memory-based PIM architectures have enabled effective ML acceleration due to their ability to perform energy-efficient matrix-vector multiplication operations. However, these dev…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Automated Discovery Has No Universally Superior Harness

2026-07-20 · Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen

General AI

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive …

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Nonexistence of Simultaneously EF1 and Pareto Optimal Allocations for Submodular Valuations

2026-07-20 · Harish Chandramouleeswaran, Prajakta Nimbhorkar

General AI

The existence of allocations of indivisible goods that are simultaneously fair (envy-free up to one item (EF1)) and efficient (Pareto optimal (PO)) when agents have monotone submodular valuations has been a longstanding open problem. We settle this question negatively by giving an example with two agents where no alloc…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Per Astronomix ad Astra: High-Order Differentiable (Magneto)hydrodynamics with Energy-Conserving Self-Gravity

2026-07-20 · Leonard Storcks, Nils Thuerey, Tobias Buck

General AI

We present astronomix, a performant differentiable (magneto)hydrodynamics simulator written in Python/JAX. We demonstrate how automatic differentiation, validated against hand-derived analytical functional derivatives and finite differences, enables inverse modeling over millions of parameters and allows for sensitivit…

Review
pending
Role
unreviewed
Read
later