Research Paper Cockpit

Daily Digest - 2026-07-17

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-07-25.

Papers

56 visible entries

arxiv Score 29.4

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

2026-07-13 · Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu

Research Track A · General AI

Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired k…

Review
pending
Role
unreviewed
Read
now
arxiv Score 27.4

An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering

2026-07-13 · Mai A. Shaaban, Tausifa Jan Saleem, Alaa Mohamed, Dilnaz Utemissova, Ufaq Khan, Mohammad Yaqub

Research Track A · General AI

Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tasks without forgetting previously acquired knowledge. Continual learning (CL) provides a practical framework for this setting. Despite rapid progress in medical vision-language model…

Review
pending
Role
unreviewed
Read
now
arxiv Score 27.4

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation

2026-07-16 · Yao He, Gan Sun, Wenqi Liang, Fazeng Li, Yang Cong

Research Track A

Similar to the natural capabilities of humans to sequentially learn new tasks, robots with Vision-Language-Action (VLA) models should possess lifelong learning ability to learn a new task when deployed in open-world environments. However, most recently proposed lifelong learning models aim to effectively learn the curr…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.4

Gate-Zero Growth: A Geometric Framework for Function-Preserving Continual Learning

2026-07-16 · Dante Lok

Research Track A

We introduce \emph{gate-zero growth}, a function-preserving (FP) operator for continual learning that adds new residual blocks through a zero-initialised gate. Under a transversality condition, gate-zero growth induces \emph{rank separation} in the functional Jacobian: old directions are unchanged, new-weight direction…

Review
pending
Role
unreviewed
Read
now
huggingface Score 24.0

Spectral Rewiring for Exploration, Purification, and Model Merging

2026-07-03 · Zhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

General AI

Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilit…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.2

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

2026-07-15 · Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu

General AI

Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core…

Review
pending
Role
unreviewed
Read
now
huggingface Score 21.4

GRASP: GRanularity-Aware Search Policy for Agentic RAG

2026-07-11 · Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan

General AI

Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matching or semantic similarity, and how to co…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.4

Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

2026-07-13 · Said Elnaffar, Farzad Rashidi

Research Track B · General AI

Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready w…

Review
pending
Role
unreviewed
Read
now
huggingface Score 21.4

UniVR: Thinking in Visual Space for Unified Visual Reasoning

2026-07-14 · Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin

General AI

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRP…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.2

SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning

2026-07-13 · Jinxiu Liu, Jianru Li, Tanqing Kuang, Xuanming Liu, Kangfu Mei, Yandong Wen, Weiyang Liu

Research Track A · General AI

Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn cumulatively and evolve autonomously, which is a limitation we term the "perpetual novice"…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.2

Mask-Aware Policy Gradients for Diffusion Language Models

2026-07-16 · Haran Raajesh, Kulin Shah, Adam Klivans, Philipp Krähenbühl

General AI

Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predic…

Review
pending
Role
unreviewed
Read
now
huggingface Score 19.4

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

2026-07-16 · Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang

General AI

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely u…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.2

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

2026-07-15 · Zihao Yu, Xiu Yuan, Chongjie Zhang

Research Track A · General AI

Long-horizon robot planning requires more than predicting what actions will do next; it also requires memory of the embodied experience that makes future goals interpretable. People do not plan from the present scene alone: they draw on remembered places, object-state changes, prior procedures, and regularities reveale…

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.4

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

2026-07-14 · Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu

General AI

Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

2026-07-16 · Paul Kassianik, Blaine Nelson, Yaron Singer

General AI

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry quer…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

2026-07-16 · Shaoxiong Zhan, Shi Hu, Boyu Feng, Hai Lin, Andrew Gong, Zhengda Zhou, Jiaying Zhou, Yunyun Hou, Hao Su, Hai-Tao Zheng

General AI

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscu…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

Plover: Steering GUI Agents through Plan-Centric Interaction

2026-07-16 · Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu

Research Track B · General AI

Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and n…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.9

Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents

2026-07-14 · Richmond Alake, Cesare Bernardis, Paul Cayet, Luca Engel, Damien Hilloulin, Sungpack Hong, Allen Hosler, Nickolas Kavantzas, Ingo Kossyk, Son Le, Rhicheek Patra, Kartik Talamadupula, Valentin Venzin

Research Track A · General AI

Agent memory is a systems problem for long-horizon agents. Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumulation of procedural knowledge from prior outcomes. These requirements extend beyond document retriev…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.4

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

2026-07-16 · Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao

General AI

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on interm…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

2026-07-16 · Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

General AI

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standard…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

2026-07-16 · Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz

General AI

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the prese…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.9

Traceback Translators Against Forgetting in Continual Fake Speech Detection

2026-07-14 · Enrico Gottardis, Mattia Tamiazzo, Simone Milani

Research Track A

Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors

2026-07-16 · Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis

General AI

The reliability of deepfake detectors frequently degrades under black-box adversarial transfer, as these models often rely on fragile, architecture-dependent forensic cues. Existing transfer attacks often lack semantic awareness and struggle to maintain effectiveness under strict no-query constraints, particularly when…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

2026-07-16 · Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu

General AI

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint, which commits every retrieval direction …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search

2026-07-16 · Debayan Mukhopadhyay, Utshab Kumar Ghosh, Shubham Chatterjee

General AI

Retrieval systems are trained and evaluated on a static idea of usefulness: hand a document and a question to a reader model, see whether the answer improves, and score the document accordingly. The idea holds up when a document is read on its own. It breaks when a language model works as a search agent, issuing severa…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Hierarchical Denoising For Multi-Step Visual Reasoning

2026-07-16 · Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang

General AI

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradig…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

2026-07-16 · Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma, Yao Yao, Weiran Qi, Chuyan Jin, Guiyu Ma, Xingzhong Xu, Kai Yang, Ji-Rong Wen, Zhicheng Dou

General AI

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can beco…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

2026-07-16 · Hoang-Loc Cao, Van Pham, Truong Thanh Hung Nguyen, Phuc Truong Loc Nguyen, Phuc Ho, Veronica Whitford, Hung Cao

General AI

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In depression-related datasets, labels are often assigned without structured evidence, symptom-level justification, or traceable alignment with the criteria of the Diagnostic a…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.9

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

2026-07-15 · Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi

Research Track A · General AI

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new fai…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

2026-07-14 · Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu, Chenrui Zhao, Tongshu Bian, Xiaoman Liang, Meishan Zhang, Baotian Hu, Min Zhang

General AI

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from e…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

2026-07-16 · Sushant Gautam, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Steven A. Hicks

General AI

Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented systems for question answering and explanation quality. Parameter-efficient adaptation of…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

2026-07-16 · Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu

General AI

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-st…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

2026-07-14 · Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang

Research Track B · General AI

Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether i…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy

2026-07-15 · Yiheng Huang, Zhijia Zhao, Bihuan Chen, Susheng Wu, Zhuotong Zhou, Yiheng Cao, Kun Hu, Xin Hu, Xin Peng

General AI

Open source software is vulnerable to supply-chain attacks through transitive dependencies, especially malicious code injected into NPM packages. Existing detectors often inadequately model obfuscated behavior, overlook JavaScript's object-centric features, poorly coordinate static and dynamic analysis, and lose semant…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Can We Trust Item Response Theory for AI Evaluation?

2026-07-16 · Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang

General AI

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT es…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Learning in Infinitesimal Non-Compositional Sketches

2026-07-16 · Sridhar Mahadevan

General AI

This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting. Machine learning problems are specified as sketches: graphs with commutat…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.2

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

2026-07-16 · Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, Jürgen Schmidhuber

General AI

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scient…

Review
pending
Role
unreviewed
Read
now
huggingface Score 9.4

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

2026-07-15 · Rui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong

General AI

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reason…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning

2026-07-16 · Pengcheng Zhou, Xuanyu Liu, Yanchen Yin, Bobo Li, Shengqiong Wu, Mong-Li Lee, Wynne Hsu

General AI

Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-localization, yet existing models remain susceptible to landmark bias, causing them to overlook geographical cues or form spurious correlations, ultimately resulting in inaccurate localization. To systematically investigate this issu…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

2026-07-16 · Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang

General AI

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, Diffusion…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.9

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

2026-07-16 · Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner

Research Track A

In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In parti…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.4

MORFEO: Advancing Towards Final Design

2026-07-14 · Lorenzo Busoni, Guido Agapito, Marco Bonaglia, Alfio Puglisi, Marco Xompero, Matteo Aliverti, Francesca Annibali, Carmelo Arcidiacono, Natalia Auricchio, Nicolò Azzaroli, Andrea Balestra, Alessandro Ballone, Louis Barbier, Andrea Baruffolo, Federico Battaini, Maria Bergomi, Andrea Bianco, Michele Cantiello, Giulio Capasso, Giulia Carlà, Enrico Cascone, Ed Chapin, Manal Chebbo, Simonetta Chinellato, Vincenzo Cianniello, Paolo Ciliegi, Mirko Colapietro, Jean-Jacques Correia, Giuseppe Cosentino, Elia Costa, Matteo D'ambrogio, Vincenzo De Caprio, Giuseppe De Luca, Nicholas Devaney, Ivan Di Antonio, Amico Di Cianno, Simone Di Filippo, Benedetta Di Francesco, Ugo Di Giammatteo, Chiara Di Prospero, Gianluca Di Rico, Andrea Di Rocco, Daphne Diretto, Christian Eredia, Simone Esposito, Jacopo Farinato, Italo Foppiani, Takashi Funakawa, Fulvio Gianotti, Laurence Gluck, Davide Greggio, Sylvain Guieu, Marco Gullieuszik, Yuuichi Harikane, Masahiro Ikoma, Laurent Jocou, Dan Kerley, Mikio Kurita, Salvatore Lampitelli, Tommaso Lapucci, Fulvio Laudisio, Yves Magnard, Demetrio Magrin, Hossein Mahmoodzadeh, Dheeraj Malik, Luca Marafatto, Laurence Michaud, Christophe Michel, Satoshi Miyazaki, Kentaro Motohara, David Mouillet, Thibaut Moulin, Matteo Munari, Kentaro Nagamine, Sylvain Oberti, Fabrice Pancher, Giorgio Pariani, Sophie Penger, Amedeo Petrella, Laurent Pinard, Cédric Plantet, Elisa Portaluri, Kalyan Radhakrishnan, Roberto Ragazzoni, Edoardo Redaelli, Edgar Renault, Colin Richardson, Marco Riva, Sylvain Rochat, Gabriele Rodeghiero, Luca Rosignoli, Bernardo Salasnich, Benoit Sassolas, Salvatore Savarese, Marcello Scalera, Pietro Schipani, Danilo Selvestrel, Mahshid Shiri, Mina Sibalic, Malcolm Smith, Sebastian Soler, Rosanna Sordo, Alessandro Tacchini, Alessio Taranto, Ludovico Teodori, Gabriele Umbriaco, Yoshinori Uzawa, Angelo Valentini, Jean-Pierre Véran

Research Track A

The Multiconjugate adaptive Optics Relay For ELT Observations (MORFEO) is a first-generation adaptive optics module for the Extremely Large Telescope (ELT), designed to deliver a diffraction-limited, highly uniform 53x53 arcsec field of view to the MICADO near-infrared camera. As the project advances toward its Final D…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.4

MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox

2026-07-14 · Zhenwen Miao, Honglin Wang, Mingheng Mi

Research Track A · General AI

As LLM technology advances, the space of model families, compute hardware, quantization schemes, parallelization strategies, and specialized optimization kernels continues to expand, sharply increasing the code complexity and maintenance cost of general-purpose inference frameworks. Conventional software engineering us…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

2026-07-16 · Ziren Gong, Xiaohan Li, Fabio Tosi, Ninghui Xu, Stefano Mattoccia, Jianfei Cai, Matteo Poggi

General AI

This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feed-forward model from the 3R family to process RGB videos and regress local point maps, and on a merging model, MAGMA, that combines loc…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

SceneBind: Binding What and Where Across Vision, Audio and Language

2026-07-16 · Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

General AI

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addr…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.4

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

2026-07-15 · Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang

General AI

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quali…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.4

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

2026-07-15 · Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li

General AI

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV se…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

2026-07-14 · Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal

General AI

Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 2…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

AutoSynthesis: An agentic system for automated meta-analysis

2026-07-16 · Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano, Francesco Pierri, Stefan Feuerriegel

General AI

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficult to scale. Here, we introduce AutoSynthesis, an end-to-end multi-agent system for automated meta-analysis. Given a res…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

Online Neural Space Time Memory for Dynamic Novel View Synthesis

2026-07-16 · Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo

General AI

Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models man…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

TikStance: A Multimodal and Hierarchical Dataset for Multi-target Stance Analysis in TikTok Political Conversations

2026-07-16 · Yazhi Zhang, Fuqiang Niu, Bowen Zhang

General AI

Political discourse has increasingly moved to short-video platforms, yet computational analysis of such content remains constrained by the scarcity of datasets that jointly preserve audiovisual information and hierarchical conversations. Here we present TikStance, a multimodal and context-aware dataset comprising 161 v…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

2026-07-16 · Qiwei Li, Jorge Ortiz

General AI

Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observational and collected without interventions, which makes causal questions such as "How would rain change traffic density?" difficult to answer. We present teLLMe, a system for explora…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 5.4

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

2026-07-15 · Sietse Schelpe

General AI

We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored, by graft, into a fresh inference context. The restore is bit-exact: under a pinne…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.4

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

2026-07-16 · Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang

General AI

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.4

WanSong v1.0 Technical Report

2026-07-16 · Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

General AI

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.2

Disintegration Temporal Logic for Probabilistic Hyperproperties

2026-07-16 · Mishel Carelli, Bernd Finkbeiner

General AI

We introduce Disintegration Temporal Logic (DTL), a new probabilistic temporal logic that can express a wide range of probabilistic hyperproperties, including probabilistic non-interference and perfect indistinguishability. DTL is based on the notion of measure disintegration from probability theory, which allows for c…

Review
pending
Role
unreviewed
Read
later