Research Paper Cockpit

Daily Digest - 2026-07-13

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-07-25.

Papers

40 visible entries

arxiv Score 22.8

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

2026-07-10 · Nirjhar Das, Md. Al-Mamun Provath

General AI

We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.0

Interference and Retention in Continual Learning

2026-07-10 · Julius Störk

Research Track A · General AI

Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that forgetting should instead be modeled directly as interference between tasks. In the frozen-feature regime, forgetting from learning a new task is exactly the interference energy induc…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

2026-07-10 · Kaiji Zhou, Ales Leonardis, Yue Feng

General AI

Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of expert models or tools, while overlooking critical factors s…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

2026-07-10 · Shravan Murlidaran, Miguel P. Eckstein

General AI

Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding …

Review
pending
Role
unreviewed
Read
now
huggingface Score 19.0

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

2026-07-07 · Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang, Yufeng Chen, Yang Zhao, Ling Shao, Hao Tang

General AI

Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation …

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.0

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

2026-07-09 · Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang

General AI

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

2026-07-10 · Cláudio Lúcio do Val Lopes, Lucca Machado da Silva

General AI

Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome this without distortive data resampling, we propose the Semant…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

2026-07-10 · Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj, Maanak Gupta

General AI

Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testing approaches. While recent adoptions of Large Language Model (LLM) agents have demonstrated promise in penetration test…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.5

A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

2026-07-10 · Linhui Xiao, Guiping Cao, Mingyue Guo, Xianchao Guan, Fan Yang, Ming Tao, Xin Li, Yuxin Peng, Yaowei Wang

Research Track A · General AI

The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability. This survey provides a comprehensive overview of the green development of la…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

Mosaic: Runtime-Efficient Multi-Agent Embodied Planning

2026-07-10 · Kunjal Panchal, Saayan Mitra, Sunav Choudhary, Victor Bursztyn, Somdeb Sarkhel, Hui Guan

General AI

LLM-based multi-agent embodied planning remains impractical due to prohibitively high execution latency. We identify failed actions as the dominant bottleneck, stemming from two core challenges: inaccurate state tracking under partial observability and inefficient coordination that produces redundant or conflicting act…

Review
pending
Role
unreviewed
Read
now
huggingface Score 13.0

Self-Guided Test-Time Training for Long-Context LLMs

2026-07-10 · Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu

General AI

Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades, indicating that models still struggle to identify and use the evidence most relevan…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

From Raw IDs to Semantic Planning: How Recommender Systems Utilize Information at Scale

2026-07-10 · Changhong Jin, Shiqiu Yang, Roger Zhe Li, Yingjie Niu, Aghiles Salah, Mete Sertkan, Zheng Ju, Xingsheng Guo, Huifeng Guo, Ruihai Dong, Barry Smyth

General AI

The evolution of recommender systems can be explored by asking how they utilize information at scale. Throughout most of the historical period under consideration during the past two decades, industrial systems have relied on raw IDs, which are discrete, globally unique, and semantically opaque identifiers that enable …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories

2026-07-10 · Xiangxin Zhao, Han Li, Shuaiting Li, Tianyi Zhao, Earl T. Barr, Federica Sarro, He Ye

General AI

Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based environments, making their reliability a growing concern. Existing empirical studies investigate why coding agents fail, yet they largely treat failure as a final outcome rather than a…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility

2026-07-10 · Filippo Ziliotto, Luciano Serafini, Lamberto Ballan, Tommaso Campari

General AI

A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap. We demonstrate that VGGT implicitly encodes co-visibility as an emergent behavior: without any supervision for this ta…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.0

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation

2026-07-08 · Seulbin Hwang, Kiyoung Om, Daejung Kim, Jinhan Lee

General AI

Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and recent methods have optimized accordingly, leaving diversity underexplored. We introduce Flow-ERD, a multi-agent simulator that pursues realism and diversity jointly. Its …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.0

Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift

2026-07-10 · Dan C. Hsu, Luke Lu

Research Track A · General AI

Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational experience while the model, tools, and harness remain fixed. Over lon…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

Promptable Concept Segmentation from Above: Evaluating SAM 3's Zero-Shot and One-Shot Capabilities in Remote Sensing

2026-07-10 · Mohammad Dabaja, Turgay Celik

General AI

The deployment of large-scale foundation models, such as the Segment Anything Model 3 (SAM 3), promises a transition toward open-vocabulary, training-free computer vision. However, their capacity to generalize out-of-distribution to the complex, top-down geometric structures of Earth Observation imagery remains largely…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems

2026-07-10 · Refat Ishrak Hemel, Ehsan Hallaji, Roozbeh Razavi-Far

General AI

The emergence of metaverse platforms has created virtual economies that introduce new challenges related to fraud, bot activity, and illicit financial behavior. Despite growing interest in trustworthy metaverse analytics, existing datasets typically focus on user behavior, authentication, or financial transactions in i…

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.0

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

2026-07-09 · Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong

General AI

Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the \textbf{Knowing--Using Gap}, characterized by an accuracy gap and a temporal lag between memorization and generalization. To und…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works

2026-07-10 · Miguel Arana-Catania, Gillian Pink, Glenn Roe

General AI

Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper explores the application of machine learning to automatic thematic in…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

LLM for EDA in Front-End Design: Challenges and Opportunities

2026-07-10 · Kangwei Xu, Bing Li, Ulf Schlichtmann

General AI

As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great potential in Electronic Design Automation (EDA). Beyond specification understanding, LLMs show the potential to serve as a unifie…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning

2026-07-10 · Ivan Ilin, Philip Zmushko, Peter Richtárik

General AI

Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused to choose where a model should adapt. We propose Super, a sparse parameter-efficient fine-…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

2026-07-09 · Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung

General AI

The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

2026-07-10 · Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen

General AI

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary for comprehensive scene understanding. Existing radar-camera …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

2026-07-10 · Jiawen Li, Tian Guan, Huijuan Shi, Xitong Ling, Mingxi Fu, Anjia Han, Chao He, Yonghong He

General AI

Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial scales, fragmenting complementary expertise across separate backbones. Here we present ALICE, a unified foundation model trained through multi-stage agglomerative distillati…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.3

Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory

2026-07-10 · Chengkai Zhu, Ziao Tang, Guocheng Zhen, Yimeng Cao, Yusheng Zhao, Ranyiliu Chen, Xuanqiang Zhao, Lei Zhang, Xin Wang

General AI

Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correction. Formalizing its coding theorems requires connecting finite-block protocols, analytic inequalities, and asymptotic limits within…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.0

Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator

2026-07-10 · Li Hengyu

Research Track A · General AI

The attention matrix of a causal transformer is row-stochastic, iterated over depth, and non-normal by construction. For non-normal operators, eigenvalues control only asymptotic behavior; finite-depth behavior is controlled by resolvent quantities such as pseudospectra and Kreiss constants. We test, under pre-register…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

2026-07-10 · Mohadeseh Mollapour, Koorosh Aslansefat, Zeinab Dehghani, Bhupesh Kumar Mishra, Tejal Shah, Zhibao Mian

General AI

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather t…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.0

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

2026-07-09 · Zanyi Wang, Xin Lin, Haodong Li, Dengyang Jiang, Yijiang Li

General AI

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and geometric priors. Existing generative and editing approaches reuse these priors by casting dense prediction as target generation: annotations such as depth, normals, al…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.0

A Sovereign, Open-Source Foundation Model for German and English

2026-07-10 · The Soofi-Team, Benedikt Droste, David Fitzek, Ruben Härle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom Röhr, Sebastian von Rohrscheidt, Jörg Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim Köhler, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max Lübbering

General AI

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over den…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Balancing Usefulness and Naturalness: An LLM-based Curation Pipeline for Code Review Comments

2026-07-10 · Oussama Ben Sghaier, Martin Weyssow, Houari Sahraoui

General AI

Code review is a cornerstone of software development, where reviewers provide feedback through written comments to ensure code quality, maintainability, and correctness. The effectiveness of this process hinges on the quality of review comments. As large language models (LLMs) gain traction in automating code review ta…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

PanoWorld: Real-World Panoramic Generation

2026-07-10 · Haoyuan Li, Dizhe Zhang, Yuemei Zhou, Xiangkai Zhang, Haoran Feng, Xiaofan Lin, Wenjie Jiang, Bo Du, Ming-Hsuan Yang, Lu Qi

General AI

In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, where rotation can be treated as an implicit geometric transformation.Building on this insight, we propose PanoWorld, which simplifies camera t…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

2026-07-10 · Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh, Nikolai Rozanov, Kentaro Inui

General AI

Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal knowledge or a gap between internal representations and verbalized outputs. Training simple probes on activations from four VLMs across five…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.0

Trust Region Policy Distillation

2026-07-06 · Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang

General AI

Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.0

KronQ: LLM Quantization via Kronecker-Factored Hessian

2026-07-08 · Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda

General AI

Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PTQ methods, including GPTQ, construct quantization objectives exclusively from input activation statistics, effectively assuming that all output channels contribute equa…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis

2026-07-10 · Ren Takahashi, Emre Yusuf, Jayabrata Bhaduri

General AI

Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achieving a state-of-the-art area under the receiver operating characteristic curve (AUC) of approximately 0.70 on the DREAM database (Wong et al., 2025, Nature Communications). We introduc…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Scalable Visual Pretraining for Language Intelligence

2026-07-10 · Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen

General AI

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.0

Video Generation Models are General-Purpose Vision Learners

2026-07-10 · Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu

General AI

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training para…

Review
pending
Role
unreviewed
Read
later
arxiv Score 4.8

Inunda: A GPU-Native, Agent-enabled, Differentiable Solver for High-Resolution Flood Inundation Modeling

2026-07-10 · Zhi Li

General AI

Predicting where floodwater goes and how deep it gets, at high resolution and across large domains, remains computationally expensive with conventional hydraulic solvers, while purely data-driven surrogates are fast but lack physical guarantees and generalize poorly beyond their training events. We present Inunda, a GP…

Review
pending
Role
unreviewed
Read
later
arxiv Score 4.8

Kleene Algebra with Transitive Commutativity Conditions

2026-07-10 · Han Xu, Chenyu Zhou, Zachary Kincaid, David Walker

General AI

Kleene algebra (KA) provides a foundational algebraic framework for reasoning about program structure and control flow. To capture equivalences arising from reordering or independence of actions, Kozen [1996] purposed that KA can be extended with commutativity conditions, that is, equations of the form { ab = ba | (a,b…

Review
pending
Role
unreviewed
Read
later