Research Paper Cockpit

Daily Digest - 2026-09-09

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-09-22.

Papers

69 visible entries

arxiv Score 27.5

ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

2026-09-05 · Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong

Research Track A · General AI

Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, a…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.5

Continual Learning Mechanisms Compose for Long-Horizon Memorization

2026-09-07 · Zheyuan Zhang, Alvin Zhang, Daniel Khashabi, Tianmin Shu

Research Track A

Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training …

Review
pending
Role
unreviewed
Read
now
huggingface Score 21.0

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

2026-09-04 · Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang

General AI

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introd…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.0

KanAdapter: A Kolmogorov-Arnold Network-based Plug-and-Play Module for Efficient Fine-tuning of Foundation Speech Models

2026-09-04 · Phuong Tuan Dat, Phuong Khai Minh, Tran Huy Dat

Research Track A

Fully fine-tuning self-supervised learning (SSL) speech models for downstream tasks is computationally prohibitive, and existing parameter-efficient fine-tuning approaches predominantly rely on MLP-based adapters whose fixed activation functions limit their representational expressiveness under tight parameter budgets.…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.5

SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction

2026-08-31 · Bowei He, Xiaokun Zhang, Meng Ding, Xue Liu

Research Track B · General AI

Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step, but they treat the skill library as a flat or two…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.4

Training-Free Task Vectors for LLM Behavioral Control

2026-09-08 · Gabriel J. Perin, Lucas Boscaini, André Araujo, Nina S. T. Hirata

Research Track A · General AI

Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.0

Selective Knowledge Control for Continual GUI Agent Learning over Application Streams

2026-09-06 · Zirui Shang, Xin Shu, Yang Liu, Zhi Gao, Xinxiao Wu, Lifeng Fan

Research Track A · Research Track B · General AI

Continual learning is a crucial capability for Graphical User Interface (GUI) agents to adapt to evolving applications while retaining knowledge acquired from previous applications. Such application streams pose a challenging knowledge modeling problem: new applications often share underlying knowledge with past ones, …

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

Drive by Hindsight and Foresight: Tool-Grounded Synergistic Reasoning over Hierarchical Memory for Autonomous Driving

2026-09-08 · Baojie Chen, Zijun Jia, Jing Zhong

Research Track A · General AI

VLMs have shown promise for autonomous driving, yet still suffer from hallucination, weak spatio-temporal perception, and limited generalization. Recent methods improve reasoning and decision-making through CoT explanations, retrieval-augmented generation or the static injection of tool outputs. Although these mechanis…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents

2026-09-08 · Boyu Yang, Jiazheng Sun, Zilong Lu, Zhi Qiu, Xin Peng, Jun Zheng

General AI

Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences and task knowledge across extended interactions. Conventional retrieval mechanisms optimize semantic compatibility rather than downstream utility, frequently introducing outdated, misleading, or conflicting evide…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.2

ReCite: Agentic Reasoning for Faithful Citation

2026-09-08 · Yuyang Huang, Bobo Li, Jiajia Song, Yuzhe Ding, Chong Teng, Fei Li, Donghong Ji

General AI

Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating core claims. However, manually navigating the growing volume of scientific literature is increasingly difficult, prompting reliance on automatic citation recommendation. While modern retrieval-augmented architectu…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.8

SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation

2026-09-03 · Qi Liu, Qinzheng Wang, Yiming Bie

Research Track A · General AI

As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-evolving agent built around the Simulation of Urb…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.2

PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation

2026-09-08 · Yixuan Liu, Zilong Zhen, Yin Wu, Yi Li

General AI

As Large Language Model (LLM) agents increasingly automate offensive operations across the cyber kill chain, their efficacy in complex local post-exploitation tasks remains inadequately quantified. Among these, Linux privilege escalation is a key step between initial access and full system compromise. However, existing…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.0

TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning

2026-09-07 · Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann

Research Track A · General AI

The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly on edge hardware using local user data. However, this shift requires optimization of deep learning training on resource-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.5

When and What to Teach: Budget-Aware Online Adaptation for Web Agents

2026-08-31 · Jianwei Zhang, Sihan Cao, Pengcheng Zheng, Ya Wen, Pei Ke, Kuien Liu, Shen Gao, Wei Dong, Yang Yang, Chaoning Zhang

Research Track B · General AI

Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners must rely on lightweight local models that evolve p…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

Evaluation of Contextual Understanding in Large Language Models

2026-09-08 · Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Uthayasanker Thayasivam, Kamal Premaratne

General AI

Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs ext…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.2

Measuring LLM Sycophancy under Sustained Multi-Turn Pressure

2026-09-08 · Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang

General AI

Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in whi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts

2026-09-07 · Kai Guo, Chuanbin Liu, Peng Hu, Hao Wang, Xi Peng

Research Track A · General AI

Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. However, existing methods typically assume that the set of modalities per task is predefined and fixed. In this paper, we investigate a more …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.5

Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery

2026-09-06 · Keyu Lin, Fei Ye, Qihe Liu, Shijie Zhou, Jiguo Yu

Research Track A

Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, a…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.2

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

2026-09-08 · Jaewoo Lim, Sungbok Shin, Sanghyun Hong

General AI

Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?

2026-09-07 · Daniel Gomm, Maarten de Rijke, Madelon Hulsebos

General AI

Democratizing access to the knowledge held in large corpora of tables such as data lakes is emerging as a central research challenge. Research in this space is advancing and broadening in scope, increasingly supplying the components to satisfy a person's insight need end-to-end. Yet these efforts remain fragmented acro…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

Omni Interaction Agent Technical Report

2026-09-08 · Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddysun, Steveyves, Zhou Zhao, Bryanytian

General AI

In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enab…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.2

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

2026-09-08 · Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan Ö. Arık

General AI

Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order, and under which conditions. As trajecto…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses

2026-09-07 · Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu

General AI

Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). Unlike direct prompt injection, IPI hides adversarial instructions in untrusted tool outputs and can covertly alter the execution of a legitimate task. Defenses trained on fixed attacks may fail as an attacker changes its strategy, inject…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

2026-09-08 · Mar Gonzàlez I Català, Haitz Sáez de Ocáriz Borde, Davide Murari, Carola-Bibiane Schönlieb, Pietro Liò, George Montañez

General AI

Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves ov…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

2026-09-08 · Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen

General AI

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective sup…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

2026-09-08 · Jingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun, Hang Zhang, Jianhua Zhu, Yufeng Wang

General AI

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. F…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

2026-09-08 · Wenbo Gao, Zhaomou Song, Zhiyuan Ji, Renxi Liu, Xing Li, Xianzhi Yu, Xiaoguang Li, James Chung-wai Cheung, Weizhe Lin, Yaoyuan Wang

Research Track A · General AI

Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without s…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation

2026-09-08 · Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler, Manuel Schaaf, Markus J. Hofmann

General AI

This study examines whether individual text corpora (ICs) from search histories can be used to simulate individual knowledge. We collected ICs from 316 adults, who answered 36 multiple-choice knowledge items, and compared several large language models (LLMs) on this task, of which only Qwen3-1.7B proved viable. After t…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.2

PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

2026-09-08 · Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie

General AI

While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of …

Review
pending
Role
unreviewed
Read
now
huggingface Score 13.0

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

2026-09-07 · Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He

General AI

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning

2026-09-05 · Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong

General AI

Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module he…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.2

ExecCritic: Learn to Test, Test to Improve for Coding Agents

2026-09-08 · Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao

General AI

Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create f…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

2026-09-04 · Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi

General AI

Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression metho…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding

2026-09-07 · Chia-Hui Chen, Shih-Ying Yeh, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai

General AI

In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

2026-09-02 · Suyoung Lee, Myungsub Choi

General AI

Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instr…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.4

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

2026-09-08 · Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun

Research Track A · General AI

Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet convention…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment

2026-09-08 · Zhan-Lun Chang, Dong-Jun Han, Seyyedali Hosseinalipour, Mung Chiang, Christopher G. Brinton

General AI

Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments

2026-09-08 · Xilin Wang, Guoxi Zhang, Hongming Xu, Zhuofan Zhang, Tianxu Wang, Lifeng Fan

General AI

Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exploration, an LN agent must consolidate experience from earlier stages and reuse it in later stages, often through persistent scene represent…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Prior-free relative 6D pose estimation of multiple object instances

2026-09-08 · Behdad Khodabandehloo, Andrea Caraffa, Davide Boscaini, Fabio Poiesi

General AI

Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving from explicit 3D models to multi-view object captures to single reference images. We take this progression to its extreme by introducing prior-free relative 6D pose estimation, which lifts the assumption of kn…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning

2026-09-08 · Tim Tomashevskiy

General AI

Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before they lead to unsafe behavior. Existing approaches typically rely on safety constraints defined at design time or updated reactively during execution, assuming that such constraints remain valid over time. Howeve…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.2

ToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback

2026-09-08 · Min Zeng, Yuzhou Liu, Zhenyu Cao, Hanxiu Chen, Heng Li, Caiquan Liu, Yafei Wen, Xiaoxin Chen

General AI

High-quality tool-use data is critical for training language models to interact effectively with external tools. However, existing synthetic approaches typically follow a generate-then-filter paradigm with static post-hoc verification, often yielding inefficient data with imbalanced feature distributions. We propose To…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

Shifting from Injection to Interaction: Rethinking Web Security in the Age of LLMs and Beyond

2026-09-03 · Nivedita Singh, Alsharif Abuadbba, Yansong Gao, Surya Nepal, Hyoungshick Kim

Research Track B · General AI

Large language models (LLMs) are becoming integral to web applications and browser agents, transforming online interactions while introducing new attack vectors and reshaping longstanding web vulnerabilities. Classical threats such as cross-site scripting (XSS) can be amplified through LLM-mediated interactions, while …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.5

MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning

2026-09-04 · Guanglong Sun, Kanglei Zhou, Liyuan Wang, Qi Cheng, Hongwei Yan, Shuang Cui, Hang Su, Jun Zhu, Yi Zhong

Research Track A

General continual learning (GCL) aims to learn from evolving data streams without task identities, explicit boundaries, or repeated access to previous data, making it a realistic yet challenging setting for continual intelligence. Although pretrained models (PTMs) provide rich prior knowledge for addressing the limited…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.5

Streaming Hierarchical Inference with Tabular Foundation Models

2026-09-07 · Vitor Crista, Afonso Lourenço, Diogo Martinho, Goreti Marreiros

Research Track A · General AI

Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency. We propose \textit{HINT}, a hierarchical inference framework that combines edge-based…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.2

Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

2026-09-08 · Mariia Drozdova, Stéphane Liem Nguyen, François Fleuret

General AI

Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph c…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.2

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

2026-09-08 · Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu

General AI

While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this g…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware

2026-08-12 · Ayushman Bhattacharya, Nihal Gazi

Research Track A · Research Track B · General AI

AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers grounded in live web results. We developed OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference. Its local search, caching, ses…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.4

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

2026-09-08 · Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre

General AI

Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

2026-09-08 · Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur

General AI

Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harnes…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

2026-09-08 · Jiacheng Xu, Feng Chen, Xiuneng Xu, Bo An

General AI

Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To m…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

2026-09-08 · Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen, Yang Liu

General AI

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting

2026-09-04 · Maryam Fakhari, Mehran Safayani

General AI

Cryptocurrency markets exhibit extreme volatility and non-stationary dynamics that challenge conventional forecasting methods. Although Large Language Models (LLMs) have shown promise for time series forecasting, the combined effects of adaptation choices remain largely unexplored in financial settings. This study intr…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.2

DYAD: A Multimodal Dataset of Co-Located Human Assistance

2026-09-08 · Akhil Ajikumar, Mahya Qorbani, Sakib Reza, Sean Andrist, Mohsen Moghaddam

General AI

An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions

2026-08-29 · Yucheng Du, Xiyang Hu

General AI

Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evaluating (1).startswith("1"), instead of abstaining. We ask whether this failure reflects missing recognition or failed routing from recognition to abstention. Across instruction-tuned models from 1.7B to 70B param…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation

2026-09-06 · Mingwei Li, Yi Yang, Hehe Fan

General AI

Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth no…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

2026-09-07 · Xuechao Zou, Yi Zhou, Kai Li, Shun Zhang, Yuhui Chen, Congyan Lang, Junliang Xing

General AI

The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitt…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

2026-09-08 · Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

General AI

Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation,…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.2

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

2026-09-08 · Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov, Özgün Turgut, Michelle Espranita Liman, Lisa Steinhelfer, Rickmer Braren, Daniel Rueckert

General AI

The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inhe…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.0

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

2026-09-07 · Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos

General AI

Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rat…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.0

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

2026-09-07 · Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao

General AI

World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tigh…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

Copying explains the collective behavior of AI agents in the wild

2026-09-08 · Giordano De Marzo, Nicola Albore, David Garcia

General AI

In June 2026, thousands of AI agents found that a small public wiki would accept edits from inside their sandboxes, and started using it to help one another pass a timed test. Each agent lived for about an hour and remembered nothing afterwards. Nobody asked them to cooperate, and the wiki had not been built for them. …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

2026-09-08 · Yunpeng Xu, Kun Zheng

General AI

Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention

2026-09-08 · Raito Kiya, Satoki Ohashi, Kosuke Sato, Go Kamoda, Ryosuke Takahashi, Yuji Yamamoto, Daiki Shiono, Keisuke Sakaguchi, Goro Kobayashi

General AI

Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently co-occur, and MAs can pose challenges for low-bit quantization. In this study, we analyze the factors underlying AS and MAs that emerge at t…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.2

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

2026-09-08 · Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, Dhruv Shah

General AI

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjus…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 5.4

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

2026-09-08 · Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai

General AI

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp a…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.4

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

2026-09-08 · Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao

General AI

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smooth…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.2

Point4D: Long-range 4D Motion Reconstruction

2026-09-08 · Minsik Jeon, Jay Karhade, Deva Ramanan, Shubham Tulsiani

General AI

We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is able to reliably infer dense per-point 3D trajectories across multi-hundred-frame videos, unlike existing 4D methods that are limited to short input windows of at most a few dozen frames. A key innovation that ena…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.0

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise

2026-09-07 · Mika Okamoto, Gabriele Sarti

General AI

A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute t…

Review
pending
Role
unreviewed
Read
later
arxiv Score 4.8

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

2026-09-04 · Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam

General AI

Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a reference image. Recent approaches perform well when the reference image contains a single object, effectively capturing a global style that encomp…

Review
pending
Role
unreviewed
Read
later