Research Paper Cockpit

Daily Digest - 2026-10-02

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-10-04.

Papers

82 visible entries

arxiv Score 37.0

SPACE-LoRA: Allocating Activation-Subspace Protection for Continual Learning

2026-09-28 · Seunghyun Yoo, Kiseok Kim, Hyeontae Joo, Junyeop Bang, Hwangnam Kim

Research Track A

This study addresses the catastrophic forgetting problem that occurs when sequentially learning successive tasks using Low-Rank Adaptation (LoRA) from a lifelong learning perspective. While existing approaches have primarily constrained parameter updates or learning subspaces to reduce interference with past knowledge,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 30.3

Task-Oriented Rank Adaptation for Continual Learning in Text Classification

2026-10-01 · Rey Sanchez Lopez, Eduardo Morales Manzanares, Hugo Jair Escalante

Research Track A · General AI

Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact repres…

Review
pending
Role
unreviewed
Read
now
arxiv Score 30.0

ChainLoRA: Geometry-Preserving Task Vector Merging for Continual Learning in LLMs

2026-09-30 · Hang Yin, Haozhe Wang, Yuhua Luo, Zhangqi Pan, Xiaoxing Wang, Junchi Yan

Research Track A · General AI

Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new tasks, and strict parameter budgets. We present \textbf{ChainLoRA}, a replay-free continual merging framework built on chain-updated task-vector geometry. From a paramete…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.6

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

2026-10-01 · Haibo Wang, Jiteng Mu, Jialu Li, Jingru Yi, Yuanjun Xiong, Jianming Zhang, Lifu Huang, Mingze Xu

General AI

We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process:…

Review
pending
Role
unreviewed
Read
now
arxiv Score 23.5

One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents

2026-09-28 · Ziqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang, Zihuan Jiang, Linqiang Guo, Siobhan Reid, Zhi Liu, Yang Wang

Research Track B · General AI

GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployment. Because GUI acti…

Review
pending
Role
unreviewed
Read
now
huggingface Score 23.0

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

2026-09-30 · Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu

General AI

Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.6

VISTA: A Visual Harness for Reasoning in an Interactive World

2026-10-01 · Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

General AI

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly …

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.6

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

2026-10-01 · Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris

General AI

On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Rece…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.3

Learn the Directions, Normalize the Gains: Post-Training Normalization for LoRA

2026-10-01 · Zailong Tian, Yanzhe Chen, Zhuoheng Han, Houfeng Wang, Lizi Liao

Research Track A

While Low-Rank Adaptation (LoRA) enables efficient task specialization, its learned updates can compromise capabilities beyond the target task. We identify \textbf{adaptation imbalance}: a few singular directions dominate the trained update, leaving its performance sensitive to how gains are allocated. We argue that \t…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.6

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

2026-10-01 · Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao

General AI

Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an a…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.9

From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs

2026-09-26 · Yangzhe Peng, Xiaoyang Wang, Yiyang Zhao, Lijun Wu, Kun He

Research Track B · General AI

Comparative Process Reward Models (PRMs) provide critical step-level guidance for autonomous web agents by evaluating state-conditioned preferences between candidate actions. However, existing preference training data synthesized via multi-policy sampling suffers from a severe scarcity of Grounded Minimal Contrastive P…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

2026-10-01 · Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi, Hiroki Itoh, Kotaro Funakoshi

Research Track B · General AI

Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks und…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.6

From Knowledge Access to Source Learning: Developing Source-Specific Competence

2026-10-01 · Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

General AI

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source i…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.6

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

2026-10-01 · Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer

General AI

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world …

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.6

Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

2026-10-01 · Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR

General AI

Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation pro…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.3

Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations

2026-10-01 · Giulio Schiavi, Andrei Cramariuc, Michael Pantic, Roland Siegwart

Research Track A · General AI

Most current grasp synthesis systems are trained offline and remain fixed during deployment. While this works well when deployment conditions resemble the training data, performance can degrade when robots encounter conditions they have not seen before, such as unfamiliar objects. In this work, we present a continual-l…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.2

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

2026-09-24 · Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasović

Research Track B · General AI

Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such work…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.0

Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model

2026-09-30 · Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie

Research Track A · General AI

Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

AdaptArena: Evaluating Test-Time Personalization of Web Agents

2026-09-29 · Dongchan Shin, Xing Han Lù, Jiaqi Deng, Jay Gala, Tomás Vergara Browne, Jaewon Moon, Fengyuan Liu, Alexandre Drouin, Siva Reddy, Alexandre Lacoste

Research Track B · General AI

Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent prefere…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.6

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

2026-10-01 · Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang

General AI

Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors whil…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.8

TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories

2026-09-29 · Yu-Shu Chen, Yu-Jung Liang, Pengtao Xie

General AI

Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.6

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

2026-10-01 · Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

General AI

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehous…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.3

Continual Reinforcement Learning with Neuroevolution

2026-10-01 · Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi

Research Track A · General AI

Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search di…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

2026-09-30 · Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren

General AI

As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic environment over time…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.6

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

2026-10-01 · Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn

General AI

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced t…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.6

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

2026-10-01 · Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu

General AI

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLM…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.0

JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces

2026-09-30 · Haoyang Su, Weiran Huang

General AI

LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement limits autonomous task solving, where th…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.0

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

2026-09-30 · Xinge Peng, Yiting Lu, Tianwu Zhi, Wen Wen, Jianzhao Liu, Xin Li, Zhibo Chen

General AI

Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inh…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.8

AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

2026-10-01 · Sumin Lee, Sukmin Cho, Suengjae Lim, Youngjin Kwon

General AI

Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

CLASP: Continual Low-rank Adapters for Spatially Placed Concepts from One Hypernetwork

2026-10-01 · Wojciech Gromski, Patryk Krukowski, Jan Miksa, Maciej Zieba, Przemysław Spurek

Research Track A · General AI

Continual personalization of text-to-image diffusion models requires sequentially acquiring new concepts while retaining previously learned ones. However, existing methods either suffer from catastrophic forgetting or rely on storing additional concept-specific parameters and spatial components, causing their parameter…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.6

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

2026-10-01 · Arman Behnam, Binghui Wang

Research Track A · General AI

Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.6

Form and Void: Entangled Composition through an Autonomous AI Agent

2026-10-01 · Shiwen Wang, Jian Yang, Xu Wang, Xincan Wang, Weiming Dong

General AI

Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.6

Hierarchical Continuous Diffusion Language Models

2026-10-01 · Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing

General AI

Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistica…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.6

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

2026-10-01 · Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui

General AI

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.5

Dynamic LoRA-Experts and Prototype-Ensemble Matching for Class-Incremental Learning

2026-09-30 · Hongwei Zhao, Rui Liu, Yansong Liu

Research Track A · General AI

Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. Parameter-efficient fine-tuning with pre-trained models reduces parameter overhead but can suffer from cumulative interference and suboptimal alignment between inference samples and specialized modu…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.3

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

2026-10-01 · Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal

Research Track A · General AI

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Probe to Act: Elevating Browser-Use Agent via Active Visual Probing

2026-09-27 · Keliang Li, Heng Wang, Chen Hu, Daxin Jiang, Hong Chang, Shiguang Shan

Research Track B · General AI

Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alig…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

2026-09-28 · Ruozhao Yang, Mingfei Cheng, Xiaofei Xie

Research Track B · General AI

LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still re…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.8

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

2026-10-01 · Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim

General AI

Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.6

CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

2026-10-01 · Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao

General AI

Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.6

MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

2026-10-01 · Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu

General AI

Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training reci…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.6

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

2026-10-01 · Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi

General AI

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model wei…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.5

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

2026-09-28 · Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou

Research Track B · General AI

Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowle…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.8

Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs

2026-10-01 · Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi

General AI

Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, which conflates robus…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.8

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

2026-10-01 · Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu

General AI

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.6

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

2026-10-01 · Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong

General AI

Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, a…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.6

Finetuning with Sampling: SFT Learns Better Than You Think

2026-10-01 · Aayush Karan, Sitan Chen, Yilun Du

Research Track A · General AI

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is …

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.6

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

2026-10-01 · Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin, Jinhyeok Choi, Junyoung Seo, Seungryong Kim

General AI

How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, they lose direct evidence of its evolution, often failing to preserve its state and dynami…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.0

WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents

2026-09-28 · Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev, Alena Fenogenova

Research Track B · General AI

We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emit typed events with …

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.0

Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention

2026-09-30 · Siddeshwar Raghavan, Ziqin Yuan, Fengqing Zhu, Byung-Cheol Min

Research Track A · General AI

Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills. However, retaining task performance does not ensure the behavior remains grounded in language because policies may rely on scene cues, object associations, or memorized task structure. We introdu…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.8

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

2026-10-01 · Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen, Yansong Shi, Junhao Zhou, Yifei Li, Jun Zhang, Chuanyu Qin, Chenxu Yang, Xinlei Yu, Kun Ouyang, Yuchen Shao, Qianshan Wei, Changhai Zhou, Jun Gao, Jiaqi Wang, Limin Wang

General AI

Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.6

ROWBench: Do Video Models Render What the Program Specifies?

2026-10-01 · Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang

General AI

Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instructio…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.4

Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding

2026-09-25 · Valeriy Vyaltsev, Anton Andreychuk, Taisia Zlotnikova, Konstantin Yakovlev, Aleksandr Panov, Alexey Skrynnik

General AI

Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same contex…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.4

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

2026-09-26 · Sitao Cheng, Xunjian Yin, Zhiyuan Sun, Yuxuan Li, Ruiwen Zhou, Xiangru Jian, Victor Zhong

Research Track B · General AI

Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less fully than its conten…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.0

CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning

2026-09-29 · Wenbin Hu, Huihao Jing, Haochen Shi, Yuxuan Liu, Haoran Li, Yangqiu Song

General AI

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation.…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 11.0

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

2026-09-30 · Fedor Rodionov, Aleksandar Cvejic, Michael Birsak, John Femiani, Peter Wonka

General AI

Furnished floor plans support real-estate visualization, interior design, and architectural workflows, yet automatic furnishing remains challenged by limited real-world data and the need to satisfy interacting geometric and functional constraints. We ask whether professional furnishing knowledge can be learned from rea…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.8

Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

2026-10-01 · Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou

General AI

Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with select…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.6

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

2026-10-01 · Jichao Jiang, Cristian McGee, El Houcine Bergou, Hanqin Cai, Aritra Dutta

General AI

Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently int…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.6

Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling

2026-10-01 · Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona, Yen Ting Lin

General AI

There has been a proliferation of sampling algorithms based on Wasserstein gradient flows (WGF) and forward-only diffusion processes (FODP), often accompanied by theoretical guarantees of exponentially fast convergence to the target distribution. These guarantees are frequently interpreted as evidence that such methods…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.5

Structured Interaction, Visual Localization, and Robust Execution for Complex Web Tasks: A Technical Report on the WebRetriever Challenge

2026-09-28 · Ziqi Zhang, Shaohui Li, Bing Li

Research Track B · General AI

This report presents the web agent system developed for the WebRetriever Challenge. The system follows a structuredinteraction- first strategy, using semantic webpage information for routine browser operations and invoking visual perception only when structured representations are insufficient. Three key designs are in…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Architectural Degradation: How to Measure and to Remediate

2026-10-01 · Noman Ahmad, Ruoyu Su, Matteo Esposito, Andrea Janes, Valentina Lenarduzzi, Davide Taibi

Research Track A · General AI

Context. Architectural degradation undermines software maintainability, evolvability, and quality. However, existing research remains fragmented across measurement approaches, metrics, tools, and remediation strategies, limiting our understanding of how these elements relate across the degradation lifecycle. Aim. We co…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.6

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

2026-10-01 · Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin

General AI

Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determin…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.6

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

2026-10-01 · Yinheng Li, Justin Wagle

General AI

Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fin…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.6

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

2026-10-01 · Zilin Du, Bowen Yang, Boyang Albert Li

General AI

Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valua…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.5

Zengram-Lite: An In-Browser Agentic-Memory Framework - Semantic Knowledge, Session Tracking, and Token-Budgeted Context

2026-09-03 · Gene Zhang

Research Track B · General AI

AI agents increasingly run in the browser, and they need somewhere to keep what they learn. The client-side state of the art, however, is a vector index - nearest-neighbor search over embeddings - with the rest of an agent's memory left to application code: the conversation history is an array in localStorage, context …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.0

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

2026-09-28 · Jiayi Li, Ruizhe Li

General AI

Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collater…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.0

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

2026-09-29 · Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov

General AI

Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of w…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.0

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

2026-09-29 · Minoo Kim, Vasileios Lampos, George Drayson

General AI

Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded mo…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.0

LOCI: Spatial Linear Memory for Streaming World Models

2026-09-30 · Ji Xia, Tingting Liao, Xuezhi Liang, Hao Li, Guangyi Liu

General AI

When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but co…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

2026-10-01 · Junkang Liu

General AI

Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

Learning to structure data from user-generated thematic corpora

2026-10-01 · Elishay Avram, Oren Glickman, Elad Yom-Tov

General AI

Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, do…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models

2026-10-01 · Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen

General AI

A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but the…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

2026-10-01 · Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev

General AI

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

2026-10-01 · Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen, Kiet A. Nguyen, Adheesh Sunil Juvekar, Ismini Lourentzou

General AI

High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models

2026-10-01 · Jianhong Li, Jiahao Chen, Yuwen Pu, Chunyi Zhou, Oubo Ma, Zhou Feng, Hangtao Zhang, Jichao Bi, Chunqiang Hu

General AI

Beyond adapting Large Language Models (LLMs) to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same distribution, i.e., th…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Analysis of Quantized and Efficiently Adapted Protein Language Models

2026-09-30 · Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee

General AI

Background: Protein language models (PLMs) are increasingly used for sequence generation and property prediction, but their size makes fine-tuning and deployment expensive. The effects of quantization and parameter efficient fine-tuning on performance, representations and generation remain insufficiently characterized.…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

Embedding Prediction Helps Image Generation

2026-10-01 · Sihan Xu, Ji Xie, Zilin Wang, Hui Shen, Stella X. Yu

General AI

In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

2026-10-01 · Hang Zou, Chao Zhang, Yuzhi Yang, Yu Tian, Samson Lasaulce, Mérouane Debbah

General AI

Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the commun…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

2026-10-01 · Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari

General AI

Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.6

A comprehensive simulation framework for multi-modal kilonova observations from all-sky surveys

2026-10-01 · Felipe Fontinele Nunes, Andrew Toivonen, Farhana Taiyebah, Leonard Lupin-Jimenez, Skylar Callis, Malhar Kulkarni, Soumi De, Michael W. Coughlin

General AI

Current all-sky surveys such as Vera Rubin's Legacy Survey of Space and Time and the Zwicky Transient Facility (ZTF) promise a wealth of scientific gains that are contained in the millions of astrophysical transient candidates produced each night. Kilonovae, one such transient of interest, will be challenging to identi…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Moore, Escher, Penrose: A Conformal Golden Braid

2026-10-01 · Sophia Feldman, Assaf Shocher

General AI

I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved universe.'' So wrote M.C. Escher about his 1956 lithograph Pr…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation

2026-09-30 · Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi

General AI

Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the Nation…

Review
pending
Role
unreviewed
Read
later