Research Paper Cockpit

Daily Digest - 2026-09-17

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-09-22.

Papers

64 visible entries

arxiv Score 25.0

Uncertainty-Aware Continual Learning for Open-World Intent Discovery Under an evolving Label Space

2026-09-15 · Pisante Aida, Formentin Simone

Research Track A

Real-world intelligent systems increasingly operate under open-world conditions, where user intents are not fixed or exhaustively known a priori and may evolve as new interaction patterns emerge. This paper proposes a unified uncertainty-aware probabilistic framework for continual new intent discovery under an evolving…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.5

CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework

2026-09-15 · Yunxiang Fu, Meng Lou, Zicheng Liao, Yizhou Yu

Research Track A · General AI

Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a data stream without catastrophic forgetting. While leveraging pretrained models has significantly advanced continual learning, existing methods exhibit a scalability bottl…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.9

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

2026-09-11 · Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal

Research Track B · General AI

Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

Clueing up LLMs with Tool-Augmented Deductive Reasoning

2026-09-16 · Rebecca Ansell, Autumn Toney-Wails

General AI

Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that require integrating evidence across multiple reasoning steps, maintaining consistency with prior inferences, and updating beliefs under new constraints …

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.2

AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery

2026-09-14 · Junhao Qiu, Qinglong Hu, Xialiang Tong, Mingxuan Yuan, Liyong Lin, Qingfu Zhang

General AI

Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, blocks cross-paradigm transfer, and discards valuable execution feedback. We propos…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.0

DR.WILSS: Diffusion-Based Replay for Weakly Supervised Continual Semantic Segmentation

2026-09-16 · Leon Arthur Marx, Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh

Research Track A · General AI

Weakly supervised class-incremental semantic segmentation (WILSS) aims to train a segmentation model over multiple steps, each introducing new concepts to be learned with only image-level supervision. We introduce DR.WILSS, an innovative approach to address catastrophic forgetting in continual learning using diffusion-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents

2026-09-16 · Izumi Takahara, Kazunori Nishio, Akira Aiba, Shigeru Kobayashi, Takao Nakajima, Taro Hitosugi, Teruyasu Mizoguchi

General AI

Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of optimized samples, with the measurements reduced to predefined scalar objectives and the reasons behind success left unarticulated. Here we present SynA…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

2026-09-16 · Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang

General AI

Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience b…

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.0

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

2026-09-15 · Caiqi Zhang, Xiaochen Zhu, Chengzu Li, Yulong Chen, Dharshan Kumaran, Nigel Collier

General AI

Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however, share one design premise: they only read the cur…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.9

LifeMem: Enabling Lifelong Experience Reuse for LLM Agents

2026-09-11 · Yuli Qiu, Yutong Li, Wei Su, Zeming Liu, Wanxiang Che, Heyan Huang, Haifeng Wang, Yuang Guo

Research Track A · General AI

Large language model agents are expected to continuously adapt to new tasks and environments over their lifetime by reusing past experience. However, existing memory-based agents struggle to transfer reusable experience across environments and suffer from catastrophic forgetting as experience accumulated. To address th…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

2026-09-16 · Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li

General AI

Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these p…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.0

EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

2026-09-15 · Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang

General AI

Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk u…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

2026-09-15 · ZhuoXin Liu, Zhiming Ma, Ying Zhang, Mengzheng Yang, Yifan Wang, Zhengqi Huang, Yanhan Zhou, Zekun Lin, Jun Zhang, Shun Zhang, Yue Chen, Qiao Zhao, Peng Chen

Research Track B · General AI

Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages s…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

2026-09-15 · Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li, Zhaochen Yu, Jinchao Ding, Qiang Gao, Pengyu Zhan, Yuntong Zhang, Tian Cheng, Zhenfei Yin, Yingcheng Wu, Ling Yang

Research Track A · General AI

We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks …

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

2026-09-16 · João Meneses dos Santos, Arlindo L. Oliveira

General AI

Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action proposer with a slower planner, using two modular cognitive extensions: an Adaptiv…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.5

Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation

2026-09-15 · Hojin Lee, Yunho Lee, Daniel A Duecker, Cheolhyeon Kwon

Research Track A

Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods of…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination

2026-09-16 · Matteo Golinelli, Idilio Drago, Matteo Boffa, Francesco Bergadano, Bruno Crispo

General AI

AI agents for security inspect web pages, source code, logs, configuration files, and command outputs. These environments may contain deceptive artifacts that influence the agent's behavior. We call this adversarial task contamination. Whereas prompt injection relies on attacker-supplied instructions, task contaminatio…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

ERAF4XRD: A multimodal agentic framework for constructing validated experimental X-ray diffraction databases from scientific literature

2026-09-16 · Afnan Mostafa, William Ratcliff, Simon J. L. Billinge, Niaz Abdolrahim

General AI

The scientific literature contains decades of experimental measurements that remain difficult to access as structured data for modern AI and data-driven research. Much of this information is distributed across figures, captions, text, and tables, requiring experimental data and their context to be identified, connected…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Which LLM is Best for Translating Natural Language Goals to PDDL

2026-09-16 · Tomas Balyo, Lukas Chrpa, G. Michael Youngblood

General AI

Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts accessibility to non-experts. This paper empirically evaluates whether current Large Language Models (LLMs) can reliably translate natural language testin…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling

2026-09-15 · Xiaoyang Liu

General AI

Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, ma…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

2026-09-15 · Cai Ke, Xin Liu, Han Zhang, Jiangyue Yan, Zike Yuan, Ling Deng, Yue Yu, Hui Wang, Ruifeng Xu

General AI

Lifelong conversational agents rely on memory systems to maintain deep, context-aware interactions with users. However, existing explicit textual memory pipelines suffer from a severe information bottleneck, often losing subtle behavioral patterns and emotional shifts. Furthermore, being typically static post-deploymen…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

2026-09-16 · Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng

General AI

Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason abou…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.5

MCLC-NET: Multimodal Continual Learning for Leaf Counting

2026-09-16 · Ruchi Bhatt, Pratibha Kumari, Shreya Bansal, Vedant Agnihotri, Dwarikanath Mahapatra, Mukesh Saini

Research Track A · General AI

Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield. Most existing methods rely on RGB images, but their performance is often affected by occlusion, lighting variations, and other real-world challenges. Additional modalities, such as depth and thermal images, ca…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

Artificial Intelligence-Enabled Space Robot Operations: Technologies, Challenges and Prospects

2026-09-15 · Zeyuan Huang, Gang Chen, Zixuan Hao, Guoqin Tang, Junyi Zong, Guoyou Ban, Jiale Wang, Haoyang Lv, Chaoqian Ren, Sitong Liu

Research Track A · General AI

Space robots are increasingly expected to perform long-duration, contact-rich, and multi-stage operations with limited human intervention. Recent advances in artificial intelligence (AI), robot learning, and embodied foundation models provide new opportunities to improve the autonomy and adaptability of such systems, b…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents

2026-09-16 · Jiaxuan Jiang, Liyuan He, Zhixuan Fang

General AI

Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during post-training and prevents agents from ac…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

2026-09-16 · Peter Potash

General AI

We evaluate six frontier language models on the two-agent $\log(N)$-Questions game. A questioner sees $N$ Wikipedia lead paragraphs and must identify a secretly chosen target using exactly $\log_2 N$ yes/no questions. An answerer sees only the target and the question, and replies with one word. Both roles run on the sa…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.4

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

2026-09-14 · Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang

General AI

Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

Agora: Git as Shared Memory for Collective AutoResearch

2026-09-16 · Yifan Zhang, Yunheng Zou, Shaokun Zhang, Jian Hu, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong

General AI

Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared memory for such agents: research is recorded as an…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

2026-09-16 · Yizhuo Li, Jianhao Yan, Yun Luo, Zhi Wang, Futing Wang, Rong-Xi Tan, Kanghui Tian, Ganqu Cui, Ning Ding, Peilin Zhao, Yafu Li, Yu Cheng

General AI

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Car…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

A Zeroth-Order Paradigm for LLM Preference Alignment

2026-09-16 · Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin

General AI

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

2026-09-16 · Dunyao Xue, Chengshuo Du, Zhengbo Wang, Wenlin Dai, Cheng Meng

General AI

We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. Existing selection strategies rely predominantly on scalar probabilities, ignoring geometric semantic relationships and causing candidate redundancy.…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection

2026-09-16 · Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang

General AI

Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing methods face a fundamental trade-off between accuracy and efficiency: direct full-image inference often fails to captur…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

When Agents Look Like Beacons: NIDS Evasion by Model Context Protocol Traffic

2026-09-16 · Muhammad Abdullah Sohail

General AI

The Model Context Protocol (MCP) standardizes communication between autonomous Artificial Intelligence (AI) agents and remote tools over Streamable HTTP. This shift introduces a class of machine-generated, authenticated, and high-frequency JSON-RPC traffic directly into enterprise networks. Enterprise network defenders…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.4

Token Efficient Task Execution via Application Behavior Modeling for Web Agents

2026-09-11 · Alexandru Ianta, Eleni Stroulia

Research Track B · General AI

The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-applicatio…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.0

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

2026-09-16 · Yu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng

General AI

Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

Lumen: Parameter-Efficient Alignment of Pretrained Vision and Language Encoders for Zero-Shot Computational Pathology

2026-09-15 · Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti, Javier Garcia-Baroja, Philipp Zens, Branislav Zagrapan, Yuri Tolkach, Martin D. Berger, Aurel Perren, Bastian Dislich, Inti Zlobec, Amjad Khan

General AI

Pathology vision-language models are commonly built by pretraining or fine-tuning large encoders on paired image-caption data. We asked whether a pathology vision-language model can instead be assembled by parameter-efficient alignment of frozen unimodal foundation models, leaving their pretrained representations untou…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.0

Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches

2026-09-15 · Vivek Kalyanarangan

General AI

When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan in which each query decides how many bits of each k…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

2026-09-16 · Toqeer Ehsan, Thamar Solorio

General AI

Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper, we present a multi-step framework to enhance the annotation quality of NER datasets by employing automated techniques. We propose a frequency-based i…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory

2026-09-16 · Michael M. Craig, Riley J. Hickman, Yingshan Ma, Rémi Piché-Taillefer, Christine Allen, Pauric Bannigan

General AI

Self-emulsifying drug delivery systems (SEDDS) can improve the oral bioavailability of poorly soluble drugs, but identifying high-performing formulations remains experimentally intensive. We present Andromeda 2, an agentic system that reasons over structured in-house experimental evidence and invokes computational and …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging

2026-09-16 · Johannes Kaiser, Florian Braunmiller, Daniel Rückert, Georgios Kaissis

General AI

AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) architectures currently fail to satisfy. Expert-based routing improves in-domain learning but may sacrifice cross-modal signals, which appear par…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

In-Context Robot Learning with VLM Agents

2026-09-16 · Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu

Research Track A · General AI

Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.2

LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys

2026-09-14 · Md Khalid Syfullah, Alvi Ataur Khalil

General AI

Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generaliza…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.0

Characterizing Network Centralization and Observability in the Remote MCP Ecosystem

2026-09-16 · Muhammad Abdullah Sohail

Research Track A · General AI

The Model Context Protocol (MCP) has emerged as the dominant interface for connecting autonomous agents to external data sources and execution environments. The ecosystem's transition from local process execution to remote Streamable HTTP deployments introduces unmeasured architectural and security constraints at scale…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

Affora: A Design System for Agent-Friendly Interfaces

2026-09-16 · Jin Gao

General AI

Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a design system that supports both readers while preserving visual freedom and familiar human workflows. Three controlled studies examine component imple…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

2026-09-16 · Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid

General AI

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense caption…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

2026-09-16 · Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski

General AI

World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, w…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.5

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence

2026-09-15 · Xingxuan Zhang, Gang Ren, Hao Yuan, Hao Zou, Hongze Tan, Hui Wang, Jianhao Song, Jiansheng Li, Jiayao Zhang, Jinghan Zhang, Kaifang Li, Lang Mo, Li Mao, Mingchao Hao, Nuo Xu, Rui Ding, Ruiji Zhang, Shuyang Li, Siyu Mei, Tianyang Zhang, Weiyang Mu, Yancheng Dong, Yongxian Wei, Yuan Xue, Yuanrui Wang, Yue He, Zijia Yang, Ziyun Li, Dongzhe Li, Fuqiang Wang, Jiandong Liu, Jiawei Chen, Jiaxin Du, Kaijie Cheng, Kehan Li, Lei Sun, Linjun Zhou, Ningbo Dai, Qi Wang, Renzhe Xu, Shaoxing Du, Shumeng Yang, Wang Lu, Wenjing Chu, Xiannan Huang, Xiaoyu Lin, Xing Ai, Xinyan Han, Xuanyue Li, Xuanyue Su, Xukun Zhang, Yan Lu, Yaxin Zhang, Yi Qin, Yifei Huang, Yihan Xu, Yongle Lv, Yuanyuan Jiang, Yushan Han, Peng Cui

Research Track A

We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of i…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM

2026-09-15 · Adam Zachary Wasserman, David Beauchemin

General AI

We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation) p…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

2026-09-16 · Pranaya Jajoo

General AI

Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy's value? We show that it can when the logger depends on history. For every horizon $H \ge 3$, we construct two POMDPs with at most two latent states per stage, three actions, and a common logger with …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

2026-09-16 · Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen

General AI

Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distingui…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.4

SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization

2026-09-13 · Zian Liu, Yiwen Hu, Zican Dong, Tian Xie, Wayne Xin Zhao, Yucheng Ding, Ran Tao, Bryan Dai

General AI

Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dy…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs

2026-09-16 · Mohan Shi, Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang, Eray Eren, Abeer Alwan

General AI

Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting them to domain-shifted speech, such as ch…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Flag Game: A Toy Model for Mechanistic Swarm Interpretability

2026-09-16 · Elizabeth Pavlova, Hidenori Tanaka

General AI

Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. To this end, we introduce the Flag Game, a toy model for st…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

2026-09-16 · Matteo Marchi, João Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada

General AI

Large Language Models (LLMs) are now routinely trained using synthetic data, since high-quality human data has been exhausted by the ever increasing needs of larger and larger models. However, recursive training on synthetic data frequently induces model collapse, a degenerative feedback loop where models progressively…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Track, Articulate, Act: Generating Articulation from Casual Human Videos

2026-09-16 · Jiaming Zhang, Homanga Bharadhwaj

General AI

Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes in object state. In this work, we study articulated objects such as doors, drawers, cabinets, laptops, ovens, and hinged containers that are ubiquitous in daily life and…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.3

Quantum Computing in Next-Gen Smart Grid Operations: A Comprehensive Review

2026-09-16 · Md Habib Ullah

General AI

The rapid proliferation of grid-edge distributed energy resources has significantly increased the operational complexity of modern power systems. Consequently, conventional computational techniques face growing scalability and computational-efficiency challenges in addressing large-scale optimization and control, uncer…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

2026-09-15 · Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang

General AI

Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Designing Grid-Aware Dynamic Specifications for Large Data Center Loads

2026-09-16 · Ashutossh Gupta, Vassilis Kekatos

General AI

As data center (DC) loads increasingly penetrate the power grid, there is an urgent need for grid operators to provide clear dynamic specifications to DC owners to ensure safe grid operation. To this end, we study two salient behaviors of large language model (LLM) training loads: abrupt ramps at job initiation and ter…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

2026-09-16 · Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa

General AI

Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in contact-rich tasks where appropriate contact forces are essential for success. In this work, we expl…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

EarStreAM: A Closed-Loop Earable System for Personalized Stress-Adaptive Meditation

2026-09-16 · Jonas Hummel, Luisa Faust, Elias Müller, Eva Bertog, Valeria Zitz, Marius Johannes Prill, Luca L. Bennardo, Luisa Weber, Tobias Röddiger, Michael Beigl

General AI

We present EarStreAM, a closed-loop earable system for stress-adaptive meditation that integrates in-ear physiological sensing with personalized, real-time intervention. Leveraging OpenEarable 2.0's multimodal sensing, EarStreAM continuously monitors physiological signals and detects elevated stress from heart rate and…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Objective vs. Search: Decomposing What Makes a Good Tokeniser

2026-09-16 · Ahmetcan Yavuz, Clara Meister, Tiago Pimentel

General AI

Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these …

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Position Anchor Tuning: Towards Efficient Adaptation of Pre-Trained Point Cloud Transformers

2026-09-16 · Zheng Liu, Xin Gao, Jinchao Zhu, Gao Huang

General AI

Parameter-efficient fine-tuning (PEFT) has recently emerged as a pivotal research direction for adapting pre-trained point cloud transformers to diverse downstream tasks. Although existing methods achieve excellent fine-tuning performance with high parameter efficiency, they ignore inference efficiency. To tackle this …

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.4

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

2026-09-13 · Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao, Shafiq Joty

General AI

Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in common use, and each grows differently: expert dis…

Review
pending
Role
unreviewed
Read
later
huggingface Score 5.4

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

2026-09-14 · Xingyang Li, Dongyun Zou, Shining Zhang, Jiacheng Chen, Haocheng Xi, Lvmin Zhang, Jun-Yan Zhu, Song Han, Zhekai Zhang, Yujun Lin, Muyang Li

General AI

Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entries, leaving typical e…

Review
pending
Role
unreviewed
Read
later