Research Paper Cockpit

Today Inbox

Fresh papers from the latest digest window that still need a decision.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-08-18.

Papers

62 visible entries

arxiv Score 29.5

Continual-learning rules shape representational drift

2026-08-17 · Yikai Si, Shanshan Qin

Research Track A

Lifelong learning requires acquiring new knowledge without erasing the old. Yet neural population codes for familiar stimuli and behaviors change over days and weeks. This coexistence of stable memory and changing internal codes may depend on how a learning system prevents forgetting. We therefore tested whether differ…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.3

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

2026-08-17 · Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao

General AI

Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks, prompt injection, backdoors, poisoning, or…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.3

TDD-Agent: Test-Driven Reasoning for Code Generation

2026-08-17 · Hongyue Yu, Kefan Li, Jiakun Li, Hongzheng Chai, Yuan Yuan, Rui He, Junyi Wei

General AI

Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which limits their ability to guide implementation and may introduce misleading…

Review
pending
Role
unreviewed
Read
now
arxiv Score 21.3

Geometry of Forgetting: Representation Flux in Continual Learning

2026-08-16 · Maksim A. Kazanskii

Research Track A · General AI

Catastrophic forgetting remains a fundamental obstacle to continual learning, where neural networks lose previously acquired knowledge while learning new tasks. Existing methods primarily mitigate forgetting through parameter regularization or experience replay, while the representation-space dynamics associated with f…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.6

Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents

2026-08-15 · Tianxin Wei, Zhan Shi, Minhua Lin, Bing He, Zewen Liu, Yisi Sang, Yuanchen Bei, Xuying Ning, Jiaru Zou, Ting-Wei Li, Xiao Lin, Yanjun Zhao, Chi Wang, Benoit Dumoulin, Dakuo Wang, Jingrui He, Hanqing Lu

Research Track B · General AI

Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.3

LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing

2026-08-17 · Ruoqi Shu, Xuhui Wang, Isaac Wang, Yanming Mai, Bo Wan

General AI

Financial document validation in production, such as payroll auditing, tax compliance, and loan underwriting, demands exceptional accuracy, consistency, and reproducibility under strict enterprise constraints. In practice, documents arrive with heterogeneous layouts and formats, semantically rich and context-dependent …

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.3

When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

2026-08-17 · Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu

Research Track A · General AI

Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code generation, driving their gradual evolution from text generation models into the core of agents capable of perceiving environments, invoking tools, and executing tasks. Traditional LL…

Review
pending
Role
unreviewed
Read
now
huggingface Score 20.0

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

2026-08-12 · Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng

General AI

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME leaderboard, suggesting that conventional single-turn video understanding tasks are becoming increasin…

Review
pending
Role
unreviewed
Read
now
huggingface Score 19.0

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

2026-08-16 · Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng

Research Track B · General AI

Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive

2026-08-16 · Brian B. Moser, Ahmed Anwar, Tobias Christian Nauen, Shishir Muralidhara, Federico Raue, René Schuster, Stanislav Frolov, Andreas Dengel

Research Track A

Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. Per-parameter looks more flexible than per-layer, but each layer's diagonal Fisher is a weak summary of its actual curvature, missing the top-eig…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.3

Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection

2026-08-15 · Yoon Gyo Jung, Jaewoo Park, Kuan-Chuan Peng, Seongdeok Bang, Octavia Camps

Research Track A · General AI

Greedy sampling produces a compact yet representative summary of normal data, which is essential for reliable anomaly detection that relies on measuring distance from normality. For continual anomaly detection where tasks arrive sequentially, extending greedy sampling is straightforward with unbounded memory through co…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.3

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

2026-08-17 · Reza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville

Research Track A · General AI

The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static memory throughout the entire sequence. Because early tokens face no compression pre…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.3

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

2026-08-17 · Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu

General AI

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations natu…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.3

LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents

2026-08-17 · Xingjun Wang, Gongsheng Li, Qi Fan, Yunlin Mao, Luyan Su, Yingda Chen

General AI

LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexe…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.3

PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation

2026-08-17 · Yuji Wang, Yuheng Chen, Teng Hu, Ran Yi, Yijia Hong, Han Feng, Weijian Cao, Chengjie Wang, Lizhuang Ma, Jiangning Zhang

General AI

Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across …

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.6

SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion

2026-08-15 · Yu He, Weikai Yang

General AI

Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgments, which may merge superficially related but behaviorally incompa…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.5

TRACER: Balancing Stability-Plasticity-Cognitivity Trilemma for LLM Enhanced Continual Recommendation

2026-08-17 · WooJoo Kim, HyunSik Yoo, JunYoung Kim, JaeHyung Lim, SeongKu Kang, HwanJo Yu

Research Track A

Continual recommendation aims to capture evolving user interests from streaming data but struggles with sparsity. LLM enhancers mitigate this with semantic knowledge, but naive integration creates a new conflict. We identify this as the Stability-Plasticity-Cognitivity (SPC) Trilemma, where generalized LLM semantic pri…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.3

AutoSR: Automatic Symbolic Regression by Searching Research States

2026-08-17 · Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang

General AI

We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numerically competitive expressions that imply very different behavior outsi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.3

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

2026-08-17 · Bingxin Xu, Yuzhang Shang, Emilio Ferrara

General AI

Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising recipe fr…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.3

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

2026-08-17 · Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen, Yang Zhang

General AI

Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. This can give strategically optimized pages visibility disproportionate to their authority or relevance and even make weak or false information appear well supported. Unlike con…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.8

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

2026-08-16 · Peng Chunyi, Xu Zhipeng, Yan Yukun, Liu Zhenghao, Yu Shi, Mei Sen, Sun Yubo, Zhang Yongheng, Zhou Jie, Gu Yu, Yu Ge, Sun Maosong

General AI

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual de…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.3

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

2026-08-17 · Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison

General AI

Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their variance reduction strategy. Group-relative methods like GRPO reduce gradient variance by sampling multiple rollouts per prompt, but provide only sequence-level credit. Training is also blocked by straggler rollouts, r…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.0

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

2026-08-14 · Yuhao Zhan, Bingxiang He, Zecong Tang, Chaojun Xiao

General AI

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchm…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.5

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

2026-08-17 · Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, Yijiang Li

General AI

Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that th…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse

2026-08-07 · Yingtao Ren, Ziyi Zhao, Yiwei Fu, Xiao Luo, Yu-Cheng Chang, Chin-Teng Lin

General AI

Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty

2026-08-17 · Yan Ma, Lizhuo Zhang

General AI

Multimodal large language models (MLLMs) are widely used for automated annotation, yet their per-class accuracy varies widely (e.g., 12%-98% across the 13 classes of three classroom sub-datasets) and is expensive to measure: evaluating one 27B MLLM on 5,416 validation images takes roughly 14 hours, whereas a frozen-CLI…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

2026-08-17 · Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang

General AI

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter

2026-08-17 · Vignesh Nagarajan, Sriram Venkatapathy

General AI

A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number into readable prose can quietly add claims the model never made. MIRROR is a research prototype built to separate those failures. It chains a multi-label classifier, a Grad-CAM loca…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

2026-08-17 · Minh-Ha Nguyen, Cathy Shyr

Research Track A · General AI

Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evalu…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

2026-08-17 · Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma

General AI

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic v…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

Quipu: A Governed Bitemporal Knowledge Graph Store

2026-08-17 · Steve Brown

General AI

Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware. These four defaults are individually conve…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

2026-08-17 · David Eric Austin, Kaheer Suleman, Jackie Chi Kit Cheung

General AI

Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natur…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.0

Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks

2026-07-28 · Kiran N. Kumar, Santhosh K. Saminathan

Research Track B · General AI

Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, B…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.0

Task-Anchored Representation Shaping for Pre-Trained Model-Based Continual Learning

2026-08-17 · Zhiming Xu, Huiyu Yi, Zhen-Hao Xie, Baile Xu, Furao Shen, Jian Zhao, Suorong Yang

Research Track A

Pre-trained models (PTMs) provide a strong foundation for continual learning by offering stable representations that facilitate lightweight adaptation to new tasks. However, adapting well to each task does not ensure reliable inference over all learned tasks. Since task boundaries are often artificial and semantically …

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.5

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

2026-08-17 · Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo

General AI

As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose \textbf{Ventor-QT…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.3

Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis

2026-08-17 · Reza Fayyazi, Michael Zuzak, Shanchieh Jay Yang

General AI

Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main criteria that must be met when using LLMs in cybersecurity, that is, trust in the generated outputs. As Agentic AI is in…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.5

Twin: Playing an Unknown Game with a Test-Time Digital Twin

2026-08-14 · Alexy Skoutnev, Kirill Acharya, Gaston Longhitano, Madeleine Udell, Kevin Ellis, Iddo Drori

Research Track A · General AI

We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditional approaches hand-engineer such models, one custom design per task. Each game hides its rules and goal, and our system…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 11.3

$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

2026-08-17 · Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou

General AI

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 11.3

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

2026-08-17 · Zi Haur Pang, Casey Kennington, Tatsuya Kawahara

General AI

Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 11.3

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

2026-08-17 · Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth

General AI

Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the s…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.8

SkillComposer: Learning Reusable Skills for Natural-Language Robot Programming

2026-08-14 · John Woods, Hasti Seifi

General AI

Natural-language interfaces can lower the barrier to programming robots, but existing systems struggle when users request complex tasks. While large language models (LLMs) perform well with simple commands, they often struggle to generate code for multi-step tasks, decompose high-level instructions, or reuse prior solu…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.8

Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach

2026-08-14 · Sheng Hong, Xuanqi Wang, Jiacheng Wang, Yuwei Wang

General AI

Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams. Off-the-shelf models struggle to filter noise from these streams and show limited structured-output reliab…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.5

Drive, Pack, Fly: The Travelling Thief Problem with Drone

2026-08-17 · Kabir Murjani, Abhay Sobhanan

General AI

In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onboard drone can offset this penalty by retrieving outlying items, thereby shortening the makespan and increasing operational profit. However, travel time remains load-dependent, and …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.5

Mechanical-microstructural correlation on SPS-fabricated NiTi alloy

2026-08-17 · Tadeáš Těhan, Jaromír Kopeček, Elizaveta Iaparova, Eduardo Alarcón, Sneha Samal

Research Track A

Mechanical properties and dynamical mechanical analysis were performed on compact Spark Plasma Sinter samples. It has been observed that the sample with less porosity reflects the behavior of superelasticity response. Other samples show failure during first cycles that may be due to porosity. Compaction of metallic pow…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.5

Two-Level Decorrelated Coded Modulation on the $D_4$ Lattice

2026-08-17 · Leopold Bertholet, Chloe Makdad, Stephen Mackes, Daniel Chew, Matthew Robinso

Research Track A

We propose \textit{two-level decorrelated coding} (TLDC), a novel coded modulation scheme for the $D_4$ lattice that combines Voronoi shaping with a two-stage decoding process to achieve lattice shaping and coding gains at low complexity. In TLDC, the decoded values of the first level allow the several random variables…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 10.0

Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

2026-08-02 · Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu Pan, Jason Eisner

General AI

Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--wher…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.3

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

2026-08-17 · Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Orăsan, Chrysoula Zerva, Ricardo Rei, Frédéric Blain, André F. T. Martins, Marco Turchi, Matteo Negri, Rajen Chatterjee, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya

General AI

Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe:…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.3

Model Hypnosis: Strong control of AI via additive subliminal effects

2026-08-17 · Enric Boix-Adsera, Benedict Tessler

General AI

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.3

SoftModel: A Neural Model That Grows Its Own Topology -- Governed Structural Growth for Continual In-Service Learning

2026-08-17 · Zhoumin Xie

General AI

Today, a neural system is almost always used in two phases -- trained, then deployed -- and in that regime it freezes twice: training ends, and the topology itself was never a degree of freedom. We take the opposite premise as an axiom -- total plasticity: no part of a model, including its structure, is ever frozen -- …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA

2026-08-15 · Juseok Jeon, Ramy E. Ali, Doyun Kwon, Myungbeom Her, Jinhwi Kim, Jinhyun So

General AI

Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of large language models, but its factorized parameterization creates a tension between accurate aggregation of local updates and continuity of locally optimized factors. Factor-wise aggregation incurs aggregation mismatch but better preserves factor co…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.5

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

2026-08-17 · Parsa Mazaheri, Kasra Mazaheri

General AI

Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a comple…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.3

Evaluating Beyond the Screen: Collective Assessment of AI-Generated Business Plans with Resource-Constrained Entrepreneurs

2026-08-17 · Qi Zhao, Marjory Pineda, Ketul Chhaya, Aakash Gautam, Yasmine Kotturi

General AI

Entrepreneurs increasingly use end-user generative AI technologies such as ChatGPT for high-stakes documents like loan applications and business plans, where AI-generated errors---a wrong price, a fabricated product---can affect loan or funding outcomes. Current approaches to supporting evaluation of AI-generated text …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.3

Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation

2026-08-17 · Suraj Yadav

General AI

Universal visual representations require adaptation mechanisms that adapt across heterogeneous domains without fragmenting knowledge into domain-specific modules. Parameter-efficient fine-tuning adapts frozen visual foundation models efficiently, but standard low-rank adapters use a fixed subspace for all inputs, which…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.8

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

2026-08-15 · Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi

General AI

Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We addre…

Review
pending
Role
unreviewed
Read
later
huggingface Score 7.5

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

2026-08-17 · Kwan Yun, Serin Yoon, Sunjin Jung, Jung Eun Yoo, Inyup Lee, Junyong Noh

General AI

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging rece…

Review
pending
Role
unreviewed
Read
later
arxiv Score 7.3

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

2026-08-17 · Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi

General AI

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-establishe…

Review
pending
Role
unreviewed
Read
later
arxiv Score 7.3

Q-based Variational Inverse Reinforcement Learning

2026-08-17 · Ondrej Bajgar, Peter Tisnikar, Alessandro Abate, Konstantinos Gatsis, Maike Osborne

General AI

The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is often infeasible. Inverse reinforcement learning (IRL) addresses this challenge by inferring preferences, represented as reward functions, f…

Review
pending
Role
unreviewed
Read
later
arxiv Score 7.3

Sample Complexity of Peer Prediction

2026-08-17 · Abdellah Aznag, Robin Bowers, Rachel Cummings, Jason Hartline, Matthew vonAllmen, Bo Waggoner

General AI

Peer prediction seeks to incentivize agents to truthfully report an observed signal by rewarding joint sets of reports without observing a ground truth. Following the generalization of information-theoretic mutual information introduced in Kong and Schoenebeck (2019), we call a function of a joint distribution over sig…

Review
pending
Role
unreviewed
Read
later
arxiv Score 7.3

The ultimate carbon cost of a ChatGPT query

2026-08-17 · Paul Kron

General AI

This paper reviews and combines findings from the fields of product and life-cycle analysis [36, 38], the usage of modern transformer- based large language models (LLM) [6], as well as on greenhouse gas emissions and the ultimate cost of their subsequent consequences for future generations [2]. In this paper, it is sho…

Review
pending
Role
unreviewed
Read
later
arxiv Score 7.3

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

2026-08-17 · Benjamin Belay

General AI

A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled arch…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.8

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form

2026-08-15 · Parsa Mazaheri

General AI

Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Anchor-Regularized Adaptation for Generalizable AI-Generated Image Detection with DINOv3

2026-08-15 · Hyeongjun Choi, Juhun Lee, Davide Cozzolino, Luisa Verdoliva, Simon S. Woo

General AI

Recent works in AI-generated image detection have shown that careful training data alignment can improve generalization by removing spurious correlations. However, linear probes on frozen DINOv3 representations achieve remarkably strong performance even when trained on misaligned datasets. Motivated by this result, we …

Review
pending
Role
unreviewed
Read
later