Research Paper Cockpit

Daily Digest - 2026-08-14

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-08-18.

Papers

45 visible entries

huggingface Score 32.0

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

2026-08-13 · Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen

General AI

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another lin…

Review
pending
Role
unreviewed
Read
now
arxiv Score 27.8

Intern-S2-Preview: Scientific Agentic Foundation Model

2026-08-13 · Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv, Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun, Yu Sun, Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang, Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu, Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang, Chao Zhang, Chen Zhang, Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou, Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou

General AI

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support …

Review
pending
Role
unreviewed
Read
now
arxiv Score 23.8

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

2026-08-13 · Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi

Research Track A · General AI

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to asse…

Review
pending
Role
unreviewed
Read
now
arxiv Score 23.8

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

2026-08-13 · Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck

General AI

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) ag…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.5

SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents

2026-08-12 · Ruitao Wang, Yuwen Hao, Menglin Yang

Research Track B · General AI

Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data synthesis methods reduce manual annotation, but they still face two key limitations: they often fail to cover the full functionality of a website, and without sufficient website prior…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.5

Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

2026-08-13 · Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu

Research Track B · General AI

Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an …

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

2026-08-13 · Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu, Jin Yao, David I. Inouye, Jing Gao, Danhua Guo, Burhan Yaman

General AI

Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. …

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation

2026-08-13 · Dongfang Li, Zixuan Liu, Junmai Wang, Jiahe Huang, Fuhao Li, Bonian Jia, Baotian Hu, Min Zhang

Research Track A · General AI

Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to extract, summarize, or update memories. This design makes memory construction increasingly costly as conversations grow…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.8

When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory

2026-08-13 · Ruizhe Li, Licheng Zhang, Benfeng Xu, Mingxuan Du, Zheren Fu, Weidong Chen

Research Track A · General AI

Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ask how much of that benefit comes from the structure itself, rather than from competent retrieval over the raw history.…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.0

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

2026-08-13 · Yuxuan Zhang, Haozhong Xiong, Yubo Huang, Jiayi Song, Jinpeng Yu, Haofan Wang, Jiaming Liu, Ruihua Huang, Liwei Wang

General AI

Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding …

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models

2026-08-13 · Mohammed Ayman Habib, Rylan Hart, Morteza Fayazi

General AI

Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert intuition. Among recent developments, LLMs have introduced a promising approach by bringing natural language reasoning to circuit design tasks. The majority of conventional LLM-bas…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.5

Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning

2026-08-13 · Zeyang Zhang, Tieliang Gong, Junyan Lu, Weizhan Zhang

Research Track A · General AI

Plasticity loss has emerged as a critical challenge in continual learning that significantly hinders the acquisition of sequential tasks. While optimizing activation designs offers a potential solution, current fixed-form functions suffer from an inherent spectral bias towards low-frequency variations, whereas learnabl…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

2026-08-13 · Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ruizhu Zhou, Lingcheng Meng, Lining Hu, Ting Liu, Yuzhuo Fu

General AI

Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code presents a greater challenge: a model must jointly recover visual structure, ground the requested change, generate compilable code, and preserve…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.0

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

2026-08-11 · Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li

General AI

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.0

Mind the Context: Continual Learning of Socially Appropriate Robot Actions via Environmental-Social Disentanglement

2026-08-13 · Rafal Robert Karpinski, Fethiye Irmak Dogan, Nikhil Churamani, Yiming Luo, Maartje M. A. de Graaf, Davide Dell'Anna, Hatice Gunes

Research Track A · General AI

Social robots are expected to operate across diverse environments, where similar arrangements can imply different socially appropriate actions, e.g., starting a conversation may be acceptable in a crowded home but disruptive in an office meeting. Because such norms and environments cannot all be anticipated in advance,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

2026-08-12 · Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar

Research Track A · General AI

Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation. Additionally, surfacing reasoning mistakes t…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

2026-08-13 · Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li

General AI

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to dr…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

2026-08-13 · Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech

General AI

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is train…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

2026-08-07 · Tao Feng, Fangxu Yu, Haozhen Zhang, Zhongjie Dai, Liangqi Yuan, Zijie Lei, Weizhi Zhang, Kunlun Zhu, Haodong Yue, Keyang Xuan, Ge Liu, Jiaxuan You

General AI

No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a s…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

2026-08-13 · Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook

General AI

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context pass…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

2026-08-13 · Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

General AI

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems …

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

2026-08-13 · Zixuan Lan, Yanhong Li, Jiawei Zhou

Research Track A · General AI

Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative sl…

Review
pending
Role
unreviewed
Read
now
huggingface Score 13.0

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

2026-07-31 · Xinyan Guan, Jiali Zeng, Chunlei Xin, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, Fandong Meng

General AI

Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this futile reasoning phenomenon through systematic analysis, revealing universal capability overreach and…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.0

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

2026-08-13 · Jiaqian Li

Research Track A · General AI

Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

Vero: Can AI Agents Build Formally Verified Software Repositories?

2026-08-13 · Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song

General AI

AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated software. Existing …

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.5

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

2026-08-13 · Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel

Research Track A · General AI

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S.…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

2026-08-13 · Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

General AI

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically eval…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.0

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

2026-08-13 · Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao

General AI

Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.0

Integration-First Structural Coverage for Embedded Software:Trace-Based Evidence, Hybrid Runtime Analysis, and Cross-Variant Consolidation

2026-08-13 · Alexander Weiss, Albert Schulz, Michael Wittner

Research Track A

Structural coverage is widely used as evidence that testing is complete, yet in embedded projects it is predominantly collected at unit level, simply because that is where instrumentation and observability are inexpensive. This produces a mismatch. The most representative completeness signal would come from integration…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles

2026-08-13 · Md Wasiul Haque, Sagar Dasgupta, Mizanur Rahman, Md Rayhanur Rahman

General AI

Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming exploitability requires executable test artifacts that are difficult …

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.8

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

2026-08-13 · Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi

General AI

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skati…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.0

Distributed and Dynamic Hub Network Operation Planning in a Hyperconnected Less-Than-Truckload Operating System

2026-08-13 · Tiankuo Zhang, Jihye Jung, Paria Nourmohammadi, Benoit Montreuil, Alan Erera, Sahrish Jaleel Shaikh

Research Track A · General AI

The less-than-truckload (LTL) industry plays a vital role in enhancing the efficiency and sustainability of logistics systems, as LTL shipments offer greater consolidation opportunities than full-truckload shipments. Despite of this flexibility, the average cost of LTL shipments remains considerably higher due to less …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

CAPRI: Contract-Aware Proof Repair for Isabelle

2026-08-13 · Jim Woodcock, Gabriel Leite, Augusto Sampaio, Ran Wei

General AI

We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM changed only what the developer authorised. We present CAPRI, a contract-aware repair workflow in which Isabelle checks the proof and an independe…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective

2026-08-13 · Nestor R. Barraza, Gabriel Pena

General AI

Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds. In this work we e…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

QuoteBench: How Matched Scores Can Hide Command-Path Failures

2026-08-13 · Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang

General AI

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks fro…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

TabSOM: A tabular-to-image encoding method based on self-organizing maps

2026-08-13 · David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara, Francisco J. Lara-Abelenda, Luis Zhinin-Vera, Diego H. Peluffo-Ordóñez

General AI

Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.g., t-SNE…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.4

Full-bandwidth transformer

2026-08-09 · Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford

General AI

Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.4

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

2026-08-10 · Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou

General AI

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

2026-08-13 · Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu

General AI

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences betwe…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

2026-08-13 · Xingqi Cui, Chieh-Jan Mike Liang, Ziang Tang, Jiarong Xing, Haoran Qiu

General AI

Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs). Autoscaling is the key mechanism for cluster resource management, yet a basic system design question is open for serving LLMs: what sho…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.5

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

2026-08-13 · Ebenezer Tarubinga

General AI

Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.0

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

2026-08-12 · Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo

General AI

We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving ris…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.8

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

2026-08-13 · Weihan Meng, Hongzhu Guo, Yi Jing, Dewen Liu, Zijun Yao, Xiaozhi Wang, Lei Hou, Juanzi Li

General AI

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from co…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.0

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

2026-08-13 · Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang

General AI

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSw…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.8

Joint Communication-Control Strategy Optimization with Partially Nested Information Structures: The Linear-Quadratic Case

2026-08-13 · Haoyi You, Kaiqing Zhang

General AI

In this paper, we formalize a joint communication-control strategy optimization (JCCO) problem in multi-agent linear systems with quadratic costs, under the common-information-based (CIB) framework from decentralized stochastic control. For computational tractability, we focus on such JCCO problems with partially neste…

Review
pending
Role
unreviewed
Read
later