arxiv
Score 25.5
2026-08-27 · Dezheng Han, Anbang Zhang, Zhihao Zhu, Shuaishuai Guo
Research Track A · General AI
To mitigate catastrophic forgetting in downstream continual learning (CL) for large language models (LLMs), existing methods typically constrain parameter updates or introduce task-specific adaptation modules. However, these methods often rely on explicit task boundaries during training, limiting their applicability to…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 20.5
2026-08-26 · Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang
General AI
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem. Full-library prompting preserves coverage at high context cost, vector retrieval returns compact neighborhoods but treats skills as independ…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 20.5
2026-08-27 · Yibo Feng
Research Track A · General AI
Rehearsal-free class-incremental learning (CIL) with LoRA adapters remains challenging because the low-rank subspaces updated across tasks evolve without geometric control, causing unstable shared representations and repetitive collapse of task-specific updates into previously occupied directions. We introduce Geo-LoRA…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 20.3
2026-08-27 · Wendong Li, Jochen Garcke
General AI
Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. We propo…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 17.5
2026-08-22 · TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin, Yongdong Luo, Yibo Hu, Meiguang Jin, Junfeng Ma, Weihang Pan, Jiaxin Zhao, Zulong Chen
General AI
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, prompts, and tools can be updated independently of model wei…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 17.5
2026-08-27 · Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur
Research Track A · General AI
Robotic and edge intelligence systems operate in dynamic environments where data arrives continuously, requiring models to adapt while preserving previously learned knowledge under strict memory and energy constraints. While parameter-efficient fine-tuning has shown promise for continual learning with vision transforme…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 17.3
2026-08-27 · Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
General AI
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.3
2026-08-27 · Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li
General AI
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordi…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.5
2026-08-27 · Maximilian Du, Zhanyi Sun, Chen Xu, Paarth Shah, Masha Itkina, Shuran Song
Research Track A · General AI
Robot policies deployed in the wild should have the capability to continually learn new tasks without forgetting existing behaviors. A common approach to combat such catastrophic forgetting is to train on new task data with a replay buffer of previously learned task data. Although this buffer is commonly sampled random…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 15.5
2026-08-27 · Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
General AI
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe edit…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.3
2026-08-27 · Hengyuan Xu, Wei Cheng, Yumeng Ji, Xuanyang Zhang, Xianfang Zeng, Gang Yu, Xingjun Ma
General AI
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce \textbf{Aphanta}, an automated task-discovery and closed-loop diag…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.3
2026-08-27 · Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen
General AI
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseu…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.0
2026-08-27 · Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofrè, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz
Research Track A · General AI
The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an information, economic and power asymmetry between developers and the diverse user base of modern AI. Recent public discourse acknowledges this concern, calling for SovereignAI (an …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.3
2026-08-27 · Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam, Nadir Durrani
General AI
Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yield consistent judgments across modality (text vs. speech) and language (English vs. Arabic). We intr…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 13.5
2026-08-27 · Yang Xiao, Yusong Sun, Haoyi Wu, Wenyang Hui, Wen Da, Zhaokai Luo, Mu Chuan, Yao Hu, Wenjie Li, Chengyue Jiang
General AI
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should inst…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.3
2026-08-27 · Yufan Wu, Yinghui He, Zhengyi Hu, Lang Wei, Ruichen Li, Qifan Yang, Ting Zhu
Research Track A · General AI
Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoni…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.3
2026-08-27 · Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung
General AI
Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.3
2026-08-26 · Wei Sun, Marie-Francine Moens
General AI
Large language models (LLMs) exhibit uneven multilingual performance, especially when dealing with low-resource languages. Inference-time intervention offers a lightweight way to improve cross-lingual transfer by modifying the hidden states produced by the LLMs during the forward pass, without updating model parameters…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.3
2026-08-27 · Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
General AI
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approa…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.3
2026-08-27 · Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng
General AI
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code re…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 11.5
2026-08-24 · Abhilash Nandy, Rahul Seetharaman, Aman Bansal, Rounak Saha, Manav Nitin Kapadnis, Millon Madhur Das, Pawan Goyal, Niloy Ganguly
Research Track A · General AI
Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and implicit relationships across image and text mod…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.3
2026-08-27 · Xingyu Shen, Huishuai Zhang, Peng Li, Yinchun Wang, Dongyan Zhao
General AI
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-par…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.3
2026-08-27 · Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao
Research Track A · General AI
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert tas…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 10.5
2026-08-26 · Youtian Lin, Yikang Yang, Zhanpeng Hu, Mengqi Zhou, Feihu Zhang, Xun Cao, Jiaheng Liu, Yao Yao
General AI
Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where a machined object should be sharp, it carries no part decomposition, and it exposes no parameter a user could edit. To address this, we explore the paradigm of 3D shape as code, leveraging and scaling t…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.3
2026-08-27 · Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jianxing Yu, Zibin Zheng
General AI
To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task success does not guarantee high-quality supervision: successful trajectories…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.3
2026-08-27 · Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
Research Track A · General AI
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 9.3
2026-08-26 · Harshavardhan Adepu, Li Zhang, Sanjiv Kumar, Vikas Singh
General AI
Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $d$ is the model's hidden dimension and $r$ is the rank. Our proposal…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.3
2026-08-27 · Yisen Xi
General AI
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply. We present Persona-Execution Separation (PES): persona and execution re…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 9.3
2026-08-27 · Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, Tu Vu
General AI
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 9.0
2026-08-27 · Ke Shu, Kira Hinderks, Eetu Mäkelä, Mikko Tolonen
Research Track A · General AI
This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses on a candidate set centered on essays by eighteenth-century Scottish philosopher David …
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 8.5
2026-08-26 · Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang, Yixing Ma, Zhenglin Wan, Kaipeng Zhang, Wangbo Zhao, Yang You
General AI
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, comp…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 8.5
2026-08-27 · Shiyi Zhang, Mushui Liu, Yunze Tong, Wanggui He, Siyu Zou, Jinlong Liu, Yunlong Yu, Jian Song, Hao Jiang, Pipei Huang, Bo Zheng
General AI
On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory signals, has achieved significant success in Large Language Models (LLMs) and has recently been adapted to flow matching models. However, this paradigm suffers from two major issues: First, training a sep…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.3
2026-08-27 · Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
General AI
Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Supp…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.3
2026-08-27 · Peiling Yi
General AI
Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information is framed, omitted, contextualised, and communicated. Yet less research has focused on how such misleadingness arises and shapes the interpretations formed by readers. To address th…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.5
2026-08-26 · Kexin Sun, Qiang Liu, Minfu Feng, Mingchao Cai
Research Track A
Physics-Informed Neural Networks (PINNs) have recently gained considerable attention as a mesh-free framework for solving partial differential equations. Nevertheless, their performance deteriorates when applied to strongly coupled multiphysics systems, such as Biot's consolidation model, due to severely ill-conditione…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.3
2026-08-27 · Kechen Liu, Ola Shorinwa
General AI
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment actio…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.3
2026-08-27 · Prachi Chaturvedi, Shahnawaz Ahmad, Ehsan Nowroozi, Muhammad Waqas, George Loukas, Alireza Jolfaei, Lucas Cordeiro, Pierre Dantas
General AI
Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, and through what mechanisms can responsibility be traced, explained,…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.3
2026-08-27 · Maria Djurić, Ralph Schönrich
General AI
The vertical structure of our Galaxy has commonly been assumed to follow a pseudo-isothermal distribution. However, there is no \textit{a priori} reason to expect this form to arise from scattering by giant molecular clouds (GMCs), since GMCs are confined to a narrow layer around the Galactic midplane and therefore do …
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 6.5
2026-08-27 · Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
General AI
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this di…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.3
2026-08-27 · Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal
General AI
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlled, artifact-backed benc…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.3
2026-08-27 · Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
General AI
Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \textit{de novo} generation of product molecules or through heuristic graph edits that operate directly on molecular topology. We introduce MAELLE (\textbf{M}ech\textbf{A}nistic \textbf…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 5.5
2026-08-25 · Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li
General AI
Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied …
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.3
2026-08-27 · Hai-tao Yu, Nan Min, Zheng Fang, Hongyu Zhan, Yusen Tan, Yuhan Wang, Jun Xia
General AI
Inferring molecular structures from multimodal spectroscopic measurements requires integrating complementary yet highly heterogeneous signals. However, the common paradigm of directly concatenating multispectral sequences can exhibit anomalous performance degradation, primarily due to pronounced heterogeneity and the r…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 5.3
2026-08-27 · Mahmud Hasan Saikot, Sydney Spiegel, Sudheera Akalanka Kariyawasam, Andrew Stefka, Josh Chrisler, Jianguo Zhao
General AI
Robots that can change their morphologies and behaviors for different tasks and environments hold great promise for adaptable, multifunctional systems. Modular reconfigurable robots (MRRs) can achieve such functionalities by docking and rearranging individual units, but most rely on rigid modules that lack structural c…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 5.3
2026-08-27 · Eleni Tselepi, Cristian Sestito, Shady Agwa, Themis Prodromakis
General AI
Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal models has made them widely associated with text-to-image applications running on large datacentres. However, vision generative models are equally needed in applications th…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 4.3
2026-08-27 · Ting Yan
General AI
AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 4.3
2026-08-27 · Hiep V. Dang, Antonios Mamalakis
General AI
Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity. We introduce SimCast-S2S, a generative …
- Review
- pending
- Role
- unreviewed
- Read
- later