arxiv
Score 27.5
2026-09-05 · Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong
Research Track A · General AI
Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, a…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 24.5
2026-09-07 · Zheyuan Zhang, Alvin Zhang, Daniel Khashabi, Tianmin Shu
Research Track A
Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training …
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 21.0
2026-09-04 · Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
General AI
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introd…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 20.0
2026-09-04 · Phuong Tuan Dat, Phuong Khai Minh, Tran Huy Dat
Research Track A
Fully fine-tuning self-supervised learning (SSL) speech models for downstream tasks is computationally prohibitive, and existing parameter-efficient fine-tuning approaches predominantly rely on MLP-based adapters whose fixed activation functions limit their representational expressiveness under tight parameter budgets.…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.5
2026-08-31 · Bowei He, Xiaokun Zhang, Meng Ding, Xue Liu
Research Track B · General AI
Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step, but they treat the skill library as a flat or two…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.4
2026-09-08 · Gabriel J. Perin, Lucas Boscaini, André Araujo, Nina S. T. Hirata
Research Track A · General AI
Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.0
2026-09-06 · Zirui Shang, Xin Shu, Yang Liu, Zhi Gao, Xinxiao Wu, Lifeng Fan
Research Track A · Research Track B · General AI
Continual learning is a crucial capability for Graphical User Interface (GUI) agents to adapt to evolving applications while retaining knowledge acquired from previous applications. Such application streams pose a challenging knowledge modeling problem: new applications often share underlying knowledge with past ones, …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.2
2026-09-08 · Baojie Chen, Zijun Jia, Jing Zhong
Research Track A · General AI
VLMs have shown promise for autonomous driving, yet still suffer from hallucination, weak spatio-temporal perception, and limited generalization. Recent methods improve reasoning and decision-making through CoT explanations, retrieval-augmented generation or the static injection of tool outputs. Although these mechanis…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.2
2026-09-08 · Boyu Yang, Jiazheng Sun, Zilong Lu, Zhi Qiu, Xin Peng, Jun Zheng
General AI
Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences and task knowledge across extended interactions. Conventional retrieval mechanisms optimize semantic compatibility rather than downstream utility, frequently introducing outdated, misleading, or conflicting evide…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.2
2026-09-08 · Yuyang Huang, Bobo Li, Jiajia Song, Yuzhe Ding, Chong Teng, Fei Li, Donghong Ji
General AI
Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating core claims. However, manually navigating the growing volume of scientific literature is increasingly difficult, prompting reliance on automatic citation recommendation. While modern retrieval-augmented architectu…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 17.8
2026-09-03 · Qi Liu, Qinzheng Wang, Yiming Bie
Research Track A · General AI
As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-evolving agent built around the Simulation of Urb…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 17.2
2026-09-08 · Yixuan Liu, Zilong Zhen, Yin Wu, Yi Li
General AI
As Large Language Model (LLM) agents increasingly automate offensive operations across the cyber kill chain, their efficacy in complex local post-exploitation tasks remains inadequately quantified. Among these, Linux privilege escalation is a key step between initial access and full system compromise. However, existing…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 17.0
2026-09-07 · Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann
Research Track A · General AI
The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly on edge hardware using local user data. However, this shift requires optimization of deep learning training on resource-…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.5
2026-08-31 · Jianwei Zhang, Sihan Cao, Pengcheng Zheng, Ya Wen, Pei Ke, Kuien Liu, Shen Gao, Wei Dong, Yang Yang, Chaoning Zhang
Research Track B · General AI
Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners must rely on lightweight local models that evolve p…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.2
2026-09-08 · Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Uthayasanker Thayasivam, Kamal Premaratne
General AI
Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs ext…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.2
2026-09-08 · Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang
General AI
Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in whi…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.8
2026-09-07 · Kai Guo, Chuanbin Liu, Peng Hu, Hao Wang, Xi Peng
Research Track A · General AI
Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. However, existing methods typically assume that the set of modalities per task is predefined and fixed. In this paper, we investigate a more …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.5
2026-09-06 · Keyu Lin, Fei Ye, Qihe Liu, Shijie Zhou, Jiguo Yu
Research Track A
Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, a…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-09-08 · Jaewoo Lim, Sungbok Shin, Sanghyun Hong
General AI
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.8
2026-09-07 · Daniel Gomm, Maarten de Rijke, Madelon Hulsebos
General AI
Democratizing access to the knowledge held in large corpora of tables such as data lakes is emerging as a central research challenge. Research in this space is advancing and broadening in scope, increasingly supplying the components to satisfy a person's insight need end-to-end. Yet these efforts remain fragmented acro…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 14.4
2026-09-08 · Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddysun, Steveyves, Zhou Zhao, Bryanytian
General AI
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enab…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.2
2026-09-08 · Yuxing Lu, Yicheng Chen, Shanchan Wu, Sercan Ö. Arık
General AI
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order, and under which conditions. As trajecto…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.8
2026-09-07 · Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu
General AI
Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). Unlike direct prompt injection, IPI hides adversarial instructions in untrusted tool outputs and can covertly alter the execution of a legitimate task. Defenses trained on fixed attacks may fail as an attacker changes its strategy, inject…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-09-08 · Mar Gonzàlez I Català, Haitz Sáez de Ocáriz Borde, Davide Murari, Carola-Bibiane Schönlieb, Pietro Liò, George Montañez
General AI
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves ov…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-09-08 · Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen
General AI
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective sup…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-09-08 · Jingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun, Hang Zhang, Jianhua Zhu, Yufeng Wang
General AI
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. F…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-09-08 · Wenbo Gao, Zhaomou Song, Zhiyuan Ji, Renxi Liu, Xing Li, Xianzhi Yu, Xiaoguang Li, James Chung-wai Cheung, Weizhe Lin, Yaoyuan Wang
Research Track A · General AI
Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without s…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-09-08 · Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler, Manuel Schaaf, Markus J. Hofmann
General AI
This study examines whether individual text corpora (ICs) from search histories can be used to simulate individual knowledge. We collected ICs from 316 adults, who answered 36 multiple-choice knowledge items, and compared several large language models (LLMs) on this task, of which only Qwen3-1.7B proved viable. After t…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-09-08 · Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie
General AI
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of …
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 13.0
2026-09-07 · Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He
General AI
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.8
2026-09-05 · Huiyi Wang, Daijiao Liu, Lina Yao, Dong Gong
General AI
Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module he…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.2
2026-09-08 · Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao
General AI
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create f…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.0
2026-09-04 · Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi
General AI
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression metho…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.0
2026-09-07 · Chia-Hui Chen, Shih-Ying Yeh, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
General AI
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.8
2026-09-02 · Suyoung Lee, Myungsub Choi
General AI
Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instr…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.4
2026-09-08 · Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun
Research Track A · General AI
Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet convention…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-09-08 · Zhan-Lun Chang, Dong-Jun Han, Seyyedali Hosseinalipour, Mung Chiang, Christopher G. Brinton
General AI
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-09-08 · Xilin Wang, Guoxi Zhang, Hongming Xu, Zhuofan Zhang, Tianxu Wang, Lifeng Fan
General AI
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exploration, an LN agent must consolidate experience from earlier stages and reuse it in later stages, often through persistent scene represent…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-09-08 · Behdad Khodabandehloo, Andrea Caraffa, Davide Boscaini, Fabio Poiesi
General AI
Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving from explicit 3D models to multi-view object captures to single reference images. We take this progression to its extreme by introducing prior-free relative 6D pose estimation, which lifts the assumption of kn…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-09-08 · Tim Tomashevskiy
General AI
Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before they lead to unsafe behavior. Existing approaches typically rely on safety constraints defined at design time or updated reactively during execution, assuming that such constraints remain valid over time. Howeve…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-09-08 · Min Zeng, Yuzhou Liu, Zhenyu Cao, Hanxiu Chen, Heng Li, Caiquan Liu, Yafei Wen, Xiaoxin Chen
General AI
High-quality tool-use data is critical for training language models to interact effectively with external tools. However, existing synthetic approaches typically follow a generate-then-filter paradigm with static post-hoc verification, often yielding inefficient data with imbalanced feature distributions. We propose To…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.8
2026-09-03 · Nivedita Singh, Alsharif Abuadbba, Yansong Gao, Surya Nepal, Hyoungshick Kim
Research Track B · General AI
Large language models (LLMs) are becoming integral to web applications and browser agents, transforming online interactions while introducing new attack vectors and reshaping longstanding web vulnerabilities. Classical threats such as cross-site scripting (XSS) can be amplified through LLM-mediated interactions, while …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 10.5
2026-09-04 · Guanglong Sun, Kanglei Zhou, Liyuan Wang, Qi Cheng, Hongwei Yan, Shuang Cui, Hang Su, Jun Zhu, Yi Zhong
Research Track A
General continual learning (GCL) aims to learn from evolving data streams without task identities, explicit boundaries, or repeated access to previous data, making it a realistic yet challenging setting for continual intelligence. Although pretrained models (PTMs) provide rich prior knowledge for addressing the limited…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 10.5
2026-09-07 · Vitor Crista, Afonso Lourenço, Diogo Martinho, Goreti Marreiros
Research Track A · General AI
Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency. We propose \textit{HINT}, a hierarchical inference framework that combines edge-based…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.2
2026-09-08 · Mariia Drozdova, Stéphane Liem Nguyen, François Fleuret
General AI
Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph c…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.2
2026-09-08 · Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu
General AI
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this g…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 9.8
2026-08-12 · Ayushman Bhattacharya, Nihal Gazi
Research Track A · Research Track B · General AI
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers grounded in live web results. We developed OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference. Its local search, caching, ses…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 9.4
2026-09-08 · Nhat-Tan Bui, Varshini Elangovan, Arun Reddy Anugu, Sreyas Mohan, Wei Ye, Dilin Wang, JQ Huang, Rakesh Ranjan, Aviral Chharia, Fernando De la Torre
General AI
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.2
2026-09-08 · Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur
General AI
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harnes…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.2
2026-09-08 · Jiacheng Xu, Feng Chen, Xiuneng Xu, Bo An
General AI
Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To m…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.2
2026-09-08 · Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen, Yang Liu
General AI
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.8
2026-09-04 · Maryam Fakhari, Mehran Safayani
General AI
Cryptocurrency markets exhibit extreme volatility and non-stationary dynamics that challenge conventional forecasting methods. Although Large Language Models (LLMs) have shown promise for time series forecasting, the combined effects of adaptation choices remain largely unexplored in financial settings. This study intr…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.2
2026-09-08 · Akhil Ajikumar, Mahya Qorbani, Sakib Reza, Sean Andrist, Mohsen Moghaddam
General AI
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They …
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 8.0
2026-08-29 · Yucheng Du, Xiyang Hu
General AI
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evaluating (1).startswith("1"), instead of abstaining. We ask whether this failure reflects missing recognition or failed routing from recognition to abstention. Across instruction-tuned models from 1.7B to 70B param…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 8.0
2026-09-06 · Mingwei Li, Yi Yang, Hehe Fan
General AI
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth no…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.8
2026-09-07 · Xuechao Zou, Yi Zhou, Kai Li, Shun Zhang, Yuhui Chen, Congyan Lang, Junliang Xing
General AI
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitt…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.2
2026-09-08 · Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
General AI
Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation,…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.2
2026-09-08 · Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov, Özgün Turgut, Michelle Espranita Liman, Lisa Steinhelfer, Rickmer Braren, Daniel Rueckert
General AI
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inhe…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 7.0
2026-09-07 · Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos
General AI
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rat…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 7.0
2026-09-07 · Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao
General AI
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tigh…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-09-08 · Giordano De Marzo, Nicola Albore, David Garcia
General AI
In June 2026, thousands of AI agents found that a small public wiki would accept edits from inside their sandboxes, and started using it to help one another pass a timed test. Each agent lived for about an hour and remembered nothing afterwards. Nobody asked them to cooperate, and the wiki had not been built for them. …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-09-08 · Yunpeng Xu, Kun Zheng
General AI
Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-09-08 · Raito Kiya, Satoki Ohashi, Kosuke Sato, Go Kamoda, Ryosuke Takahashi, Yuji Yamamoto, Daiki Shiono, Keisuke Sakaguchi, Goro Kobayashi
General AI
Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently co-occur, and MAs can pose challenges for low-bit quantization. In this study, we analyze the factors underlying AS and MAs that emerge at t…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-09-08 · Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, Dhruv Shah
General AI
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjus…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 5.4
2026-09-08 · Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai
General AI
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp a…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-09-08 · Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao
General AI
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smooth…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.2
2026-09-08 · Minsik Jeon, Jay Karhade, Deva Ramanan, Shubham Tulsiani
General AI
We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is able to reliably infer dense per-point 3D trajectories across multi-hundred-frame videos, unlike existing 4D methods that are limited to short input windows of at most a few dozen frames. A key innovation that ena…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.0
2026-09-07 · Mika Okamoto, Gabriele Sarti
General AI
A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute t…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 4.8
2026-09-04 · Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam
General AI
Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a reference image. Recent approaches perform well when the reference image contains a single object, effectively capturing a global style that encomp…
- Review
- pending
- Role
- unreviewed
- Read
- later