arxiv
Score 25.0
2026-09-15 · Pisante Aida, Formentin Simone
Research Track A
Real-world intelligent systems increasingly operate under open-world conditions, where user intents are not fixed or exhaustively known a priori and may evolve as new interaction patterns emerge. This paper proposes a unified uncertainty-aware probabilistic framework for continual new intent discovery under an evolving…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 24.5
2026-09-15 · Yunxiang Fu, Meng Lou, Zicheng Liao, Yizhou Yu
Research Track A · General AI
Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a data stream without catastrophic forgetting. While leveraging pretrained models has significantly advanced continual learning, existing methods exhibit a scalability bottl…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.9
2026-09-11 · Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal
Research Track B · General AI
Web agents can utilize reusable tools to reduce the cost and latency of low-level browser interaction, but automatically discovered tool collections can be large, redundant, and poorly aligned with user demand. We present AutoTailor, a meta-agentic framework for constructing and maintaining a compact set of trajectory-…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.8
2026-09-16 · Rebecca Ansell, Autumn Toney-Wails
General AI
Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that require integrating evidence across multiple reasoning steps, maintaining consistency with prior inferences, and updating beliefs under new constraints …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.2
2026-09-14 · Junhao Qiu, Qinglong Hu, Xialiang Tong, Mingxuan Yuan, Liyong Lin, Qingfu Zhang
General AI
Large language models have advanced automated algorithm discovery by synthesizing executable code, but existing frameworks trap them in rigid search pipelines with pre-defined control flows. This limitation restricts adaptive reasoning, blocks cross-paradigm transfer, and discards valuable execution feedback. We propos…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.0
2026-09-16 · Leon Arthur Marx, Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh
Research Track A · General AI
Weakly supervised class-incremental semantic segmentation (WILSS) aims to train a segmentation model over multiple steps, each introducing new concepts to be learned with only image-level supervision. We introduce DR.WILSS, an innovative approach to address catastrophic forgetting in continual learning using diffusion-…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.8
2026-09-16 · Izumi Takahara, Kazunori Nishio, Akira Aiba, Shigeru Kobayashi, Takao Nakajima, Taro Hitosugi, Teruyasu Mizoguchi
General AI
Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of optimized samples, with the measurements reduced to predefined scalar objectives and the reasons behind success left unarticulated. Here we present SynA…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.8
2026-09-16 · Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang
General AI
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience b…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 18.0
2026-09-15 · Caiqi Zhang, Xiaochen Zhu, Chengzu Li, Yulong Chen, Dharshan Kumaran, Nigel Collier
General AI
Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however, share one design premise: they only read the cur…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.9
2026-09-11 · Yuli Qiu, Yutong Li, Wei Su, Zeming Liu, Wanxiang Che, Heyan Huang, Haifeng Wang, Yuang Guo
Research Track A · General AI
Large language model agents are expected to continuously adapt to new tasks and environments over their lifetime by reusing past experience. However, existing memory-based agents struggle to transfer reusable experience across environments and suffer from catastrophic forgetting as experience accumulated. To address th…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.8
2026-09-16 · Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
General AI
Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these p…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 16.0
2026-09-15 · Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang
General AI
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk u…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.8
2026-09-15 · ZhuoXin Liu, Zhiming Ma, Ying Zhang, Mengzheng Yang, Yifan Wang, Zhengqi Huang, Yanhan Zhou, Zekun Lin, Jun Zhang, Shun Zhang, Yue Chen, Qiao Zhao, Peng Chen
Research Track B · General AI
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages s…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.8
2026-09-15 · Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li, Zhaochen Yu, Jinchao Ding, Qiang Gao, Pengyu Zhan, Yuntong Zhang, Tian Cheng, Zhenfei Yin, Yingcheng Wu, Ling Yang
Research Track A · General AI
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.8
2026-09-16 · João Meneses dos Santos, Arlindo L. Oliveira
General AI
Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action proposer with a slower planner, using two modular cognitive extensions: an Adaptiv…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.5
2026-09-15 · Hojin Lee, Yunho Lee, Daniel A Duecker, Cheolhyeon Kwon
Research Track A
Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods of…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.8
2026-09-16 · Matteo Golinelli, Idilio Drago, Matteo Boffa, Francesco Bergadano, Bruno Crispo
General AI
AI agents for security inspect web pages, source code, logs, configuration files, and command outputs. These environments may contain deceptive artifacts that influence the agent's behavior. We call this adversarial task contamination. Whereas prompt injection relies on attacker-supplied instructions, task contaminatio…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.8
2026-09-16 · Afnan Mostafa, William Ratcliff, Simon J. L. Billinge, Niaz Abdolrahim
General AI
The scientific literature contains decades of experimental measurements that remain difficult to access as structured data for modern AI and data-driven research. Much of this information is distributed across figures, captions, text, and tables, requiring experimental data and their context to be identified, connected…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.8
2026-09-16 · Tomas Balyo, Lukas Chrpa, G. Michael Youngblood
General AI
Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts accessibility to non-experts. This paper empirically evaluates whether current Large Language Models (LLMs) can reliably translate natural language testin…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.8
2026-09-15 · Xiaoyang Liu
General AI
Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, ma…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.8
2026-09-15 · Cai Ke, Xin Liu, Han Zhang, Jiangyue Yan, Zike Yuan, Ling Deng, Yue Yu, Hui Wang, Ruifeng Xu
General AI
Lifelong conversational agents rely on memory systems to maintain deep, context-aware interactions with users. However, existing explicit textual memory pipelines suffer from a severe information bottleneck, often losing subtle behavioral patterns and emotional shifts. Furthermore, being typically static post-deploymen…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.8
2026-09-16 · Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
General AI
Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason abou…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.5
2026-09-16 · Ruchi Bhatt, Pratibha Kumari, Shreya Bansal, Vedant Agnihotri, Dwarikanath Mahapatra, Mukesh Saini
Research Track A · General AI
Leaf counting is an important task in plant phenotyping for monitoring plant growth and estimating crop yield. Most existing methods rely on RGB images, but their performance is often affected by occlusion, lighting variations, and other real-world challenges. Additional modalities, such as depth and thermal images, ca…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.8
2026-09-15 · Zeyuan Huang, Gang Chen, Zixuan Hao, Guoqin Tang, Junyi Zong, Guoyou Ban, Jiale Wang, Haoyang Lv, Chaoqian Ren, Sitong Liu
Research Track A · General AI
Space robots are increasingly expected to perform long-duration, contact-rich, and multi-stage operations with limited human intervention. Recent advances in artificial intelligence (AI), robot learning, and embodied foundation models provide new opportunities to improve the autonomy and adaptability of such systems, b…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.8
2026-09-16 · Jiaxuan Jiang, Liyuan He, Zhixuan Fang
General AI
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during post-training and prevents agents from ac…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.8
2026-09-16 · Peter Potash
General AI
We evaluate six frontier language models on the two-agent $\log(N)$-Questions game. A questioner sees $N$ Wikipedia lead paragraphs and must identify a secretly chosen target using exactly $\log_2 N$ yes/no questions. An answerer sees only the target and the question, and replies with one word. Both roles run on the sa…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.4
2026-09-14 · Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang
General AI
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.0
2026-09-16 · Yifan Zhang, Yunheng Zou, Shaokun Zhang, Jian Hu, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong
General AI
Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared memory for such agents: research is recorded as an…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.0
2026-09-16 · Yizhuo Li, Jianhao Yan, Yun Luo, Zhi Wang, Futing Wang, Rong-Xi Tan, Kanghui Tian, Ganqu Cui, Ning Ding, Peilin Zhao, Yafu Li, Yu Cheng
General AI
In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Car…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.8
2026-09-16 · Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
General AI
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.8
2026-09-16 · Dunyao Xue, Chengshuo Du, Zhengbo Wang, Wenlin Dai, Cheng Meng
General AI
We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. Existing selection strategies rely predominantly on scalar probabilities, ignoring geometric semantic relationships and causing candidate redundancy.…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.8
2026-09-16 · Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang
General AI
Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing methods face a fundamental trade-off between accuracy and efficiency: direct full-image inference often fails to captur…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.8
2026-09-16 · Muhammad Abdullah Sohail
General AI
The Model Context Protocol (MCP) standardizes communication between autonomous Artificial Intelligence (AI) agents and remote tools over Streamable HTTP. This shift introduces a class of machine-generated, authenticated, and high-frequency JSON-RPC traffic directly into enterprise networks. Enterprise network defenders…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.4
2026-09-11 · Alexandru Ianta, Eleni Stroulia
Research Track B · General AI
The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-applicatio…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 11.0
2026-09-16 · Yu Lin, Yiming Wang, Runyuan Cai, Hanze Liu, Xiaodong Zeng
General AI
Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exi…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.8
2026-09-15 · Kiarash Tajbakhsh, Abdelrahman Faqieh, Michael Jopiti, Javier Garcia-Baroja, Philipp Zens, Branislav Zagrapan, Yuri Tolkach, Martin D. Berger, Aurel Perren, Bastian Dislich, Inti Zlobec, Amjad Khan
General AI
Pathology vision-language models are commonly built by pretraining or fine-tuning large encoders on paired image-caption data. We asked whether a pathology vision-language model can instead be assembled by parameter-efficient alignment of frozen unimodal foundation models, leaving their pretrained representations untou…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 10.0
2026-09-15 · Vivek Kalyanarangan
General AI
When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan in which each query decides how many bits of each k…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.8
2026-09-16 · Toqeer Ehsan, Thamar Solorio
General AI
Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper, we present a multi-step framework to enhance the annotation quality of NER datasets by employing automated techniques. We propose a frequency-based i…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.8
2026-09-16 · Michael M. Craig, Riley J. Hickman, Yingshan Ma, Rémi Piché-Taillefer, Christine Allen, Pauric Bannigan
General AI
Self-emulsifying drug delivery systems (SEDDS) can improve the oral bioavailability of poorly soluble drugs, but identifying high-performing formulations remains experimentally intensive. We present Andromeda 2, an agentic system that reasons over structured in-house experimental evidence and invokes computational and …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.8
2026-09-16 · Johannes Kaiser, Florian Braunmiller, Daniel Rückert, Georgios Kaissis
General AI
AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) architectures currently fail to satisfy. Expert-based routing improves in-domain learning but may sacrifice cross-modal signals, which appear par…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.8
2026-09-16 · Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, Tong Wu
Research Track A · General AI
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.2
2026-09-14 · Md Khalid Syfullah, Alvi Ataur Khalil
General AI
Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generaliza…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.0
2026-09-16 · Muhammad Abdullah Sohail
Research Track A · General AI
The Model Context Protocol (MCP) has emerged as the dominant interface for connecting autonomous agents to external data sources and execution environments. The ecosystem's transition from local process execution to remote Streamable HTTP deployments introduces unmeasured architectural and security constraints at scale…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.8
2026-09-16 · Jin Gao
General AI
Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a design system that supports both readers while preserving visual freedom and familiar human workflows. Three controlled studies examine component imple…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.8
2026-09-16 · Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid
General AI
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense caption…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.8
2026-09-16 · Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski
General AI
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, w…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 8.5
2026-09-15 · Xingxuan Zhang, Gang Ren, Hao Yuan, Hao Zou, Hongze Tan, Hui Wang, Jianhao Song, Jiansheng Li, Jiayao Zhang, Jinghan Zhang, Kaifang Li, Lang Mo, Li Mao, Mingchao Hao, Nuo Xu, Rui Ding, Ruiji Zhang, Shuyang Li, Siyu Mei, Tianyang Zhang, Weiyang Mu, Yancheng Dong, Yongxian Wei, Yuan Xue, Yuanrui Wang, Yue He, Zijia Yang, Ziyun Li, Dongzhe Li, Fuqiang Wang, Jiandong Liu, Jiawei Chen, Jiaxin Du, Kaijie Cheng, Kehan Li, Lei Sun, Linjun Zhou, Ningbo Dai, Qi Wang, Renzhe Xu, Shaoxing Du, Shumeng Yang, Wang Lu, Wenjing Chu, Xiannan Huang, Xiaoyu Lin, Xing Ai, Xinyan Han, Xuanyue Li, Xuanyue Su, Xukun Zhang, Yan Lu, Yaxin Zhang, Yi Qin, Yifei Huang, Yihan Xu, Yongle Lv, Yuanyuan Jiang, Yushan Han, Peng Cui
Research Track A
We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of i…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.8
2026-09-15 · Adam Zachary Wasserman, David Beauchemin
General AI
We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation) p…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.8
2026-09-16 · Pranaya Jajoo
General AI
Can a logged dataset visit every hidden state frequently and still be exponentially uninformative about a target policy's value? We show that it can when the logger depends on history. For every horizon $H \ge 3$, we construct two POMDPs with at most two latent states per stage, three actions, and a common logger with …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.8
2026-09-16 · Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen
General AI
Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distingui…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 7.4
2026-09-13 · Zian Liu, Yiwen Hu, Zican Dong, Tian Xie, Wayne Xin Zhao, Yucheng Ding, Ran Tao, Bryan Dai
General AI
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling. However, existing context extension approaches typically apply continued pretraining directly without modifying these layers, overlooking the spectral properties of linear attention state dy…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.8
2026-09-16 · Mohan Shi, Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang, Eray Eren, Abeer Alwan
General AI
Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting them to domain-shifted speech, such as ch…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.8
2026-09-16 · Elizabeth Pavlova, Hidenori Tanaka
General AI
Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. To this end, we introduce the Flag Game, a toy model for st…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.8
2026-09-16 · Matteo Marchi, João Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada
General AI
Large Language Models (LLMs) are now routinely trained using synthetic data, since high-quality human data has been exhausted by the ever increasing needs of larger and larger models. However, recursive training on synthetic data frequently induces model collapse, a degenerative feedback loop where models progressively…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.8
2026-09-16 · Jiaming Zhang, Homanga Bharadhwaj
General AI
Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes in object state. In this work, we study articulated objects such as doors, drawers, cabinets, laptops, ovens, and hinged containers that are ubiquitous in daily life and…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.3
2026-09-16 · Md Habib Ullah
General AI
The rapid proliferation of grid-edge distributed energy resources has significantly increased the operational complexity of modern power systems. Consequently, conventional computational techniques face growing scalability and computational-efficiency challenges in addressing large-scale optimization and control, uncer…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-09-15 · Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
General AI
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches an…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-09-16 · Ashutossh Gupta, Vassilis Kekatos
General AI
As data center (DC) loads increasingly penetrate the power grid, there is an urgent need for grid operators to provide clear dynamic specifications to DC owners to ensure safe grid operation. To this end, we study two salient behaviors of large language model (LLM) training loads: abrupt ramps at job initiation and ter…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-09-16 · Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa
General AI
Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories that lack force information, causing failures in contact-rich tasks where appropriate contact forces are essential for success. In this work, we expl…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-09-16 · Jonas Hummel, Luisa Faust, Elias Müller, Eva Bertog, Valeria Zitz, Marius Johannes Prill, Luca L. Bennardo, Luisa Weber, Tobias Röddiger, Michael Beigl
General AI
We present EarStreAM, a closed-loop earable system for stress-adaptive meditation that integrates in-ear physiological sensing with personalized, real-time intervention. Leveraging OpenEarable 2.0's multimodal sensing, EarStreAM continuously monitors physiological signals and detects elevated stress from heart rate and…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-09-16 · Ahmetcan Yavuz, Clara Meister, Tiago Pimentel
General AI
Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these …
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-09-16 · Zheng Liu, Xin Gao, Jinchao Zhu, Gao Huang
General AI
Parameter-efficient fine-tuning (PEFT) has recently emerged as a pivotal research direction for adapting pre-trained point cloud transformers to diverse downstream tasks. Although existing methods achieve excellent fine-tuning performance with high parameter efficiency, they ignore inference efficiency. To tackle this …
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-09-13 · Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao, Shafiq Joty
General AI
Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in common use, and each grows differently: expert dis…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-09-14 · Xingyang Li, Dongyun Zou, Shining Zhang, Jiacheng Chen, Haocheng Xi, Lvmin Zhang, Jun-Yan Zhu, Song Han, Zhekai Zhang, Yujun Lin, Muyang Li
General AI
Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entries, leaving typical e…
- Review
- pending
- Role
- unreviewed
- Read
- later