Research Paper Cockpit

Daily Digest - 2026-09-01

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-09-22.

Papers

60 visible entries

arxiv Score 24.3

Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security

2026-08-30 · Sanket Badhe, Deep Shah, Priyanka Tiwari, Nehal Kathrotia

Research Track A · General AI

Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agen…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.0

One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual Learning

2026-08-31 · Yunxiang Fu, Meng Lou, Yizhou Yu

Research Track A · General AI

Class-incremental learning (CIL) requires a model to incrementally learn tasks that contain new classes without accessing earlier training data while preserving the ability to recognize all seen classes. Recently, pretrained-model-based approaches have become prevalent by adapting a frozen backbone with additional ligh…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.0

Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training

2026-08-31 · Lukas Borggren, Jenny Kunz, Marco Kuhlmann

Research Track A · General AI

Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models through additional training on target-domain corpora. In this work, we investigate such continued pre-training for adapt…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.0

Fast Weight Attention for Continual Learning

2026-08-27 · Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao

Research Track A · General AI

Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-me…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.0

SIR: Self-improving Red-teaming for Compute Use Agents

2026-08-31 · Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho

Research Track B · General AI

Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate everyday digital tasks. Because they can be exposed to untrusted content while operating, they are vulnerable to indirect …

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.3

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

2026-08-31 · Can Zhang, Baofeng Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Ruirui Li

General AI

Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can cause failure before substantive reasoning begins. Prescribing a fine-grained solution procedure for…

Review
pending
Role
unreviewed
Read
now
arxiv Score 18.3

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

2026-08-31 · Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria

General AI

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We ad…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.5

SHAPE of Chain-of-Thought in Math Reasoning

2026-06-28 · Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo

General AI

Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce SHAPE, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics education:…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.5

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching

2026-08-27 · Rong Shan, Tianyi Xu, Congmin Zheng, Wenteng Chen, Jiachen Zhu, Junjie Wu, Teng Wang, Weiwen Liu, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin

General AI

Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent within personal photo collections, where users often seek compact visual stories bound by structural…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.5

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu, Xiaoyang Wang, Yueyu Wang, Qianli Ma, Fan Yang, Ran Mei, Jia Wei, Jiangpeng Hu, Xuhao Liu, Hongming Chen, Yuanbin Shao, Yiyang Lin, Ziliang Li, Liang Pan, Xinhang Liu, Yuntao Ma, Tingxiang Fan

General AI

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicite…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.3

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

2026-08-31 · Jiajun Shi, Siyuan Tao, Yuhao Wu, Zexuan Wang, Jingyuan Zhang, Jiaheng Liu, Xinping Lei, Xinrong Zhang, Siyuan Fang, Zhewen Tan, Tianle Cai, Junhao Fang, Jiameng Huang, Yueyang Wang, Jinkai Liu, Yuxuan Zhang, Jian Yang, Zhoujun Li, Shen Yan, Wenhao Huang, Ge Zhang

Research Track A · General AI

Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use tha…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.3

Learning Simple Test-Time Environments for LLM Web Agents

2026-08-29 · Junxuan Li, Zijun Liu, Ziyi Huang, Peng Li, Yuzhou Liu, Ming Yan, Yang Liu

Research Track B · General AI

Large language model (LLM) agents have demonstrated remarkable proficiency in manually constructed environments, yet their performance frequently collapses when transitioned to complex real-world settings. Existing research largely attribute this degradation to the compositional generalization gaps in LLMs on combinati…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.0

Normalized Low-Rank Adaptation

2026-08-31 · Jiale Kang, Ziyin Yue, Zheng Zhan, Yangyi Huang, Weiyang Liu

Research Track A · General AI

While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning

2026-08-28 · Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan, Sowmya Rasipuram, Shubhashis Sengupta

General AI

Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich textual narratives. While recent advancements have enabled models to reason across modalities and perfo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

2026-08-31 · Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo

General AI

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewa…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.3

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

2026-08-31 · Qiyao Yan, Chenpeng Wang, Liangming Pan

General AI

When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct …

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy

2026-08-30 · Zhirui Fang, Qingchi Yu, Ziyang Chen, Longfei Li, Haoran Ma, Keru Zhou, Xinrun Xu, Samith Va, Yuxuan Hu, Peixuan Song, Qiang Du, Bin Qian, Yongkang Deng, Xin Li, Yezhen Wang, Zhe Li, Hao Luo, Shuyan Li, Ziwei Wang, Weijian Deng, Xiu Li

Research Track A · General AI

A robot's effective ``mind'' need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocatio…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

InsightToast: Proactive Information Retrieval & Glanceable Visualization in the Side Channel of Data-Rich Meetings

2026-08-31 · Mohammad Abolnejadian, Matthew Brehmer

General AI

Missing institutional context during meetings can impede effective participation. Retrieving relevant information, often scattered across heterogeneous internal and external sources, requires costly task-switching that disrupts both individual focus and collective conversational flow, particularly detrimental during co…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

2026-08-31 · Hamed Babaei Giglou, Sören Auer, Peio Popov, Mahsa Sanaei, Jennifer D'Souza

General AI

Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural aligners to knowledge graph embedding (KGE) models and, more recently, Large Language Model (LLM)-based approaches. Although modern OA frameworks provide unified ecosystems for deploying these heterogeneous…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents

2026-08-31 · Ziyi Bai, Siqi Li, Tinglei Huang, Börje F. Karlsson

General AI

Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual observations into executable plans. However, building agents that can continually improve through interaction and rapidly adapt to their environments remains challenging. Su…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

2026-08-31 · Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang, Jintao Chen, Xuhong Zhang

General AI

Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged info…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.3

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

2026-08-31 · Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos

General AI

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this today, but at prohibitive cost. Each questi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.3

DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception

2026-08-31 · Manish Kansana, Mohammed Yusuf Mujawar, Sudip Mittal, Shahram Rahimi, Noorbakhsh Amiri Golilarz

General AI

Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Robotic Perception) https://doi.org/10.21227/rmv3-be47, a calibrated dual-arm RGB-D-IR dataset for object-centered robotic perception using two independently moving eye-in-h…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.3

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

2026-08-31 · Chengyuan Gao, Jiang Wu, Tao Lu, Jiayan Guo, Mingkun Xu, Tianyi Zang, Shangyang Li

General AI

Computational mental health screening using multimodal speech and text has shown great promise. However, existing models often assume all clinical speech protocols carry equivalent evidentiary validity. In reality, heterogeneous protocols, from free interviews to fixed reading tasks, support fundamentally different evi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.3

MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval

2026-08-31 · Seokwon Song, Sohyeon Kim, Gunhee Kim

General AI

Information retrieval (IR) increasingly targets open-ended queries that admit diverse perspectives. Existing IR benchmarks, however, focus primarily on closed-ended queries, while even open-ended benchmarks largely consist of queries whose supporting documents span a single subject domain and modality. We introduce Mul…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.0

Unsupervised Continual Learning with Growing Self-Organizing Maps and Synthetic Replay

2026-08-27 · Pujan Thapa, Alexander Ororbia, Travis Desell

Research Track A · General AI

This work presents a generative continual learning framework based on growing self-organizing maps (GSOMs) that are augmented with learned distributional statistics as well as encoder-decoder models for class-incremental learning. The proposed approach enables exemplar-free replay using distributional statistical memor…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.5

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

2026-08-29 · Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini

General AI

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful A…

Review
pending
Role
unreviewed
Read
now
huggingface Score 11.5

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

2026-08-31 · Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim

General AI

Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.3

CineForge: Self-Improving Agents for Long-Horizon Video Generation

2026-08-30 · Junxiang Liu, Lin Wang, Haiyu Shi, Hongxu Ma, Xiaoyu Yang, Chunjie Chen, Xiaoxiao Xu, Kaiqiao Zhan, Boao Wang, Shuizhou Shi, Tianyun Zhu, Jie Li, Jiangtong Li

General AI

Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt construction, rendering, and revision across interdependent scenes. Existing adaptive video systems primarily refine requests or reusable skills, leaving recurring production…

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.5

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

2026-08-29 · Yue Peng, Lanke Xia, Zihan Wang, Jiahao Ye, Ke Ning, Hongyi Wen

General AI

Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information pr…

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.5

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

2026-08-29 · Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai

General AI

Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multim…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.5

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase

2026-08-29 · Daegyu Sung, Yukyeong Lee, Geon Park, Yumin Choi, Sung Ju Hwang

Research Track A · General AI

Organizations often develop and maintain portfolios of related applications: independently deployable codebases that share substantial domain logic, interface patterns, or operational conventions. As LLM coding agents are increasingly used to generate and maintain such software, a naive application-by-application workf…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.3

A Model with No Head and Many Thoughts

2026-08-31 · Nikita Koriagin, Yaroslav Aksenov, George Bredis, Gleb Gerasimov, Nikita Balagansky, Daniil Gavrilov

General AI

Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projecto…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.3

Aspire: Can Models Self-Evolve from Vague Goals?

2026-08-31 · Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Yuxuan Zhang, Xinping Lei, Junting Zhou, Zexuan Wang, Yuchen Wu, Huan Zhou, Duo Wang, Yinzhu Piao, Yongchang Peng, Yunfeng Shi, Jin Chen, Zuo Wang, Jinkai Liu, Jiaheng Liu, Wenxuan Zhang, Shen Yan, Wenhao Huang, Ge Zhang

General AI

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically beg…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.3

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

2026-08-31 · Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu

General AI

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generato…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.3

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

2026-08-31 · Ahmed El Kady, Aravind Narayanan, Rehana Noorani, Yani Ioannou, Shaina Raza

General AI

Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BB…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.3

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

2026-08-31 · Hamed Babaei Giglou, Sören Auer, Jennifer D'Souza

General AI

The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoL…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.0

Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring

2026-08-28 · Sjoerd van Straten, Marwan Hassani

Research Track A

Predictive Process Monitoring (PPM) models are increasingly deployed in dynamic environments where concept drift causes the underlying process distribution to shift over time. While recent work has moved toward online continual learning, existing methods train compact, task-specific networks entirely from scratch, leav…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.0

All You Need Is Non-Commutative Words

2026-08-29 · Carla M. Quispe Flores, Stanley Salvatierra, Renan Cabrera

Research Track A · General AI

We represent lexical tokens as unitary matrices and encode each sentence as their ordered product. The noncommutativity of matrix product captures word order without positional encodings (PEs). The same algebra yields several capabilities, including antisymmetric self-attention with no query, key, or value projections,…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.3

CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration

2026-08-31 · Yue Jiet Chong, Yimin Wang, Zhen Wu, Zixuan Wang, Wei Zhang, Xuanyao Fong

General AI

Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high utilization, memory efficiency, and scalable performance on compute-in-memory (CIM) accelerators. This paper presents CHIPSMORE, a multi-mode …

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.3

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation

2026-08-31 · Riya Ahuja, Tim Kacprowski, Roya Shiasi Sardoabi

General AI

BioMedRAG introduced retrieval-augmented generation with a learned chunk scorer for biomedical information extraction. However, it relies on fixed-size chunking which can fragment semantic evidence. We propose a configurable semantic chunking framework that addresses this limitation by combining entity-preserving windo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.3

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

2026-08-31 · Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu, Hsin-Ling Hsu, Donghua Zhang, Chenwei Wu, Jun-En Ding, Tongze Zhang, Shihao Yang, Pengfei Hu, Fang-Ming Hung, Feng Liu

General AI

Large language models (LLMs) offer promising clinical decision support but remain vulnerable to hallucinated facts, unsupported recommendations, and citation errors. We present DIASENTINEL, a fully on-premise multi-agent system for one-year type 2 diabetes mellitus (T2DM) risk screening and guideline-grounded report ge…

Review
pending
Role
unreviewed
Read
now
arxiv Score 9.0

Sense Once, Serve Many: Common-Trace Factorized Constrained PPO for Online Sensing-Session Consolidation in Multi-Tenant ISAC Networks

2026-08-29 · Dang-Dung Vu

Research Track A

Integrated sensing and communication (ISAC) networks can serve compatible requests through shared sensing sessions, but consolidation couples admission, reuse, profile selection, sensing service-level agreements (SLAs), communication quality of service (QoS), and future commitments. We formulate this problem as a const…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.5

Dynamic Important Example Mining for Reinforcement Finetuning

2026-08-29 · Haoru Tan, Sitong Wu, Yanfeng Chen, Shizhen Zhao, Yang-Tian Sun, Tianjia Liu, Chirui Chang, Shaofeng Zhang, Samm Sun, Xiuzhe Wu, Ruobing Xie, Xiaojuan Qi

General AI

Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data-centric RFT methods rely on static or heuristic sample selection, implicitly assuming a sample's value is fixed over training. …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.5

Verification-Aware Training for Speculative Decoding

2026-08-31 · Geonmo Gu, Byeongho Heo, HeeJae Jun, Yoohoon Kang, Sangmin Lee, Sangdoo Yun, Dongyoon Han

General AI

Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, which are verified by the target model in a single forward pass. Verification proceeds sequentially and discards every position from the first rejection onward, yet existing draft training relies on toke…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.3

Agentic research is oxymoronic

2026-08-31 · Natalie B. Hogg

General AI

The use of agentic large language models obviates human interpretation of scientific results, and will lead to substantial distrust in the literature.

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.3

Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression

2026-08-31 · Tianyi Zhao, Yinhan He, Wendy Zheng, Chen Chen

General AI

Chain-of-thought (CoT) reasoning improves multi-step problem solving, but long reasoning traces inflate inference cost. Token-level CoT compression reduces this cost by pruning full reasoning chains into shorter traces for model adaptation, making token selection the central challenge. Existing methods often rely on ex…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 7.5

Evaluating the Hidden Costs of Personalization in Large Language Models

2026-08-28 · Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang

General AI

While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on personal context such as conversation history, inferred preferences, and us…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.3

Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery

2026-08-31 · Onur Izmitlioglu, Shervin Dehghani, Tarek Ghannoum, Benedikt Schworm, Nassir Navab

General AI

Surgical phase recognition is key to context-aware computer-assisted feedback in vitreoretinal procedures, yet the scarcity of synchronized multimodal intraoperative data, particularly microscope views and intraoperative OCT, limits approaches that aim to replicate the multimodal integration surgeons perform naturally.…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.5

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

2026-08-25 · Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim, Donghoon Lee

General AI

Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with variational autoencoders (VAEs) to enhance computational efficiency without compromising visual quality. However, conventional VAEs are suboptimal for video data as they empl…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.3

ECGQuest: Benchmarking and Fine-Tuning Language Models for Electrocardiography

2026-08-31 · Mohammadsina Hassannia, Matthew A. Reyna, Reza Sameni

General AI

Electrocardiogram (ECG) interpretation requires knowledge of cardiology, electrophysiology, clinical diagnosis, ECG waveforms, signal acquisition, and instrumentation. Existing language-model benchmarks, however, primarily assess broad medical knowledge or interpretation of individual ECG signals and images rather than…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 5.5

Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models

2026-08-29 · Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang, Shuchang Zhou, Ming-Hsuan Yang

General AI

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still have certain shortcomings, hindering their further development as interactive design tools. Such schemes typically ad…

Review
pending
Role
unreviewed
Read
later
arxiv Score 5.3

Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

2026-08-31 · Yisen Xi

General AI

The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their users, identity determines data-handling terms, supply-chain risk, and capability expectations. No validated methodology exists for black-box identity verification of anon…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.3

Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations

2026-08-31 · William Solow, Paola Pesantez-Cabrera, Markus Keller, Lav Khot, Sandhya Saisubramanian, Alan Fern

General AI

Accurate daily predictions of cold hardiness in woody plants are critical in regions where freezing temperatures can damage dormant buds and reduce seasonal yield. Existing biophysical, hybrid, and deep learning models have shown high predictive accuracy when trained on local data but remain largely site-specific. The …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.3

HF-SID: High-Fidelity Semantic IDs for Generative Retrieval in Location-Based Services

2026-08-31 · Haowen Lin, Jing Li, Zhibin Hao, Fangye Wang, Lihui Su, Song Yang, Xiaojiang Zhou, Pengjie Wang

General AI

Generative retrieval has attracted increasing attention in Location-Based Services (LBS), where each Point-of-Interest (POI) is represented as a Semantic ID (SID). As the SID is the only channel through which POI information reaches the generative model, whatever it fails to preserve is irrecoverable at decoding time, …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 5.3

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies

2026-08-31 · Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong

General AI

Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task semantics, leaving rewards hand-crafted and behavior drifting from what c…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 4.3

Explainable Diabetic Retinopathy Classification Using Vision Foundation Models

2026-08-28 · Abhishek Verma, Anila Krishna, Abhishek Gajanan Bankar, Juan Miguel Lopez Alcaraz

General AI

Diabetic retinopathy (DR) is a major cause of preventable blindness, creating a need for accurate and trustworthy automated screening. This study investigates an explainable DR classification framework using vision foundation models and multiple transfer learning strategies. Three backbones, DINOv2, CLIP, and Vision Tr…

Review
pending
Role
unreviewed
Read
later
arxiv Score 4.3

Channel Gains to Captions: Task-Unified Multi-Level RF Sensing with Vision-Language Models

2026-08-31 · Tianyu Hu, Zhiren Gong, Haowei Cui, Shuai Wang, Samson Lasaulce, Lingxiang Li, Wassim Hamidouche, Zhi Chen, Merouane Debbah

General AI

This letter investigates a task-unified multi-level radio-frequency (RF) sensing framework driven by vision-language models (VLMs). Existing RF sensing methods rely on task-specific designs and provide only partial environmental information, limiting their ability to handle emerging 6G applications. To address this, we…

Review
pending
Role
unreviewed
Read
later
arxiv Score 4.3

Constrained Fair Allocations via Partition Matroid Reductions

2026-08-31 · Benjamin Cookson, Nisarg Shah

General AI

We study fair allocation of indivisible goods under additive valuations and matroid constraints. A challenging open question is whether a complete and feasible envy-free up to one good (EF1) allocation exists under every matroid that admits a complete and feasible allocation. The state-of-the-art result by Biswas and B…

Review
pending
Role
unreviewed
Read
later
arxiv Score 4.3

Context-Aware Interleaved Batching for WhisperX

2026-08-31 · Carlos Bain, Max Bain

General AI

While WhisperX accelerates speech transcription via intra-audio batching, it isolates audio segments, losing the historical context needed for coherent punctuation and terminology transcription. Conversely, standard Whisper retains context sequentially but suffers from slow inference and hallucination loops. To achieve…

Review
pending
Role
unreviewed
Read
later