Research Paper Cockpit

Daily Digest - 2026-07-24

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-07-25.

Papers

45 visible entries

arxiv Score 24.6

MIRROR: Learning from the Other View for Multi-Modal Reasoning

2026-07-23 · Wen Ye, Yuxiao Qu, Aviral Kumar, Xuezhe Ma

General AI

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a pro…

Review
pending
Role
unreviewed
Read
now
arxiv Score 24.6

OpenForgeRL: Train Harness-native Agents in Any Environment

2026-07-23 · Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao

Research Track B · General AI

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.6

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

2026-07-23 · Gaurav Dadhich

Research Track A · General AI

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token c…

Review
pending
Role
unreviewed
Read
now
huggingface Score 18.8

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

2026-07-23 · Zhongyuan Peng, Dan Huang, Chuyu Zhang, Caijun Xu, Changyi Xiao, Shibo Hong, David Lo, Lin Qiu, Xuezhi Cao, Jiyuan He, Yixin Cao

General AI

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirem…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.8

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

2026-07-22 · Paul Furgale, Severin Klingler, James Nolan, Matt Staats, Gaia Di Lorenzo, Elisa Martinez Abad, Christian Schüller, Razvan Dinu, Alessio Devoto, Pascal Berard, Gal Kaplun, Elad Sarafian, Riccardo Roveri, Leon Derczynski, Ricardo Silveira Cabral

General AI

Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions th…

Review
pending
Role
unreviewed
Read
now
huggingface Score 17.8

AREX: Towards a Recursively Self-Improving Agent for Deep Research

2026-07-23 · Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang

General AI

Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply searc…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.8

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

2026-07-23 · Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang

General AI

Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.8

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

2026-07-23 · Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang

General AI

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.6

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

2026-07-23 · Qian Wu, Xinrong Zhou, Zizhan Ma, Kai Chen, Zheyao Gao, Xun Lin, Hongqiu Wu, Longfei Gou, Yixiao Liu, Ann Sin Nga Lau, Qi Dou

General AI

Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or single-turn feedback, rather than organizing an entire clinical case into a decision-centered learning trajectory. We introduce \textit{MedGame}, a framework that tran…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.6

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

2026-07-23 · Yingchao Huang, Xin Wang, Yuhan Su, Shanshan Yao

General AI

Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving patient outcomes. Speech-based CI detection has emerged as a promising non-invasive approach, as speech signals encode both linguistic and acoustic markers associated wit…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.8

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

2026-07-22 · Hanjing Ye, Tianle Zeng, Jiazhao Zhang, Shaoan Wang, Zibo Zhang, Weisi Situ, Yuchen Zhou, Yonggen Ling, Hong Zhang

General AI

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstra…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.6

Adaptive Multi-Horizon Reinforcement Learning

2026-07-22 · Manoosh Samiei, Doina Precup, Paul Masset

Research Track A · General AI

Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents ex…

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.6

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

2026-07-23 · Hongxin Zhang, Chunru Lin, Junyan Li, Zhou Xian, Tsun-Hsuan Wang, Chuang Gan

General AI

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation model…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.6

Diffusion Language Model for Recommendation

2026-07-23 · Chengyi Liu, Yongqi Zhou, Junwei Pan, Zhixiang Feng, Chengguo Yin, Haijie Gu, Jie Jiang, Yinghao Liu, Yujuan Ding, Qing Li, Wenqi Fan

General AI

Large language model (LLM)-empowered recommender systems have emerged as a promising paradigm for generative recommendation, leveraging their strong semantic reasoning and generative capacity to model complex, diverse user preferences. However, most existing approaches rely on an autoregressive paradigm that is subopti…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.8

LLMs Get Lost in Evolving User Intent

2026-07-22 · Jihoon Tack, Philippe Laban, Jennifer Neville

General AI

As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.8

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

2026-07-23 · Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, Zhijian Shao, Yuchen Shi, Shuwen Zhang, Chaofan Qiu, Linjie Che, Xiaoxi Zhao, Feng Wu, Kai Zhang, Chaofan Zhu, Yubin Qi, Xiaoyun Liang, Peijie Dong, Yunhao Zhang, Yuanjie Zhu, Ling Jiang, Xianjun Zhang, Zhehang Chu, Anyuan Sang, Zhen Feng, Sen Nie, Shi Wu, Yuanzhen Xu, Xin Li, Ning Yang, Zhiqiang Dong, Hande Dong, Qiang Lin, Yi Liu, Yunsheng Wu, Ke Li, Xing Sun

General AI

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four wo…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.6

Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

2026-07-23 · Zikui Cai, Kaushal Janga, Tan Dat Dao, Seungjae Lee, Shivin Dass, Mingyo Seo, Kaiyu Yue, Mintong Kang, Nandhu Pillai, Monte Hoover, Aadi Palnitkar, Ruchit Rawal, Ruijie Zheng, Bo Li, Yuke Zhu, Roberto Martín-Martín, Tom Goldstein, Furong Huang

General AI

Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continuously and must accumulate, retain, and selectively reuse information acquired from prior interaction…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.6

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

2026-07-23 · Vasudha Bhatnagar, Purnima Bindal, Vikas Kumar, Raj Kumari Bahl

General AI

Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluent summaries. However, underlying stochasticity of the large language models raises concerns about the stability and trustworthiness of the LLM-generated summaries. This…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.6

Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

2026-07-23 · Natan Levy, Harel Berger

General AI

AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments. This democratization enables rapid local innovation, but it also creates a reliability gap: agents that appear to users as simple productivity artifacts may depend on …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.6

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

2026-07-23 · Baihui Wang, Bernard Koch

General AI

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the b…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.6

Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning

2026-07-23 · Rogerio Guimaraes, Pietro Perona

General AI

Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the initial noise seed, many approaches spend extra compute on seed search or resampling under…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.6

Visual Contrastive Self-Distillation

2026-07-23 · Yijun Liang, Yunjie Tian, Yijiang Li, Yuqi Jia, Furong Huang, Tianyi Zhou, Di Fu

General AI

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.6

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

2026-07-23 · Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin

General AI

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.6

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

2026-07-23 · Sicheng Mo, Yuheng Li, Ziyang Leng, Krishna Kumar Singh, Bolei Zhou

General AI

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to mai…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 10.1

Benchmarking Agents for Proving Theorems in Quantum Algorithms and Quantum Information

2026-07-23 · Lei Zhang, Yusheng Zhao, Yimeng Cao, Ranyiliu Chen, Mingrui Jing, Jizhe Lai, Ziao Tang, Jingu Xie, Hongshun Yao, Xuanqiang Zhao, Guocheng Zhen, Chengkai Zhu, Xin Wang

General AI

Formal verification is becoming increasingly practical for quantum computing, yet the ability of AI agents to construct machine-checkable proofs in this domain remains unmeasured. We introduce Lean-QuantumAlg-Bench and Lean-QIT-Bench, two Lean 4 benchmarks containing 36 and 40 theorem-completion tasks for quantum algor…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.8

Robostral Navigate

2026-07-22 · Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal

General AI

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduc…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.6

3D-Aware VLMs with Implicit and Explicit Geometries

2026-07-23 · Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao, Ran Xu, Shijian Lu, Gongjie Zhang

General AI

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equi…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.6

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

2026-07-23 · Rafsan Jany, Shadab Tanjeed Ahmad, Ahsan Bulbul, Tahsinul Islam, Md Azam Hossain, Abu Raihan Mostofa Kamal

General AI

Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.4

Predictive Divergence Masks for LLM RL

2026-07-12 · Xiangxin Zhou, Jiarui Yao, Penghui Qi, Bowen Ping, Jiaqi Tang, Haonan Wang, Tianyu Pang

General AI

Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates. The dominant PPO-style approach uses the sampled-token importance ratio for two criteria: a proximity criterion, which asks whether the policy has moved too far from the behavior policy, and a…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.3

Sample-Efficient Learning from Agent Experience

2026-07-23 · Chenhui Gou, Haoqin Tu, Yunhao Fang, Jianfei Cai, Hamid Rezatofighi

Research Track A · General AI

Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is re…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.8

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

2026-07-23 · Hyunmin Cho, Jaejun Yoo, Kyong Hwan Jin

General AI

We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectral support. We reali…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.6

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

2026-07-23 · Mengfei Zhao, Dihong Huang, Yikai Tang, Peihao Li, Mingxuan Yan, Ruiqi Zhuang, Yanjia Huang, Jie Wang, Hai Zhai, Tony Zhou, Rui Zhang, Zhexi Luo, Yuchen Huang, Jianfei Yang, Jiachen Li

General AI

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalab…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity

2026-07-23 · Hongnan Ma, Yiwei Shi, Mengyue Yang, Weiru Liu

General AI

Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficiency-oriented methods can assign high importance to spurious subsequences that support the prediction wit…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs

2026-07-23 · Kaiwen Zhang, Guanjun Liu

General AI

Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing interleavings. Large language models can synthesize executable Rust tests, but their outputs often violate API preconditions, remain shallow, or reduce concurrency to accidental sequential traces. Conve…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

Improved lower bounds for the Shannon capacity of odd cycles

2026-07-23 · Nathaniel Itty, Christopher D. Rosin, Chase Carstensen, Daniel Reichman

General AI

The Shannon capacity $Θ(G)$ of a graph $G$ quantifies the maximum rate at which information can be transmitted with zero error over a noisy channel. It is lower bounded by $α(G^d)^{1/d}$ for any $d$, where $α(G^d)$ is the independence number of the $d$-th strong power of $G$. We construct independent sets of size $1347…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

2026-07-23 · Linjun Li

General AI

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.6

The Boundaries of Automation: A Theory of Persistent Human Participation

2026-07-23 · Fares Fourati, Hinrich Schütze, Eyke Hüllermeier, Iryna Gurevych

General AI

The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Implicit in this pursuit is the assumption that humans remain in the loop only because current AI systems are not yet sufficiently capable. This paper challenges that assump…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.6

How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning

2026-07-23 · Kaizhen Tan, Heqing Du, Yang Feng

General AI

A LoRA adapter is a few megabytes that almost everyone treats as a skill rather than a record of the data behind it. We put that assumption on a scale. Extending compression-based memorization analysis to the frozen-base setting, we measure directly, in bits, how much a low-rank adapter writes into a model it never cha…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Out-of-Distribution Detection in Wireless Multimodal Foundation Models for 6G ISAC

2026-07-23 · Mohammad Farzanullah, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci

General AI

The integration of Foundation Models (FMs), such as the Wireless Multimodal Foundation Model (WMFM), into 6G networks provides a unified framework for Integrated Sensing and Communication (ISAC), leveraging generalized representations to simultaneously optimize data transmission and environmental perception. However, t…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

2026-07-23 · Yu Qi, Zhang Ye, Xinyi Xu, Yuxuan Lu, Amitoj Sandhu, Boce Hu, Haojie Huang, Jonathan Tremblay, Lawson L. S. Wong

General AI

Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual \textit{instruction factors}, \textit{e.g.…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Stokes-Informed Diffusion for Robust Linear Polarization Estimation

2026-07-23 · Yidong Luo, Chenggong Li, Yuchao Feng, Boxin Shi, Junchao Zhang, Xin Yuan

General AI

Polarization cues benefit applications such as material detection and de-reflection, yet acquiring them typically requires dedicated hardware. This motivates us to estimate the linear polarization from a single RGB image. However, the task is inherently ill-posed, with the Angle of Polarization (AoP) becoming particula…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Surprisal Theory is Tautological (without Rational Grounding)

2026-07-23 · Ryan Cotterell

General AI

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose s…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Three-Pronged Spectral Control for Federated Parameter Efficient Fine Tuning

2026-07-23 · Shiva Raj Pokhrel, Dipsan Bhattarai, Anwar Walid

General AI

Federated parameter-efficient fine-tuning (PEFT) enables communication-efficient adaptation of large pretrained models on decentralized edge data, but it remains fragile under non-IID client heterogeneity. In low-rank adaptation (LoRA), different clients may learn locally useful but spectrally misaligned update subspac…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Transparent by Design, Usable in Practice? A Formative Usability Study of a Conversational Product Advisor

2026-07-23 · Kevin Schott, Dagmar Kern, Daniel Hienert

General AI

Large language models can make conversational product advisors fluent but opaque. If they hide the logic behind a ranking and the evidence for a recommendation inside natural-language replies, they challenge users' ability to understand, trust, and steer the results. One response is to build transparency into the advis…

Review
pending
Role
unreviewed
Read
later
arxiv Score 6.6

Unified Video Dense Prediction from Disjoint Data

2026-07-23 · Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee

General AI

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large …

Review
pending
Role
unreviewed
Read
later