Research Paper Cockpit

Daily Digest - 2026-07-22

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-07-25.

Papers

45 visible entries

huggingface Score 24.0

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

2026-07-20 · Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang

General AI

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggr…

Review
pending
Role
unreviewed
Read
now
arxiv Score 23.8

Agents in the Wild: Where Research Meets Deployment

2026-07-21 · Grace Hui Yang, Pranav N. Venkit, Hooman Sedghamiz, Enrico Santus, Victor Dibia, Ioana Baldini

General AI

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While aca…

Review
pending
Role
unreviewed
Read
now
arxiv Score 22.8

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

2026-07-21 · Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang

General AI

Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mode in this regime: \emph{repetitive copying}, where models extensively copy tex…

Review
pending
Role
unreviewed
Read
now
arxiv Score 20.8

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

2026-07-21 · Ziyi Wang, Yuhang Wu, Dongxu Piao, Xingyu Liu, Tianhui Zhou, Miao Liu

General AI

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal T…

Review
pending
Role
unreviewed
Read
now
arxiv Score 19.8

Continual Video-MLLM Adaptation over Evolving Domains

2026-07-21 · Rui Cheng, Meixing Shi, Yuxiang Cai, Jingcai Guo, Jianwei Yin, Zhi Chen

Research Track A · General AI

Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remains underexplored. In real-world deployments, video data often arrives continuously from heterogeneous domains, requiring the model to acquire new domain-specific knowled…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.8

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes

2026-07-21 · Daniel Pearson, Sidney Shapiro, Emiliano Sebastian Gonzalez Venegas, Sanad Al-Khatib, Aurora Pinzón Arzola

General AI

This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes. Rather than treating LangGraph, a low-level orchestration framework for stateful agents, as a model-quality benchmark target, we present three executable recipes -- SQL…

Review
pending
Role
unreviewed
Read
now
arxiv Score 17.8

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

2026-07-21 · Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou, Jalaj Bhandari, Kavosh Asadi, Daniel Jiang, Aditya Modi

General AI

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefi…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.8

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

2026-07-21 · Feinan Cheng, Dongliang Xu, Wenli Nong, Zhiheng Zhang, Ang Liu, Tianyu Wang, Yue Yao

General AI

Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments b…

Review
pending
Role
unreviewed
Read
now
arxiv Score 16.5

Code Division Modulation Layers Against Forgetting and Inference in Continual Gait Identification

2026-07-21 · Simone Milani

Research Track A

Continual learning (CL) has been recently employed in biometric identification systems thanks to its ability to integrate new knowledge within a pre-trained model and to the possibility of reducing the computational cost of training. Unfortunately, such approaches pose new challenges both in terms of final accuracy and…

Review
pending
Role
unreviewed
Read
now
huggingface Score 16.0

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

2026-07-21 · Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu

General AI

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fide…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

2026-07-21 · Yuan Wang, Yongchao Du, Mengting Chen, Jinsong Lan, Xuetao Feng, Xiaoyong Zhu

General AI

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-inten…

Review
pending
Role
unreviewed
Read
now
arxiv Score 15.8

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

2026-07-21 · Yu Chen, Caorui Li, Ziyu Xiong, Yidong Wang, Mingqi Gao, Shuman Liu, Biao Liu, Chunfeng Yang, Anxiang Zeng, Haibo Zhang, Chaofan Chen

General AI

Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity inputs. We introduce OmniReasoner, a tool-use post-training framework for Thinking with Long Audio-Video: omni-modal LLMs learn, via supervise…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.0

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

2026-07-21 · Kunlun Zhu, Xuyan Ye, Zhiguang Han, Yuchen Zhao, Bingxuan Li, Weijia Zhang, Muxin Tian, Xiangru Tang, Pan Lu, James Zou, Jiaxuan You, Heng Ji

General AI

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging …

Review
pending
Role
unreviewed
Read
now
arxiv Score 14.8

Selective State-Space Adaptation and Retrieval for Language Model Reasoning

2026-07-21 · Atahan Dokme, Larry Heck

General AI

Low-rank adaptation introduces a static learned update applied identically to every input. The update provides task-level adaptation but does not explicitly represent token-level or instance-level state variation. A family of adapters is proposed that introduces selective state-space recurrence at two complementary gra…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.0

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

2026-07-18 · Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang

General AI

Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect the NL2Pipeline gap. To bridge it, we introduce DataFlow-Harness, a platform t…

Review
pending
Role
unreviewed
Read
now
arxiv Score 13.8

Enhancing Relation Modeling with Social Attributes for Social Media Popularity Prediction

2026-07-21 · Bolun Zheng, Yuhao Luo, Wei Zhu, Ning Xu, An-An Liu, Lingyu Zhu, Canjin Wang

General AI

Recent studies highlight the critical role of retrieval-augmented mechanisms in social media popularity prediction (SMPP). Although such frameworks have improved SMPP performance by leveraging historical posts, existing methods still suffer from the low retrieval accuracy due to the oversight of relative relationships …

Review
pending
Role
unreviewed
Read
now
huggingface Score 13.0

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

2026-07-20 · Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee

General AI

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

2026-07-21 · Guy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara, Luis Belmar-Letelier

General AI

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world …

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models

2026-07-21 · Alessandro Scalese, Santhanakrishnan Narayanan, Constantinos Antoniou

General AI

The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. While Graph Neural Networks have emerged as fast, data-driven surrogates, their practical deployment is severely constrained by a spatial generalization gap. Standard models rely on transductive feature i…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.8

Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs

2026-07-21 · Alexander Manev

General AI

Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages. This leads to cross-lingual factual inconsistency, where they shift their empirical answer distributions based solely on the prompt…

Review
pending
Role
unreviewed
Read
now
arxiv Score 12.0

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents

2026-07-21 · Behzad Ousat, Nikita Turkmen, Lalchandra Rampersaud, Dillan Bailey, Amin Kharraz

Research Track B · General AI

LLM-based browser agents are rapidly changing the threat landscape for web security. Unlike traditional automation frameworks that execute predefined scripts, these agents can autonomously navigate websites, reason about page content, and interact with web interfaces using natural-language instructions. This evolution …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Staypoint Detection from Noisy Trajectory Data [Experiment Paper]

2026-07-21 · Lance Kennedy, Hossein Amiri, Yueyang Liu, Riyang Bao, Hanqi Chen, Mohammad Hashemi, Ruochen Kong, Xiaotong Liu, Joon-Seok Kim, Shengpu Tang, Liang Zhao, Andreas Züfle

General AI

Detecting staypoints from raw trajectory data is fundamental to numerous spatial computing applications. This process transforms raw numeric sequences of geolocations into semantically meaningful locations, such as homes, workplaces, or restaurants. Despite its importance for semantic trajectory analysis, staypoint det…

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Stochastic Multi-Objective Kinodynamic Planning Against Adversaries

2026-07-21 · Thomas Marshall Vielmetti, Daniel Cherenson, Dimitra Panagou

General AI

This paper addresses multi-objective kinodynamic planning in environments with stochastic hybrid adversaries that probabilistically transition to adversarial modes based on the ego state. The goal is to construct the Pareto-front of paths that trade off execution cost and the probability of safety constraint violation …

Review
pending
Role
unreviewed
Read
now
arxiv Score 11.8

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

2026-07-21 · Xilun Chen, Zhaleh Feizollahi, Ross Goodwin, Seungwhan Moon, Scott Yih, Pinar Donmez, Babak Damavandi, Luna Dong

General AI

Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factua…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

2026-07-21 · Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha

General AI

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemm…

Review
pending
Role
unreviewed
Read
now
arxiv Score 10.8

ISO: An RLVR-Native Optimization Stack

2026-07-21 · Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang

General AI

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through t…

Review
pending
Role
unreviewed
Read
now
huggingface Score 10.0

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

2026-07-21 · Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo

General AI

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors

2026-07-21 · Philip John, Eloghosa Ikponmwoba, Pinaki Pal, Opeoluwa Owoyele

General AI

This study introduces a reinforcement learning (RL) framework for generating optimal liquid-fueled reactors to improve lean blowout (LBO) predictions in gas turbine combustors. Existing approaches for determining cluster boundaries rely on manual heuristics or distance-based metrics in the input space. In contrast, the…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Breaking the Homogeneity Assumption: Specialized Multi-Generator Adversarial Learning for Rare Failure Detection in Predictive Maintenance

2026-07-21 · Alexis Lazanas, Georgios Kampouropoulos

General AI

Supervised learning models in the predictive maintenance field are regularly trained on highly imbalanced industrial datasets: machine failures occur rarely but have a disproportionate effect on operations. In addition to the clear class disparity, failure data are typically non-homogeneous, with different failure mode…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

2026-07-21 · Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang, Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang

General AI

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive mod…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models

2026-07-21 · Netanel Eliav

General AI

Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (markdown, plain text, prose, or tabular), how many simultaneous instructions a system prompt can carry before compliance degrades, and how much context a model can hold before recall …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 9.8

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

2026-07-21 · Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko

General AI

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deploym…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 9.0

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

2026-07-20 · AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao

General AI

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from te…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

Masked Visual Actions for Unified World Modeling

2026-07-21 · Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang

General AI

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 8.8

Provable diffusion-based posterior sampling for linear inverse problems via DDIM

2026-07-21 · Yuchen Jiao, Na Li, Changxiao Cai, Yuxin Chen, Gen Li

General AI

Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers either lack rigorous theoretical guarantees or incur substantial computational overhead. We propose a simple and efficient algorithm, called \pddim, for solving linear inverse proble…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

Delineate Anything v2: A Global Foundation Model for Field Delineation

2026-07-21 · Mykola Lavreniuk, Nataliia Kussul, Andrii Shelestov, Yevhenii Salii, Volodymyr Kuzin, Charlotte Julia Li-Xing Wang, Zoltan Szantoi

General AI

Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounting. While vision foundation models like SAM show remarkable zero-shot capabilities, they frequently fail in geospatial domains due to topological complexity, cropland t…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

2026-07-21 · Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang

General AI

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

2026-07-21 · Maohua Li, Qirui Li, Yanke Zhou, Yiduo Li, Zhaosheng Chi, Chao Xu, Cuifeng Shen, Yixuan Xu, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Shao-Qun Zhang

General AI

Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans,…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

EmbeddedKittens: An Evaluation of Code Embeddings for Scratch

2026-07-21 · Benedikt Fein, Gordon Fraser

General AI

The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as …

Review
pending
Role
unreviewed
Read
soon
arxiv Score 7.8

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

2026-07-21 · Haolin He, Renhe Sun, Zheqi Dai, Xingjian Du, Chunyat Wu, Zining Liang, Zhengxi Liu, Jiahe Lei, Runbang Wang, Jiayi Zhou, Mingru Yang, Xiquan Li, Yun Chen, Xie Chen, Zhiyao Duan, Weiqiang Wang, Mark D. Plumbley, Jian Liu, Qiuqiang Kong

General AI

DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Eversion-based robots can enable safe access,steering and endoscopic imaging within the spinal subarachnoid space

2026-07-21 · Zicong Wu, Panagiotis Kalozoumis, S. M. Hadi Sadati, Aminul I. Ahmed, Jonathan Shapey, Christian Baker, Thomas Booth, Wenfeng Xia, Sebastien Ourselin, Panagiotis Vartholomeos, Christos Bergeles

General AI

Safe navigation within the spinal subarachnoid space is constrained by its narrow, compliant, and delicate anatomy. Conventional catheters and continuum robots rely on proximal pushing, generating friction and shear along the tissue device interface that limit distal controllability and increase the risk of neural inju…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

Fundamental limits of distributed multiclass classification from simple binary decisions

2026-07-21 · Ioannis Papageorgiou, Srinivas Nomula, Ayalvadi Ganesh, Sidharth Jaggi, Parimal Parag

General AI

We consider the problem of constructing a $K$-class classifier from the combination of $O(\log K)$ simple binary classifiers -- this is a natural paradigm to construct a sophisticated classifier in a distributed manner with each agent performing a relatively straightforward task. We study the fundamental performance li…

Review
pending
Role
unreviewed
Read
soon
arxiv Score 6.8

ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

2026-07-21 · Chirag Vashist, Ke Li

General AI

Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold. Due to the success of diffusion models and flow matching, o…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.0

HPD-Parsing: Hierarchical Parallel Document Parsing

2026-07-21 · Shu Wei, Jingjing Wu, Lingshu Zhang, Qunyi Xie, Hao Zou, Le Xiang, Xu Fan, Yangliu Xu, Manhui Lin, Xiaolong Ma, Cheng Cui, Tengyu Du, YY

General AI

Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autoregressive trajectory,…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.0

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

2026-07-21 · Nuemaan Malik

General AI

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewAdam, an optimizer built on the observation that the three parameter pop…

Review
pending
Role
unreviewed
Read
later