Research Paper Cockpit

Daily Digest - 2026-10-03

Papers first seen in this daily snapshot.

Daily Archives

Quick jump into generated daily digests.

Research Workflow

Latest digest: 2026-10-04.

Papers

11 visible entries

huggingface Score 16.0

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

2026-09-29 · Haitong Ma, Chenxiao Gao, Rushi Qiang, Bo Dai, Na Li

General AI

Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineeri…

Review
pending
Role
unreviewed
Read
now
huggingface Score 15.8

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

2026-10-01 · Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji

General AI

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilit…

Review
pending
Role
unreviewed
Read
now
huggingface Score 14.4

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

2026-09-26 · Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu, Yue Guo, Sheng Wang, Wei Qiu, Hanwen Xu

General AI

Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist role…

Review
pending
Role
unreviewed
Read
now
huggingface Score 12.0

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

2026-09-28 · Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar

General AI

On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. W…

Review
pending
Role
unreviewed
Read
now
huggingface Score 9.8

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

2026-10-01 · Juan S. Santillana

General AI

Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a …

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

2026-09-29 · Jingxuan Wu, Yuzhe Yang, Yiqiao Huang, Chengzhi Liu, Qingni Wang, Chengxuan Qian, Shutong Wu, Jiawei Zhang, Xin Eric Wang

General AI

An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the informa…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 8.0

Video Generation Models: A Survey of Post-Training and Alignment

2026-09-30 · Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli

General AI

Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherenc…

Review
pending
Role
unreviewed
Read
soon
huggingface Score 6.4

Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration

2026-09-26 · Delong Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu

General AI

Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service orchestration to reduce this overhead while retaining service completion. The in…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.0

OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport

2026-09-29 · Guillaume Besset, Erwann Carn, Timothée Carecchio, Valentin Tordjman-Levavasseur, Fabian Schramm, Yann de Mont-Marin, Justin Carpentier, Ajay Suresha Sathya

General AI

Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differenc…

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.0

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

2026-09-30 · Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen, Yufeng Weng, Keane Ong, Ruojin Cai, Hang Hua, Justin K. W. Yeoh, Mengyu Wang

General AI

Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed …

Review
pending
Role
unreviewed
Read
later
huggingface Score 6.0

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

2026-09-30 · Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss

General AI

Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly t…

Review
pending
Role
unreviewed
Read
later