arxiv
Score 29.4
2026-07-13 · Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu
Research Track A · General AI
Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired k…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 27.4
2026-07-13 · Mai A. Shaaban, Tausifa Jan Saleem, Alaa Mohamed, Dilnaz Utemissova, Ufaq Khan, Mohammad Yaqub
Research Track A · General AI
Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tasks without forgetting previously acquired knowledge. Continual learning (CL) provides a practical framework for this setting. Despite rapid progress in medical vision-language model…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 27.4
2026-07-16 · Yao He, Gan Sun, Wenqi Liang, Fazeng Li, Yang Cong
Research Track A
Similar to the natural capabilities of humans to sequentially learn new tasks, robots with Vision-Language-Action (VLA) models should possess lifelong learning ability to learn a new task when deployed in open-world environments. However, most recently proposed lifelong learning models aim to effectively learn the curr…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 24.4
2026-07-16 · Dante Lok
Research Track A
We introduce \emph{gate-zero growth}, a function-preserving (FP) operator for continual learning that adds new residual blocks through a zero-initialised gate. Under a transversality condition, gate-zero growth induces \emph{rank separation} in the functional Jacobian: old directions are unchanged, new-weight direction…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 24.0
2026-07-03 · Zhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
General AI
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilit…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 22.2
2026-07-15 · Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu
General AI
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 21.4
2026-07-11 · Varun Gandhi, Jaewook Lee, Shantanu Todmal, Franck Dernoncourt, Ryan Rossi, Zichao Wang, Andrew Lan
General AI
Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use lexical matching or semantic similarity, and how to co…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 21.4
2026-07-13 · Said Elnaffar, Farzad Rashidi
Research Track B · General AI
Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready w…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 21.4
2026-07-14 · Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
General AI
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRP…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 21.2
2026-07-13 · Jinxiu Liu, Jianru Li, Tanqing Kuang, Xuanming Liu, Kangfu Mei, Yandong Wen, Weiyang Liu
Research Track A · General AI
Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn cumulatively and evolve autonomously, which is a limitation we term the "perpetual novice"…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 21.2
2026-07-16 · Haran Raajesh, Kulin Shah, Adam Klivans, Philipp Krähenbühl
General AI
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predic…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 19.4
2026-07-16 · Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang
General AI
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely u…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 19.2
2026-07-15 · Zihao Yu, Xiu Yuan, Chongjie Zhang
Research Track A · General AI
Long-horizon robot planning requires more than predicting what actions will do next; it also requires memory of the embodied experience that makes future goals interpretable. People do not plan from the present scene alone: they draw on remembered places, object-state changes, prior procedures, and regularities reveale…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 18.4
2026-07-14 · Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou, Pedro Vélez, Amit Raj, Aaron Germuth, Thabo Beeler, Dimitris Samaras, Di Qiu
General AI
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models (MDMs) are ideally suited to this task, yet existing samplers either decode text…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.2
2026-07-16 · Paul Kassianik, Blaine Nelson, Yaron Singer
General AI
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomplete: in operational security, every reasoning step, tool call, telemetry quer…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.2
2026-07-16 · Shaoxiong Zhan, Shi Hu, Boyu Feng, Hai Lin, Andrew Gong, Zhengda Zhou, Jiaying Zhou, Yunyun Hou, Hao Su, Hai-Tao Zheng
General AI
Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscu…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.2
2026-07-16 · Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, Dongyu Liu
Research Track B · General AI
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and n…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.9
2026-07-14 · Richmond Alake, Cesare Bernardis, Paul Cayet, Luca Engel, Damien Hilloulin, Sungpack Hong, Allen Hosler, Nickolas Kavantzas, Ingo Kossyk, Son Le, Rhicheek Patra, Kartik Talamadupula, Valentin Venzin
Research Track A · General AI
Agent memory is a systems problem for long-horizon agents. Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumulation of procedural knowledge from prior outcomes. These requirements extend beyond document retriev…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 16.4
2026-07-16 · Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao
General AI
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limited guidance on interm…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.2
2026-07-16 · Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
General AI
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standard…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 16.2
2026-07-16 · Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari, Curtis Langlotz
General AI
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the prese…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.9
2026-07-14 · Enrico Gottardis, Mattia Tamiazzo, Simone Milani
Research Track A
Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-07-16 · Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis
General AI
The reliability of deepfake detectors frequently degrades under black-box adversarial transfer, as these models often rely on fragile, architecture-dependent forensic cues. Existing transfer attacks often lack semantic awareness and struggle to maintain effectiveness under strict no-query constraints, particularly when…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-07-16 · Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu
General AI
Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint, which commits every retrieval direction …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-07-16 · Debayan Mukhopadhyay, Utshab Kumar Ghosh, Shubham Chatterjee
General AI
Retrieval systems are trained and evaluated on a static idea of usefulness: hand a document and a question to a reader model, see whether the answer improves, and score the document accordingly. The idea holds up when a document is read on its own. It breaks when a language model works as a search agent, issuing severa…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-07-16 · Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang
General AI
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradig…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-07-16 · Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma, Yao Yao, Weiran Qi, Chuyan Jin, Guiyu Ma, Xingzhong Xu, Kai Yang, Ji-Rong Wen, Zhicheng Dou
General AI
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search attempts fail to yield useful evidence, current single- and multi-agent systems can beco…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.2
2026-07-16 · Hoang-Loc Cao, Van Pham, Truong Thanh Hung Nguyen, Phuc Truong Loc Nguyen, Phuc Ho, Veronica Whitford, Hung Cao
General AI
Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health research. In depression-related datasets, labels are often assigned without structured evidence, symptom-level justification, or traceable alignment with the criteria of the Diagnostic a…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.9
2026-07-15 · Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi
Research Track A · General AI
Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new fai…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-07-14 · Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu, Chenrui Zhao, Tongshu Bian, Xiaoman Liang, Meishan Zhang, Baotian Hu, Min Zhang
General AI
OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from e…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-07-16 · Sushant Gautam, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Steven A. Hicks
General AI
Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design choices across nine documented systems for question answering and explanation quality. Parameter-efficient adaptation of…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 13.2
2026-07-16 · Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu
General AI
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded danger is the same safety problem as ordinary text-level content danger. Through hidden-st…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.2
2026-07-14 · Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang
Research Track B · General AI
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether i…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-07-15 · Yiheng Huang, Zhijia Zhao, Bihuan Chen, Susheng Wu, Zhuotong Zhou, Yiheng Cao, Kun Hu, Xin Hu, Xin Peng
General AI
Open source software is vulnerable to supply-chain attacks through transitive dependencies, especially malicious code injected into NPM packages. Existing detectors often inadequately model obfuscated behavior, overlook JavaScript's object-centric features, poorly coordinate static and dynamic analysis, and lose semant…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-07-16 · Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang
General AI
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT es…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.2
2026-07-16 · Sridhar Mahadevan
General AI
This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting. Machine learning problems are specified as sketches: graphs with commutat…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 10.2
2026-07-16 · Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu, Jürgen Schmidhuber
General AI
Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scient…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 9.4
2026-07-15 · Rui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong
General AI
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reason…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.2
2026-07-16 · Pengcheng Zhou, Xuanyu Liu, Yanchen Yin, Bobo Li, Shengqiong Wu, Mong-Li Lee, Wynne Hsu
General AI
Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-localization, yet existing models remain susceptible to landmark bias, causing them to overlook geographical cues or form spurious correlations, ultimately resulting in inaccurate localization. To systematically investigate this issu…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.2
2026-07-16 · Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo, Tianyu Pang
General AI
MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and task-specific objectives. In particular, Diffusion…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.9
2026-07-16 · Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner
Research Track A
In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In parti…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.4
2026-07-14 · Lorenzo Busoni, Guido Agapito, Marco Bonaglia, Alfio Puglisi, Marco Xompero, Matteo Aliverti, Francesca Annibali, Carmelo Arcidiacono, Natalia Auricchio, Nicolò Azzaroli, Andrea Balestra, Alessandro Ballone, Louis Barbier, Andrea Baruffolo, Federico Battaini, Maria Bergomi, Andrea Bianco, Michele Cantiello, Giulio Capasso, Giulia Carlà, Enrico Cascone, Ed Chapin, Manal Chebbo, Simonetta Chinellato, Vincenzo Cianniello, Paolo Ciliegi, Mirko Colapietro, Jean-Jacques Correia, Giuseppe Cosentino, Elia Costa, Matteo D'ambrogio, Vincenzo De Caprio, Giuseppe De Luca, Nicholas Devaney, Ivan Di Antonio, Amico Di Cianno, Simone Di Filippo, Benedetta Di Francesco, Ugo Di Giammatteo, Chiara Di Prospero, Gianluca Di Rico, Andrea Di Rocco, Daphne Diretto, Christian Eredia, Simone Esposito, Jacopo Farinato, Italo Foppiani, Takashi Funakawa, Fulvio Gianotti, Laurence Gluck, Davide Greggio, Sylvain Guieu, Marco Gullieuszik, Yuuichi Harikane, Masahiro Ikoma, Laurent Jocou, Dan Kerley, Mikio Kurita, Salvatore Lampitelli, Tommaso Lapucci, Fulvio Laudisio, Yves Magnard, Demetrio Magrin, Hossein Mahmoodzadeh, Dheeraj Malik, Luca Marafatto, Laurence Michaud, Christophe Michel, Satoshi Miyazaki, Kentaro Motohara, David Mouillet, Thibaut Moulin, Matteo Munari, Kentaro Nagamine, Sylvain Oberti, Fabrice Pancher, Giorgio Pariani, Sophie Penger, Amedeo Petrella, Laurent Pinard, Cédric Plantet, Elisa Portaluri, Kalyan Radhakrishnan, Roberto Ragazzoni, Edoardo Redaelli, Edgar Renault, Colin Richardson, Marco Riva, Sylvain Rochat, Gabriele Rodeghiero, Luca Rosignoli, Bernardo Salasnich, Benoit Sassolas, Salvatore Savarese, Marcello Scalera, Pietro Schipani, Danilo Selvestrel, Mahshid Shiri, Mina Sibalic, Malcolm Smith, Sebastian Soler, Rosanna Sordo, Alessandro Tacchini, Alessio Taranto, Ludovico Teodori, Gabriele Umbriaco, Yoshinori Uzawa, Angelo Valentini, Jean-Pierre Véran
Research Track A
The Multiconjugate adaptive Optics Relay For ELT Observations (MORFEO) is a first-generation adaptive optics module for the Extremely Large Telescope (ELT), designed to deliver a diffraction-limited, highly uniform 53x53 arcsec field of view to the MICADO near-infrared camera. As the project advances toward its Final D…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.4
2026-07-14 · Zhenwen Miao, Honglin Wang, Mingheng Mi
Research Track A · General AI
As LLM technology advances, the space of model families, compute hardware, quantization schemes, parallelization strategies, and specialized optimization kernels continues to expand, sharply increasing the code complexity and maintenance cost of general-purpose inference frameworks. Conventional software engineering us…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.2
2026-07-16 · Ziren Gong, Xiaohan Li, Fabio Tosi, Ninghui Xu, Stefano Mattoccia, Jianfei Cai, Matteo Poggi
General AI
This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feed-forward model from the 3R family to process RGB videos and regress local point maps, and on a merging model, MAGMA, that combines loc…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.2
2026-07-16 · Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman
General AI
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addr…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 7.4
2026-07-15 · Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang
General AI
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quali…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 7.4
2026-07-15 · Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
General AI
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV se…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.2
2026-07-14 · Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal
General AI
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 2…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-07-16 · Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano, Francesco Pierri, Stefan Feuerriegel
General AI
Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficult to scale. Here, we introduce AutoSynthesis, an end-to-end multi-agent system for automated meta-analysis. Given a res…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-07-16 · Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo
General AI
Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models man…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-07-16 · Yazhi Zhang, Fuqiang Niu, Bowen Zhang
General AI
Political discourse has increasingly moved to short-video platforms, yet computational analysis of such content remains constrained by the scarcity of datasets that jointly preserve audiovisual information and hierarchical conversations. Here we present TikStance, a multimodal and context-aware dataset comprising 161 v…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 6.2
2026-07-16 · Qiwei Li, Jorge Ortiz
General AI
Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observational and collected without interventions, which makes causal questions such as "How would rain change traffic density?" difficult to answer. We present teLLMe, a system for explora…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 5.4
2026-07-15 · Sietse Schelpe
General AI
We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored, by graft, into a fresh inference context. The restore is bit-exact: under a pinne…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-07-16 · Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
General AI
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-07-16 · Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou
General AI
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.2
2026-07-16 · Mishel Carelli, Bernd Finkbeiner
General AI
We introduce Disintegration Temporal Logic (DTL), a new probabilistic temporal logic that can express a wide range of probabilistic hyperproperties, including probabilistic non-interference and perfect indistinguishability. DTL is based on the notion of measure disintegration from probability theory, which allows for c…
- Review
- pending
- Role
- unreviewed
- Read
- later