arxiv
Score 24.6
2026-07-22 · Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li
General AI
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and t…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 24.6
2026-07-22 · Anmol Kankariya, Sercan Ö. Arık
General AI
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce P…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 23.6
2026-07-22 · Alexis Fox, Junlin Wang, Paul Rosu, Bhuwan Dhingra
Research Track A · General AI
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent har…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 23.3
2026-07-22 · Nabila Tasnim, Haoran Liu, Qing Cao, Saugata Ghose
Research Track A · General AI
Several edge computing platforms, such as autonomous vehicles and smart sensing devices, need to adapt to dynamic environments in real time by learning from new data in the field. Continual learning has emerged as a promising solution for edge training, by incorporating techniques that successfully combine a highly sum…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 22.6
2026-07-22 · Chang Liu, Xinyu Li, Artur Dubrawski
General AI
Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply them to gradually solve problems more effectively. We study whether Large Language Models (LLMs) can similarly benefit from such experiential abstractions. From LLMs' solution traces on the MATH training set, a st…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 20.6
2026-07-22 · Adel ElZemity, Shujun Li, Budi Arief
General AI
Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While large language models (LLMs) demonstrate impressive capabilities for technical artifact interpretation, the opacity and escalating API costs of closed-weight frontier models motivate e…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 19.8
2026-07-22 · Kailin Jiang, Lei Liu, Jian Xi, Hui Xu, Junlin Liu, Baochen Fu, Shaoqing Ren, Bin Li, Vichwang, Yu Lu, Haibo Shi
General AI
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, …
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 19.4
2026-07-17 · Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
General AI
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise activ…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 18.6
2026-07-22 · Pratyush Kumar
General AI
Configuring a computational fluid dynamics (CFD) case in OpenFOAM requires assembling a multi-directory input deck of mutually consistent solver, discretisation and boundary-condition dictionaries -- a task that remains a substantial barrier to non-specialist use of open-source CFD software. Large language models (LLMs…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 17.8
2026-07-22 · Yubiao Ma, Han Yu, Kai Guo, Changtai Lv, Zhengquan Mao, Boyang Xing, Xuemei Ren, Dongdong Zheng
Research Track A
Humans can progressively acquire highly dynamic motor skills while preserving reliable everyday motor abilities. In contrast, existing humanoid controllers face a trade-off between generalist and specialist capabilities: generalist motion tracking policies struggle to reliably execute rare highly dynamic motions, where…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 15.6
2026-07-22 · Mingqiang Tang, Haokun Wen, Meng Liu, Yupeng Hu, Weili Guan, Xuemeng Song
General AI
Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches ofte…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 15.0
2026-07-21 · Sam O'Nuallain, Nithya Rajkumar, Ramya Narayanasamy, Hanna Jiang, Shreyas Chaudhari, Andrew Drozdov
General AI
We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the representations exposed to a retrieval system. Rather than tuning retrievers, rerankers, or a small set of preprocessing hyperparameters, AutoIndex searches over programs that slice, enrich…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 14.8
2026-07-22 · Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li
General AI
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate …
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.6
2026-07-22 · Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
General AI
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-langu…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.6
2026-07-22 · Pengcheng Wang, Zhiquan Wang, Jayoung Lee, Zhuoyan Xu, Ran Xu, Saurabh Bagchi, Yin Li, Somali Chaterji
General AI
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large number of input visual tokens and the heavy computation of the large language model (LLM), remains a key barrier to practical deployment. R…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 14.6
2026-07-22 · Niqi Lyu, Pengtao Shi, Wei Qiu, Jianlin Zhong, Sicong Xia, Jianyao Ma, Yicheng Ding
General AI
Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework for token-level SLM-LLM collaborative inference. During generation, the SLM deci…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 13.8
2026-07-22 · Md Tanvirul Alam
General AI
Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduce Trace, a taxonomy-guided environment f…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 13.0
2026-07-21 · Nischay Dhankhar, Dos Baha, Abulhair Saparov
General AI
Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their use in train-time knowledge injection, whe…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.8
2026-07-22 · Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, Le Sun
General AI
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hiera…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.8
2026-07-22 · Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhichen Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zeng, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
General AI
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.6
2026-07-22 · Nethmi Muthugala, Supryadi, Surangika Ranathunga, Nisansa de Silva, Ruijie Tao, Ovindu Gunatunga, Pengyun Zhu, Shaowei Zhang, Jingting Zheng, Deyi Xiong
General AI
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 12.6
2026-07-22 · Sriprabha Ramanarayanan, Rahul G. S., Mohammad Al Fahim, Keerthi Ram, Ramesh Venkatesan, Mohanasankar Sivaprakasam
General AI
Attention Mechanism (AM) selectively focuses on essential information for imaging tasks and captures relationships between distant pixel neighborhoods to compute feature representations. Accelerated MRI reconstruction benefits from AM, as the imaging process involves Fourier domain measurements that influence image rep…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 12.4
2026-07-11 · Zhicheng Cai, Xinyuan Guo, Hanlin Wu, Mingxuan Wang, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
General AI
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential cause of PPO-Clip's failure. This work reveals the fu…
- Review
- pending
- Role
- unreviewed
- Read
- now
arxiv
Score 11.6
2026-07-22 · Sina Amirrajab, Volker Vehof, Michael Bietenbeck, Nuriye Akyol, Redouane Bouras, Khuraman Isgandarova, Alexandru Zlibut, Philipp Stalling, Ali Yilmaz
General AI
Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure, function, and pathology, but requires substantial experience in interpretation of CMR images that could be supported by artificial intelligence (AI)-based models. However, use of AI models for enhanced CMR rea…
- Review
- pending
- Role
- unreviewed
- Read
- now
huggingface
Score 10.8
2026-07-22 · Markus J. Buehler
General AI
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 10.6
2026-07-22 · Jona Carmon, Clara Bersch, Charles Fernyhough, Russell T. Hurlburt, Simone Kühn
General AI
Subjective experience is central to psychological science, yet methods for studying it force a choice between depth and scale. Classical Experience sampling, as in ecological momentary assessments (EMA), captures experience as it occurs, but it confines participants to predetermined response formats that prescribe how …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 10.6
2026-07-22 · Alejandro Gonzalez-Garcia, Wei Wang, Wei Xiao, Wilm Decre, Jan Swevers, Carlo Ratti, Daniela Rus
General AI
Aquatic self-reconfigurable robots must assemble into desired shapes while ensuring safe interactions among multiple agents. This paper proposes a hybrid framework that combines distributed Model Predictive Control (MPC) with Control Barrier Functions (CBFs) for multi-agent shape formation and reconfiguration. Given a …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 10.6
2026-07-22 · Sungrae Park, Sanghoon Kim, Gyoungjin Gim, Jungho Cho, Hyunwoong Ko, Minbyul Jeong, Minjeong Kim, Keunwoo Choi, Chaehun Shin, Chanwoong Yoon, Dongjun Kim, Eunwon Kim, Gyungin Shin, Hyeonju Lee, Hyungkyu Kang, Inseo Song, Jisu Bae, Jiyoon Han, Jiyun Lee, Joonkee Kim, Junyeop Lee, Mikyoung Cha, Sangwon Yu, Sehwan Joo, Seokyoon Kang, Seonghoon Yang, Seung Shin, Seunghyun Lee, Seungseop Lim, Seungyoun Shin, Sukyung Lee, Taegyeong Eo, Taehwan Oh, Taewhoo Lee, Wonho Song, Wonjun Oh, Wonseok Hwang, Yunsu Kim, Yura Shim, Hwalsuk Lee, Sunghun Kim, Du-Seong Chang, Kyunghyun Cho, Seungju Han, Yejin Choi, Junsuk Choe, Hwaran Lee, Minjeong Ban, Yun Taewon, Hwanjun Song, Jae-Gil Lee, KyungTae Lim, Alice Oh
General AI
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer am…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 10.6
2026-07-22 · Mahdi Nazeri, Anne-Kathrin Schmuck, Sadegh Soudjani, Alessandro Abate
General AI
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 9.8
2026-07-22 · Junhao Zhuang, Shiyi Zhang, Yuxuan Bian, Yaowei Li, Yawen Luo, Yijun Liu, Weiyang Jin, Songchun Zhang, Xianglong He, Xuying Zhang, Haoran Li, Haoyang Huang, Zeyue Xue, Nan Duan
General AI
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout stat…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.6
2026-07-22 · Yihang Gao, Vincent Y. F. Tan
General AI
Low-rank adaptation (LoRA) has become a widely used parameter-efficient fine-tuning method for large language models. Since different modules and layers may contribute unequally to downstream adaptation, allocating rank resources under a fixed parameter budget is an important problem for balancing efficiency, expressiv…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.6
2026-07-22 · Nicolas Kosanovic, Jordan Dowdy, Jean Chagas Vaz
General AI
Full-sized humanoid robot capabilities have grown exponentially in recent years, aiming towards general-purpose deployment in human environments. A popular control method used by manufacturers utilizes Virtual Reality for upper-body teleoperation and Reinforcement Learning for lower-body balance and locomotion control.…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 9.0
2026-07-21 · Vahid Satarifard, Fabian Baumann, Geetanjali Minsky, Laura Sisson, Lou M. Haux, Christophe Laudamiel, Nicholas A. Christakis
Research Track A
Perfumes are cultural artifacts and works of sensory art, composed from a finite, recombinable palette of notes that together evoke a distinctive scent impression. Here, we assemble the largest perfume corpus compiled to date, spanning multiple independent databases from 1900 to 2024, and study its evolution through a …
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.6
2026-07-22 · Amirhossein Sadr, Nima Soltani, Vahideh Moghtadaiee, Aida Pakniyat, Dara Rahmati, Saeid Gorgin
General AI
Physics-informed learning of partial differential equations (PDEs) has been dominated by multilayer perceptrons (MLPs), whose spectral bias and dense parameterization limit both accuracy and interpretability. Kolmogorov Arnold Networks (KANs) mitigate these limitations because their learnable spline activations are str…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.6
2026-07-22 · Yifan Xu, Zihao Wang, Zhixiao Wang, Jiaming Zhang, Yichun Yang, Desen Meng, Yuanxing Zhang, Pengfei Wan, Limin Wang
General AI
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing MLLMs usually generate captions directly from video inputs without exposing the perceptual evidence behind descriptions. As a r…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.6
2026-07-22 · Daniel Corva
General AI
Recent work established that under active inference, linear-Gaussian state-space models lose their epistemic drive (any incentive to act so as to gain information) "under any circumstances". The epistemic term of the Expected Free Energy becomes constant: the agent flattens to a Kalman filter whose gain sequence is fix…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 8.6
2026-07-22 · Abigail Woodring, Adrian Chan, Rana Muhammad Shahroz Khan, Sukwon Yun, Chau-Wai Wong, Tianlong Chen
General AI
Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computational costs. PortLLM is a training-free and data-free scheme used to adapt LLMs after continual pretra…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 7.8
2026-07-22 · Yang Xu, Gurpreet Singh Mukker, Raymond Wang, Jasper Gerigk, Maria Attarian, Igor Gilitschenski
General AI
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with limited spatial aware…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.6
2026-07-22 · Xiaoliang Shi, Zichen Wang, Runze Ma, Zhongyue Zhang, Shuangjia Zheng
General AI
Antibodies are essential proteins that play a central role in immune recognition by binding specific antigen molecules. Although recent protein language models have enabled progress in single-chain protein modeling and generation, they often fall short in antigen-specific antibody design, where effective modeling requi…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.6
2026-07-22 · Wael AbdAlmageed
General AI
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or rule definitions. Classical neuro-symbolic pipelines have a discrete in…
- Review
- pending
- Role
- unreviewed
- Read
- soon
arxiv
Score 7.6
2026-07-22 · Andreas Happe, Jürgen Cito, Jasmin Wachter
General AI
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and operated by trained practitioners -- agentic security tools exhibit \textit{indeterminacy} along three independent dimensions. First, their actions are drawn from a non-de…
- Review
- pending
- Role
- unreviewed
- Read
- soon
huggingface
Score 6.8
2026-07-22 · Yechan Kim, JongHyun Park, Dongho Yoon, Namhoon Jung, Moongu Jeon
General AI
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framewo…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 6.6
2026-07-22 · Rafael Gomes, Sophie Klumper, Guido Schäfer, Jens Schlöter
General AI
We study the strategic facility location problem under the egalitarian objective, where a mechanism uses the reported locations of a set of agents in Euclidean space to select a facility location that minimizes the maximum distance to any agent. We restrict our attention to strategyproof mechanisms, ensuring that no ag…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 6.6
2026-07-22 · Hiskias Dingeto
General AI
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim is never penalized.…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 6.0
2026-07-20 · Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi, Federico Alvetreti, Giorgio Strano, Donato Crisostomi, Giorgos Nikolaou, Tommaso Mencattini, Andrea Santilli, Emanuele Rodolà, Simone Scardapane, Alessio Devoto
General AI
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used…
- Review
- pending
- Role
- unreviewed
- Read
- later
arxiv
Score 5.8
2026-07-20 · Xuming Chen, Deniz Najafi, Mehrdad Morsali, Chengwei Zhou, Zahra Ghanaatianjobzari, Mahdi Nikdast, Shaahin Angizi, Gourav Datta
General AI
Silicon-photonic (SiPh) accelerators have emerged as a promising platform for Vision Transformer (ViT) inference by performing matrix multiplications on microring-resonator (MRR) banks with high throughput and energy efficiency. Extending these platforms to support on-chip fine-tuning remains challenging because backpr…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-07-15 · Dwip Dalal, Shivansh Patel, Chahit Jain, Jeonghwan Kim, Utkarsh Mishra, Alex Baratian, Hyeonjeong Ha, Heng Ji, Svetlana Lazebnik, Unnat Jain
General AI
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support visual and semantic generalization. Co-training on w…
- Review
- pending
- Role
- unreviewed
- Read
- later
huggingface
Score 5.4
2026-07-17 · Hao Liu, Chenghuan Huang, Ye Huang, Zhiying Wen, Mohan Zhang, Chen Li, Ziyang Ma, Jing Lyu, Jiangsu Du
General AI
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload …
- Review
- pending
- Role
- unreviewed
- Read
- later