Paper Detail
Luka Ribar, Jeevan Bhoot, Douglas Orr
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.
No ranking explanation is available yet.
No tags.
@misc{ribar2026llama,
title = {Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs},
author = {Luka Ribar and Jeevan Bhoot and Douglas Orr},
year = {2026},
abstract = {Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by comp},
url = {https://huggingface.co/papers/2608.21134},
keywords = {vision-language models, quantization, 2.7-bit-per-parameter format, Arm CPUs, Llama 3.2 11B Vision Instruct, visual question answering, huggingface daily},
eprint = {2608.21134},
archiveprefix = {arXiv},
}
{}