Paper Detail

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Luka Ribar, Jeevan Bhoot, Douglas Orr

huggingface Score 4.0

Published 2026-08-21 · First seen 2026-08-24

General AI

Abstract

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
later
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{ribar2026llama,
  title = {Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs},
  author = {Luka Ribar and Jeevan Bhoot and Douglas Orr},
  year = {2026},
  abstract = {Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by comp},
  url = {https://huggingface.co/papers/2608.21134},
  keywords = {vision-language models, quantization, 2.7-bit-per-parameter format, Arm CPUs, Llama 3.2 11B Vision Instruct, visual question answering, huggingface daily},
  eprint = {2608.21134},
  archiveprefix = {arXiv},
}

Metadata

{}