Apple’s LensVLM-9B Reads Compressed Pages, Then Zooms In

Apple has published the weights for LensVLM-9B, a vision-language model that reads long documents as heavily compressed images and zooms back in only on the pages it needs. The checkpoint went up on Hugging Face on September 21, 2026, alongside inference and evaluation code on GitHub. The paper behind it — a collaboration between Apple and Duke University — first appeared on arXiv in May. Its headline claim: on seven text question-answering benchmarks, LensVLM stays close to a model reading the full text at 4.3× effective compression, and beats every compression and retrieval baseline tested up to 10.1×.
Advanced
The Problem: Visual Compression Breaks Down
Rendering text as an image and feeding it to a VLM is an increasingly popular way to cut long-context cost — a single image can stand in for many more text tokens than it consumes as visual tokens. DeepSeek-OCR and Zhipu’s Glyph both build on this idea. The catch is resolution: push the compression far enough and, as the authors put it, characters “shrink below the vision encoder’s effective resolution,” becoming indistinguishable.
The paper quantifies how steep that drop is. Averaged across its seven QA benchmarks, the base Qwen3.5-9B model scores 72.4% when given the full text, but only 31.3% when the same text is rendered as images at 5× input compression. The compressed pages are still enough to tell roughly what is where; they are not enough to read the answer.
How LensVLM Works
LensVLM treats the compressed pages as a map rather than the territory. The document is deterministically rendered into images at one of three presets — roughly 5×, 10×, or 15× input compression, corresponding to 72, 48, or 24 visual tokens per image in Qwen3.5-9B’s encoder. The model skims those thumbnails, reasons about which pages likely hold the evidence, and calls an Expand tool (exposed as read_page in the released code) to pull back specific pages in uncompressed form. It can do this across up to six turns before answering; in practice it averages 1.3 calls.
What Expand returns depends on the input. For rendered text, it’s the original source text. For native documents such as PDFs, it can be OCR output or a high-resolution image crop — and here the paper finds the high-resolution zoom beats OCR, because it preserves layout cues OCR discards.
Training runs in three stages on top of Qwen3.5-9B-Base, with the vision encoder and projector frozen:
- Data construction — answer spans are tracked through pagination to label which pages hold the evidence, and only “hard” examples (ones the base model can’t answer from the thumbnails alone) are kept.
- Supervised fine-tuning on synthetic multi-turn tool-use trajectories written by Qwen3.5-397B, which was given the evidence and gold answer.
- Reinforcement learning with DAPO, using a reward of 0.7 × answer correctness plus 0.3 × correctness when the model also used the tool.
Results
At 5× input compression, LensVLM reaches 68.9% average accuracy against the 72.4% full-text bound — up from 31.3% for the base model reading compressed images, and 39.2% for the base model given the Expand tool without training. Once the tokens added by expansions are counted, that works out to 4.3× effective compression. Baselines include Glyph, LLMLingua-2 token pruning, and retrieval pipelines using BM25, BGE-M3, Jina-v4, Qwen3-Embedding, and ColPali.
Other findings worth noting:
- Memory: at 15× compression over 20 pages, LensVLM’s peak KV cache is 2,288 tokens versus 10,686 for full text — a 78.6% reduction, rising to 84.2% at 100 pages.
- Latency is the cost: a typical single-expansion query takes about 17 seconds versus 8 for single-turn full-text inference on a B200, roughly 2×.
- Robustness: training shrinks an 18-point accuracy spread across 20 font and layout configurations to under one point.
- Scale matters: 2B and 4B variants learn the tool-call syntax but pick the right page far less often (49.4% and 59.1% selection accuracy, versus 76.8% at 9B).
- Transfer: without code-specific training, the approach carries over to MMLongBench-Doc (50.5% at 5× with image zoom) and to code-understanding benchmarks, where it beats Glyph at every compression level.
- The tool alone helps: giving an untrained Claude Sonnet 4.6 the same
Expandtool improved its accuracy at every compression level.
What This Means
LensVLM reframes visual text compression as a retrieval problem the model solves for itself: a cheap, lossy overview plus the ability to ask for detail. That makes it a middle path between stuffing the whole document into context and relying on an external retriever — and on the paper’s benchmarks it beats both retrieval pipelines and pure compression at the same token budget. The trade is latency for memory, which suits long-document QA over large corpora better than interactive chat.
Two practical caveats. The weights ship under the Apple Machine Learning Research Model License, which restricts use to research, and the repository releases inference and evaluation code but not the training pipeline. The authors also stopped at 9B for compute reasons, so whether the gap to full text closes further at larger scale is still open.
Related Coverage
- Apple Open-Sources Its Foundation Models Framework, Adds Claude and Gemini — Apple’s on-device model layer from WWDC 2026
- Apple’s Simple Self-Distillation Boosts Code Generation by 30% — another Apple ML research release built on a Qwen base
- DeepSeek Open-Sources V4-Flash-Vision-Exp — DeepSeek’s latest multimodal checkpoint
- Qwen3.8-27B: Frontier Agentic Scores on a Single Consumer GPU — the Qwen family LensVLM builds on
This post was drafted with AI assistance and reviewed by RITS staff.




沪公网安备31011502017015号