LensVLM is a 9B Vision Language Model that compresses long text documents into images, then selectively expands only relevant pages to answer queries. The model uses learned tools to decompress specific sections, supporting compression ratios up to 15x while maintaining question-answering capabilities.
Apple's LensVLM-9B is a Vision-Language Model framework that maintains text recognition accuracy in compressed images by selectively expanding relevant regions using learned tools, achieving 4.3x compression while matching full-text performance on text QA benchmarks.