Apple's LensVLM-9B is a Vision-Language Model framework that maintains text recognition accuracy in compressed images by selectively expanding relevant regions using learned tools, achieving 4.3x compression while matching full-text performance on text QA benchmarks.