Laya-Vision is an open-source multimodal model that makes calibrated decisions about images with optional text, answering choice, score, and yes/no questions in a single forward pass without text generation. It replaces Laya's encoder with SmolVLM-256M-Instruct and maintains the original API and training methodology. The experimental model, trained on VQAv2, A-OKVQA, and ScienceQA datasets, achieves calibrated outputs with ~71ms latency on NVIDIA L4 hardware.