I trained a vision-language model that answers typed questions about an image. It uses also choice, score and noul, like Jev.I thought, why Jev only processes text? There should be the possibility to process an image too. I will experiment with some pictures and questions in the next days to see, how well it performs in real life.On my M1 Pro a request with six questionut 400 ms p95,

about 60 ms on a desktop GPU. The server encodes the image once and scores each

option as a short suffix against the KVLooking forward to answer your questions! :)