Liquid AI's LFM2.5-VL-3B: On-Device Vision-Language Model
Summary
Liquid AI has released LFM2.5-VL-3B, a new vision-language model with 3.1 billion parameters designed for on-device deployment. This model can read digital screens across mobile, web, and desktop platforms. It grounds objects to coordinates, parses documents and charts, and calls tools from text or image input. The model achieves an average score of 69.4 across 28 vision benchmarks. This matches InternVL-3.5-4B and is just 0.7 points behind Qwen3.5-4B, both of which are larger models. LFM2.5-VL-3B is non-reasoning, meaning it provides direct answers and maintains low latency. It fits into approximately 3 GB of memory and decodes 228 tokens per second on an Apple M5 Max. The checkpoint ships in four formats: native, GGUF, ONNX, and MLX. Day-one runtimes include llama.cpp, MLX, vLLM, SGLang, and ONNX. This makes it highly deployable for various applications. It's especially useful for indie developers, startups, and small to medium businesses with annual revenues under $10 million, as they can use it commercially at no cost. This technology could impact industries like consumer electronics, automotive, healthcare, and e-commerce.
This is an AI-generated audio summary. Always check the original source for complete reporting.