Liquid AI released LFM2.5-VL-3B, a 3-billion-parameter vision-language model optimized for on-device and edge deployment. The model pairs a SigLIP2 400M NaFlex vision encoder with Liquid AI's LFM2.5-2.6B text backbone, trained on ~34 trillion tokens with 4× more vision data than previous releases. Key features include strong screen/UI understanding across devices, improved object grounding with natural language queries, multi-image reasoning, and function calling for text-only and vision-text tasks.
LFM2.5-VL-3B is designed for real-time, low-latency inference on consumer hardware. It handles native resolution images up to 512×512 pixels without upscaling, uses intelligent patch-based processing for larger images, and delivers sub-second latency for single-turn queries. Pretraining incorporated curated and synthetic image-caption, OCR, grounding, and instruction-following data. The model supports multilingual vision understanding (Arabic, Chinese, French, German, Japanese, Korean, Spanish and more) and is available open-source on Hugging Face.
For production teams building on-device AI, LFM2.5-VL-3B signals the pace of competition in small-form-factor VLMs. At 3B parameters, it targets document processing, real-time object detection in automotive, OCR-heavy workflows, and edge inference at scale. The 4× vision data increase and strong tool-use benchmarks indicate frontier-grade architectures are now shipping at edge scale. Architects evaluating VLM inference should benchmark against LFM2.5 variants to understand whether edge latency gains justify the accuracy tradeoff versus larger cloud-based models.