Hugging Face

SmolVLM

Apache 2.0 compact vision-language model family, from 2 billion parameters down to a 256 million variant.

Small single-board computer with a camera module attached

Overview

SmolVLM is a family of compact vision-language models designed explicitly for on-device use. The flagship model is 2 billion parameters, and the SmolVLM-256M variant is described by Hugging Face as the smallest multimodal model released, aimed at very constrained hardware.

Typical use cases

  • On-device image question answering
  • Camera scene description
  • Embedded multimodal features

Deployment environment

Edge board, mobile device or small edge server

Key specifications

Parameters
2 billion flagship; 256 million compact variant
Licence
Apache 2.0
Modality
Vision-language (image and text input)
Design goal
On-device multimodal inference

Specifications taken from current manufacturer documentation and last checked on 2026-09-08.

Source: manufacturer documentation

Specifications are summarised from publicly published manufacturer material and are provided for orientation only. Always confirm current figures directly with the manufacturer before purchasing or designing in.

Thin laptop on a dark desk with an abstract neural graphic on screen
Intel

Intel Core Ultra processors (Series 2)

Client processor platform with an NPU rated up to 48 TOPS alongside an Intel Arc Xe2 GPU and hybrid CPU cores.

NPU
NPU 4.0, up to 48 TOPS
GPU
Built-in Intel Arc GPU, Xe2 architecture
CPU
Hybrid performance and low-power efficient cores
Deployment
Laptop, mini PC or industrial client device
Local assistants and transcription
Edge kiosks and clients
Industrial HMI
Silicon package on a dark reflective surface with cyan edge lighting
Qualcomm

Qualcomm Snapdragon X Elite

Arm-based 4 nm PC platform with twelve Oryon CPU cores, Adreno graphics and a Hexagon NPU for on-device AI.

CPU
12-core Qualcomm Oryon, with dual-core boost
GPU
Integrated Qualcomm Adreno GPU
NPU
Qualcomm Hexagon NPU (exact TOPS rating not published per part on the product brief)
Deployment
Laptop, thin client or embedded appliance
On-device assistants
Vision and audio processing
Always-connected clients
Mobile device running an on-device machine learning feature
Google

Google LiteRT

Apache 2.0 on-device runtime for machine learning and generative AI, the successor to TensorFlow Lite.

Type
On-device machine learning and generative AI runtime
Licence
Apache 2.0
History
Renamed from TensorFlow Lite in September 2024
Deployment
Phone, embedded device or single-board computer
Mobile and appliance inference
Microcontroller and SBC vision
Offline features