
Hugging Face
SmolVLM
Apache 2.0 compact vision-language model family, from 2 billion parameters down to a 256 million variant.
- Parameters
- 2 billion flagship; 256 million compact variant
- Licence
- Apache 2.0
- Modality
- Vision-language (image and text input)
- Deployment
- Edge board, mobile device or small edge server
On-device image question answering
Camera scene description
Embedded multimodal features