Applications

On-Device AI

Inference inside phones, PCs, wearables and appliances

On-device AI moves assistants, transcription, translation, imaging and personalisation into the device itself, improving responsiveness and keeping personal data local.

Solutions in this category

Representative platforms with the details engineering and procurement teams shortlist on.

Thin laptop on a dark desk with an abstract neural graphic on screen
Intel

Intel Core Ultra processors (Series 2)

Client processor platform with an NPU rated up to 48 TOPS alongside an Intel Arc Xe2 GPU and hybrid CPU cores.

NPU
NPU 4.0, up to 48 TOPS
GPU
Built-in Intel Arc GPU, Xe2 architecture
CPU
Hybrid performance and low-power efficient cores
Deployment
Laptop, mini PC or industrial client device
Local assistants and transcription
Edge kiosks and clients
Industrial HMI
Silicon package on a dark reflective surface with cyan edge lighting
Qualcomm

Qualcomm Snapdragon X Elite

Arm-based 4 nm PC platform with twelve Oryon CPU cores, Adreno graphics and a Hexagon NPU for on-device AI.

CPU
12-core Qualcomm Oryon, with dual-core boost
GPU
Integrated Qualcomm Adreno GPU
NPU
Qualcomm Hexagon NPU (exact TOPS rating not published per part on the product brief)
Deployment
Laptop, thin client or embedded appliance
On-device assistants
Vision and audio processing
Always-connected clients
Mobile device running an on-device machine learning feature
Google

Google LiteRT

Apache 2.0 on-device runtime for machine learning and generative AI, the successor to TensorFlow Lite.

Type
On-device machine learning and generative AI runtime
Licence
Apache 2.0
History
Renamed from TensorFlow Lite in September 2024
Deployment
Phone, embedded device or single-board computer
Mobile and appliance inference
Microcontroller and SBC vision
Offline features
Small single-board computer with a camera module attached
Hugging Face

SmolVLM

Apache 2.0 compact vision-language model family, from 2 billion parameters down to a 256 million variant.

Parameters
2 billion flagship; 256 million compact variant
Licence
Apache 2.0
Modality
Vision-language (image and text input)
Deployment
Edge board, mobile device or small edge server
On-device image question answering
Camera scene description
Embedded multimodal features

Specifications are summarised from publicly published manufacturer material and are provided for orientation only. Always confirm current figures directly with the manufacturer before purchasing or designing in.

Companies working in this category

Intel

Chip manufacturer

Supplies Core Ultra client processors that combine CPU, Arc graphics and an integrated NPU, and maintains the open-source OpenVINO toolkit for optimising inference on Intel CPUs, GPUs and NPUs.

Santa Clara, California, United States

Qualcomm

Chip manufacturer

Designs the Snapdragon X platform for Windows PCs and edge clients, pairing Oryon CPU cores, Adreno graphics and the Hexagon NPU. Qualcomm also owns the Edge Impulse edge MLOps platform.

San Diego, California, United States

AMD

Chip manufacturer

Produces Ryzen AI processors that combine Zen 5 CPU cores, RDNA 3.5 graphics and the XDNA 2 NPU, with large unified memory configurations used for local generative AI workloads.

Santa Clara, California, United States

Google

Chip manufacturer

Publishes the Coral Edge TPU accelerators for low-power on-device inference and maintains LiteRT, the on-device runtime that succeeded TensorFlow Lite, together with the open-weight Gemma model family.

Mountain View, California, United States

Rockchip

Chip manufacturer

Fabless SoC designer whose RK35xx application processors combine Arm CPU clusters, Mali graphics and an integrated NPU, widely used in single-board computers, panel devices and edge video products.

Fuzhou, China

Edge Impulse

Software provider

Edge MLOps platform, now part of Qualcomm, for collecting data, training and deploying models to microcontrollers, NPUs, CPUs and GPUs across a broad partner hardware ecosystem.

San Jose, California, United States

ggml.ai (llama.cpp)

Software provider

Maintainers of the ggml tensor library and llama.cpp, an MIT-licensed C and C++ inference engine that runs open-weight language models on CPUs and on CUDA, Metal and Vulkan backends.

Sofia, Bulgaria

Mistral AI

Software provider

Publishes open-weight language and multimodal models, including Apache 2.0 licensed Mistral Small builds designed for latency-sensitive and self-hosted deployments.

Paris, France

Hugging Face

Software provider

Hosts open-weight models and publishes the SmolVLM family of compact vision-language models designed explicitly for on-device inference under the Apache 2.0 licence.

New York, United States

On-Device AI: common questions

Does on-device AI work offline?
Fully local models do. Many shipping products use a hybrid approach where simple requests stay on device and heavier requests escalate to a server, so behaviour offline depends on the product.