Inference inside phones, PCs, wearables and appliances
On-device AI moves assistants, transcription, translation, imaging and personalisation into the device itself, improving responsiveness and keeping personal data local.
Solutions in this category
Representative platforms with the details engineering and procurement teams shortlist on.
Intel
Intel Core Ultra processors (Series 2)
Client processor platform with an NPU rated up to 48 TOPS alongside an Intel Arc Xe2 GPU and hybrid CPU cores.
Specifications are summarised from publicly published manufacturer material and are provided for orientation only. Always confirm current figures directly with the manufacturer before purchasing or designing in.
Companies working in this category
IN
Intel
Chip manufacturer
Supplies Core Ultra client processors that combine CPU, Arc graphics and an integrated NPU, and maintains the open-source OpenVINO toolkit for optimising inference on Intel CPUs, GPUs and NPUs.
Designs the Snapdragon X platform for Windows PCs and edge clients, pairing Oryon CPU cores, Adreno graphics and the Hexagon NPU. Qualcomm also owns the Edge Impulse edge MLOps platform.
Produces Ryzen AI processors that combine Zen 5 CPU cores, RDNA 3.5 graphics and the XDNA 2 NPU, with large unified memory configurations used for local generative AI workloads.
Publishes the Coral Edge TPU accelerators for low-power on-device inference and maintains LiteRT, the on-device runtime that succeeded TensorFlow Lite, together with the open-weight Gemma model family.
Fabless SoC designer whose RK35xx application processors combine Arm CPU clusters, Mali graphics and an integrated NPU, widely used in single-board computers, panel devices and edge video products.
Edge MLOps platform, now part of Qualcomm, for collecting data, training and deploying models to microcontrollers, NPUs, CPUs and GPUs across a broad partner hardware ecosystem.
Maintainers of the ggml tensor library and llama.cpp, an MIT-licensed C and C++ inference engine that runs open-weight language models on CPUs and on CUDA, Metal and Vulkan backends.
Publishes open-weight language and multimodal models, including Apache 2.0 licensed Mistral Small builds designed for latency-sensitive and self-hosted deployments.
Hosts open-weight models and publishes the SmolVLM family of compact vision-language models designed explicitly for on-device inference under the Apache 2.0 licence.
Fully local models do. Many shipping products use a hybrid approach where simple requests stay on device and heavier requests escalate to a server, so behaviour offline depends on the product.