Compilers, quantisation, profiling and deployment SDKs
Developer tools turn a trained model into something that runs efficiently on target silicon: graph compilers, quantisation and pruning tools, profilers and device SDKs.
Solutions in this category
Representative platforms with the details engineering and procurement teams shortlist on.
NVIDIA
NVIDIA JetPack SDK
Official Jetson software stack combining the Linux board support package with CUDA-accelerated AI libraries.
Type
Board support package plus AI software stack
Components
Jetson Linux (bootloader, kernel, Ubuntu, drivers, OTA) and the Jetson AI stack
Specifications are summarised from publicly published manufacturer material and are provided for orientation only. Always confirm current figures directly with the manufacturer before purchasing or designing in.
Companies working in this category
NV
NVIDIA
Chip manufacturer
Develops the Jetson family of edge system-on-modules, the DGX Spark desktop AI system and the JetPack, TensorRT and DeepStream software stacks used for edge inference, robotics and vision AI.
Supplies Core Ultra client processors that combine CPU, Arc graphics and an integrated NPU, and maintains the open-source OpenVINO toolkit for optimising inference on Intel CPUs, GPUs and NPUs.
Publishes the Coral Edge TPU accelerators for low-power on-device inference and maintains LiteRT, the on-device runtime that succeeded TensorFlow Lite, together with the open-weight Gemma model family.
Fabless company developing discrete edge AI processors, M.2 and PCIe accelerator modules and the Hailo AI Software Suite, spanning classic vision workloads and, with Hailo-10H, generative models.
Maintains the YOLO family of open-source vision models covering detection, segmentation, classification, pose estimation and oriented bounding boxes, dual-licensed under AGPL-3.0 and a commercial licence.
Provides computer vision tooling including the open-source Inference package and inference server, covering model loading, pre- and post-processing and workflow execution on edge devices.
Edge MLOps platform, now part of Qualcomm, for collecting data, training and deploying models to microcontrollers, NPUs, CPUs and GPUs across a broad partner hardware ecosystem.
Container-based platform for deploying and managing fleets of Linux edge devices, with over-the-air updates, an API and SDK, and support for a large catalogue of device types.
Maintains the MIT-licensed Ollama runtime for pulling and serving open-weight language and multimodal models locally on workstations, servers and capable edge systems.
Maintainers of the ggml tensor library and llama.cpp, an MIT-licensed C and C++ inference engine that runs open-weight language models on CPUs and on CUDA, Metal and Vulkan backends.
Originated ONNX Runtime, the MIT-licensed cross-platform inference accelerator, and publishes the Phi family of small open-weight models aimed at low-latency local inference.
Hosts open-weight models and publishes the SmolVLM family of compact vision-language models designed explicitly for on-device inference under the Apache 2.0 licence.
Lower-precision weights and activations reduce memory footprint and increase throughput, at the cost of some accuracy that must be measured against your own evaluation set.