
NVIDIA TensorRT
Inference compiler and runtime family that optimises models for NVIDIA GPUs, including Jetson.
- Type
- Inference compiler and runtime ecosystem
- Components
- TensorRT compiler, TensorRT-LLM, Model Optimizer, TensorRT for RTX, TensorRT Cloud
- Optimisations
- Graph optimisation, layer fusion, FP16 and INT8 calibration
- Deployment
- Jetson device, workstation or GPU server