
ggml.ai (llama.cpp)
llama.cpp
MIT-licensed C and C++ inference engine for open-weight language models, on CPU or GPU backends.
- Type
- LLM inference engine in C and C++
- Licence
- MIT
- Backends
- CPU plus CUDA, Metal and Vulkan
- Deployment
- Edge board, workstation or server, with or without a GPU
On-device assistants
Quantised model serving
Arm and Apple silicon inference