Software provider

ggml.ai (llama.cpp)

Maintainers of the ggml tensor library and llama.cpp, an MIT-licensed C and C++ inference engine that runs open-weight language models on CPUs and on CUDA, Metal and Vulkan backends.

Sofia, BulgariaProfile reviewed editorially

Categories

Industries served

Solutions from ggml.ai (llama.cpp)

Terminal showing token generation from a local model
ggml.ai (llama.cpp)

llama.cpp

MIT-licensed C and C++ inference engine for open-weight language models, on CPU or GPU backends.

Type
LLM inference engine in C and C++
Licence
MIT
Backends
CPU plus CUDA, Metal and Vulkan
Deployment
Edge board, workstation or server, with or without a GPU
On-device assistants
Quantised model serving
Arm and Apple silicon inference

Company information is summarised from publicly available material. If you represent ggml.ai (llama.cpp) and would like to update or expand this profile, please get in touch.