0

ggml-org/llama.cpp

View on GitHub

LLM inference in C/C++

126,67922,628C++ggmlUpdated 1d ago
README

llama.cpp

llama

LLM inference in C/C++

License: MIT Release Nightly Server Docker Winget

ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

| Backend | Target devices | | --- | --- | | BLAS | All | | BLIS | All | | CANN | Ascend NPU | | CUDA | Nvidia GPU | | HIP | AMD GPU | | Hexagon [In Progress] | Snapdragon | | IBM zDNN | IBM Z & LinuxONE | | MUSA | Moore Threads GPU | | Metal | Apple Silicon | | OpenCL | Adreno GPU | | OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs | | RPC | All | | SYCL | Intel GPU | | VirtGPU | VirtGPU APIR | | Vulkan | GPU | | WebGPU | All | | ZenDNN | AMD CPU |

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

Comments0

No comments yet. Set the tone — say what you would want to know.