l
AI Tool Profile

llama.cpp

llama.cpp is a free, MIT-licensed C/C++ engine for running language models on your own hardware, including CPUs and many GPUs. It includes an OpenAI-compatible server and sits under several local AI apps.

Website
github.com
Pricing model
Open Source
Price start
Free

Description of llama.cpp

llama.cpp is an open-source project for inference of language and vision-language models, written in plain C/C++ with no external dependencies. It is built on the ggml library and hosted under the ggml-org organization on GitHub. The licence is MIT, and there is no company pricing page: the software is free to download, build and use.

The project supports a wide range of hardware. On Apple silicon it uses Metal, Accelerate and ARM NEON; on x86 it uses AVX, AVX2, AVX512 and AMX; and RISC-V is supported. For graphics cards it lists CUDA for NVIDIA, HIP for AMD, MUSA for Moore Threads, Vulkan and SYCL for Intel, along with other backends such as BLAS, CANN for Ascend NPUs, Hexagon, OpenCL and WebGPU. Quantization from 1.5-bit to 8-bit integer formats lets larger models fit in less memory.

For everyday use, the repository shows a command that starts llama-server, an OpenAI-compatible API server with a built-in web interface, and links a REST API reference. Other desktop apps and runners are built on this engine, which is why it often appears underneath tools that look simpler.

It is aimed at technical users. You work from the command line, choose and download model files yourself, and tune settings for your hardware. A buyer who wants a ready chat window with model search and one-click downloads will usually find a packaged app easier. In return, llama.cpp offers direct control, runs fully offline once a model is downloaded, and gives you full access to the source code.

Features of llama.cpp

  1. Runs language and vision models on CPUs and on NVIDIA, AMD, Intel, Apple and other accelerators through many backends.
  2. Quantization from 1.5-bit to 8-bit integer formats helps large models fit in limited memory.
  3. llama-server provides an OpenAI-compatible API and a built-in web UI for local use.

Frequently asked questions

Is llama.cpp free?

Yes. It is open source under the MIT licence and hosted by ggml-org on GitHub. There is no paid plan or account; your cost is the computer you run it on.

Who should use llama.cpp directly?

People comfortable with the command line who want control over model files and hardware settings. Others may prefer a packaged app that is built on the engine.

Does llama.cpp have an API?

Yes. The README shows a command that launches an OpenAI-compatible API server, llama-server, and links a REST API reference. It also includes a built-in web interface.

Website: https://github.com/ggml-org/llama.cpp

Company: ggml-org (open-source project)

Key facts

Free MIT-licensed C/C++ engine for running language models on your own CPU or GPU, with an OpenAI-compatible server. Aimed at technical users; most local chat apps build on it.

Pricing model Open source
Free tier Yes · Free MIT-licensed software; no account or paid plan · no credit card required
Starting price Not verified
Data residency On-premises — Inference runs on hardware you control; the README describes local inference and a server you start yourself.
Compliance (as stated by vendor) Not verified
Integrations Not verified
Public API Yes
Self-hostable Yes
Official site Not verified
Can a general AI assistant replace it? A hosted chatbot needs no setup. llama.cpp suits offline, private or custom deployments; a packaged app like Jan hides the command line.
Last verified 2026-10-10 by editor

Alternatives & Similar Tools

Overchat AI logo
Overchat AI Freemium

Overchat AI is an all-in-one AI platform that combines 50+ leading models with 100+ mini apps for chat, image and video generation, document work, writing, and study tasks.

Jan logo
Jan Open Source

Jan is a free, open-source ChatGPT alternative from Menlo Research that runs local or online models on Windows, macOS and Linux, with an OpenAI-compatible local API server.

LM Studio Bionic logo

LM Studio, now shown as LM Studio Bionic, runs local language models and an agent that works on documents and code. The free plan covers local models, offline voice transcription and LM Link; Bionic+ is $20 a month.

Ollama logo
Ollama Open Source

Ollama downloads and runs open AI models on your own macOS, Windows or Linux computer, with a command line, local API and desktop app. Local use is free; paid plans cover Ollama's cloud models.

Mistral Vibe logo

Mistral's Vibe, which Mistral's help center says replaces the Le Chat name, is an AI assistant for research, coding and everyday tasks. A Free plan has limited messages; Pro costs $14.99 a month.

D

DeepSeek is an AI chat assistant from a Chinese research company, offered as a mobile app, web chat and desktop client. The official site calls the app free to use; the API is billed per token.

Update history

  1. 2026-10-10 · ai alternative note updated (editor)
  2. 2026-10-10 · editor verdict updated (editor)
  3. 2026-10-10 · self hostable updated (editor)
  4. 2026-10-10 · api available updated (editor)
  5. 2026-10-10 · data residency note updated (editor)
  6. 2026-10-10 · data residency updated (editor)
  7. 2026-10-10 · pricing page url updated (editor)
  8. 2026-10-10 · free tier updated (editor)
  9. 2026-10-10 · pricing model updated (editor)

Beyond the tool

Need the outcome rather than the software?