llama.cpp Optimization Tips: Speed Up Local LLM Inference on CPU and GPU

📘 Tutorials 2026-07-20 2 min read

llama.cpp is the most popular C++ implementation for running LLMs locally (76K stars). Default settings aren't optimal. These tips can double your inference speed.

💡 What You Will Learn

llama.cpp is the most popular C++ implementation for running LLMs locally (76K stars). Default settings aren't optimal. These tips can double your inference speed.

📜 Table of Contents

llama.cpp Optimization Tips: Speed Up Local LLM Inference on CPU and GPU

llama.cpp (76K GitHub stars) is the gold standard for running LLMs locally. Here are the optimization tips that make a real difference.

The Most Important Optimization: Quantization

Quant Size (7B) Speed Quality Loss
FP16 14GB 1x 0%
Q8_0 7.5GB 1.2x <1%
Q5_K_M 5.2GB 1.5x ~2%
Q4_K_M 4.2GB 2x ~3%
Q2_K 2.7GB 3x ~15%
Q4_K_M is the sweet spot. 75% smaller, 2x faster, barely noticeable quality loss.
## CPU Optimization
### Thread Count
# Auto-detect or set to physical cores
./llama-cli -m model.q4_k_m.gguf -p "Hello" -t 0

Set -t to number of physical cores, not logical threads.

Memory Bandwidth Matters Most

GPU Acceleration

NVIDIA CUDA

cmake -B build -DLLAMA_CUDA=ON
cmake --build build --config Release
./llama-cli -m model.q4_k_m.gguf -ngl 99

-ngl 99 puts all layers on GPU. For partial offloading: -ngl 24.

AMD ROCm / Apple Metal / Intel

# AMD
cmake -B build -DLLAMA_HIPBLAS=ON
# Apple
cmake -B build -DLLAMA_METAL=ON

Advanced Flags

Flag Effect
--mlock Lock model in RAM, prevent swapping
--flash-attn Flash Attention for long context
--cont-batching Continuous batching (server mode)
## Server Mode
./llama-server -m model.q4_k_m.gguf --host 0.0.0.0 --port 8080 -ngl 99 --parallel 4 --cont-batching

API at http://localhost:8080/v1/chat/completions (OpenAI compatible).

Real-World Benchmarks

Hardware Model Tokens/sec
MacBook Air M1 (8GB) Llama 3.1 8B Q4 12 t/s
MacBook Pro M3 Max (64GB) Llama 3.1 70B Q4 15 t/s
RTX 4090 (24GB) Llama 3.1 8B Q4 85 t/s
CPU only (Ryzen 7950X) Qwen3 7B Q4 9 t/s
## My Optimization Sequence
1. Pick Q4_K_M quantization
2. Use GPU with -ngl 99 if VRAM allows
3. Set threads to physical core count
4. Enable mlock
5. For server mode: add --cont-batching --parallel 4
## FAQ
Q: Is llama.cpp faster than Ollama? A: Same speed — Ollama uses llama.cpp under the hood.
Q: Can I run GGUF models from HuggingFace? A: Yes.
Q: Minimum RAM for 70B Q4? A: ~42GB model + 8GB context = 50GB minimum.

❓ FAQ

Is llama.cpp faster than Ollama?

Same speed — Ollama uses llama.cpp under the hood.

Minimum RAM for 70B Q4?

~42GB model + 8GB context = 50GB minimum.

Related Articles
2026-08-11
Speculative Decoding 2026: How Draft Models Make LLMs 2-3x Faster
2026-08-14
Ollama Models List 2026: 30+ Models Explained With Real Sizes
2026-07-16
AI Agent Tool Description Engineering 2026

Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.

💬 Comments (0)

No comments yet. Be the first!

Login to comment