llama.cpp Optimization Tips: Speed Up Local LLM Inference on CPU and GPU
llama.cpp is the most popular C++ implementation for running LLMs locally (76K stars). Default settings aren't optimal. These tips can double your inference speed.
💡 What You Will Learn
llama.cpp is the most popular C++ implementation for running LLMs locally (76K stars). Default settings aren't optimal. These tips can double your inference speed.
📜 Table of Contents
llama.cpp Optimization Tips: Speed Up Local LLM Inference on CPU and GPU
llama.cpp (76K GitHub stars) is the gold standard for running LLMs locally. Here are the optimization tips that make a real difference.
The Most Important Optimization: Quantization
| Quant | Size (7B) | Speed | Quality Loss |
|---|---|---|---|
| FP16 | 14GB | 1x | 0% |
| Q8_0 | 7.5GB | 1.2x | <1% |
| Q5_K_M | 5.2GB | 1.5x | ~2% |
| Q4_K_M | 4.2GB | 2x | ~3% |
| Q2_K | 2.7GB | 3x | ~15% |
| Q4_K_M is the sweet spot. 75% smaller, 2x faster, barely noticeable quality loss. | |||
| ## CPU Optimization | |||
| ### Thread Count |
# Auto-detect or set to physical cores
./llama-cli -m model.q4_k_m.gguf -p "Hello" -t 0
Set -t to number of physical cores, not logical threads.
Memory Bandwidth Matters Most
- DDR4-3200: ~25 GB/s → ~3 tok/s for 7B Q4
- DDR5-6000: ~50 GB/s → ~6 tok/s
- Apple M1/M2: ~100 GB/s → ~12 tok/s
GPU Acceleration
NVIDIA CUDA
cmake -B build -DLLAMA_CUDA=ON
cmake --build build --config Release
./llama-cli -m model.q4_k_m.gguf -ngl 99
-ngl 99 puts all layers on GPU. For partial offloading: -ngl 24.
AMD ROCm / Apple Metal / Intel
# AMD
cmake -B build -DLLAMA_HIPBLAS=ON
# Apple
cmake -B build -DLLAMA_METAL=ON
Advanced Flags
| Flag | Effect |
|---|---|
| --mlock | Lock model in RAM, prevent swapping |
| --flash-attn | Flash Attention for long context |
| --cont-batching | Continuous batching (server mode) |
| ## Server Mode |
./llama-server -m model.q4_k_m.gguf --host 0.0.0.0 --port 8080 -ngl 99 --parallel 4 --cont-batching
API at http://localhost:8080/v1/chat/completions (OpenAI compatible).
Real-World Benchmarks
| Hardware | Model | Tokens/sec |
|---|---|---|
| MacBook Air M1 (8GB) | Llama 3.1 8B Q4 | 12 t/s |
| MacBook Pro M3 Max (64GB) | Llama 3.1 70B Q4 | 15 t/s |
| RTX 4090 (24GB) | Llama 3.1 8B Q4 | 85 t/s |
| CPU only (Ryzen 7950X) | Qwen3 7B Q4 | 9 t/s |
| ## My Optimization Sequence | ||
| 1. Pick Q4_K_M quantization | ||
| 2. Use GPU with -ngl 99 if VRAM allows | ||
| 3. Set threads to physical core count | ||
| 4. Enable mlock | ||
| 5. For server mode: add --cont-batching --parallel 4 | ||
| ## FAQ | ||
| Q: Is llama.cpp faster than Ollama? A: Same speed — Ollama uses llama.cpp under the hood. | ||
| Q: Can I run GGUF models from HuggingFace? A: Yes. | ||
| Q: Minimum RAM for 70B Q4? A: ~42GB model + 8GB context = 50GB minimum. |
❓ FAQ
Is llama.cpp faster than Ollama?
Same speed — Ollama uses llama.cpp under the hood.
Minimum RAM for 70B Q4?
~42GB model + 8GB context = 50GB minimum.
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
