
6.4x faster than llama.cpp, 3.9x faster than MLX

BaseRT is the fastest LLM runtime on Apple Silicon. Install it with one command and run local models on your own device.
Get BaseRT · Base Compute Runtime Apple NVIDIA AMD Research Enterprise Careers Blog Contact BaseRT Meet BaseRT, the fastest runtime on Apple Silicon $ curl -LsSf https://basecompute.co/install.sh | sh Copy Technical Reports BaseRT docs GitHub Discord Benchmarks Faster than MLX and llama.cpp On Prefill, up to 6.4x vs llama.cpp and 3.9x vs MLX. Up to 1.33x on Decode. BaseRT MLX Llama.cpp Decode · tg128 Qwen3 0.6B · Q4 531 398 +33% 386 +37% Llama 3.2 1B · Q4 342 298 +15% 267 +28% Llama 3.2 3B · Q4 137 131 +5% 120 +14% Prefill Qwen3 30B-A3B · pp128 · Q4 2,478 639 +288% 1,280 +94% Gemma 4 E2B · pp2
BaseRT is the fastest LLM runtime on Apple Silicon. Install it with one command and run local models on your own device.
Get BaseRT · Base Compute Runtime Apple NVIDIA AMD Research Enterprise Careers Blog Contact BaseRT Meet BaseRT, the fastest runtime on Apple Silicon $ curl -LsSf https://basecompute.co/install.sh | sh Copy Technical Reports BaseRT docs GitHub Discord Benchmarks Faster than MLX and llama.cpp On Prefill, up to 6.4x vs llama.cpp and 3.9x vs MLX. Up to 1.33x on Decode. BaseRT MLX Llama.cpp Decode · tg128 Qwen3 0.6B · Q4 531 398 +33% 386 +37% Llama 3.2 1B · Q4 342 298 +15% 267 +28% Llama 3.2 3B · Q4 137 131 +5% 120 +14% Prefill Qwen3 30B-A3B · pp128 · Q4 2,478 639 +288% 1,280 +94% Gemma 4 E2B · pp2