BaseRT logo

BaseRT

6.4x faster than llama.cpp, 3.9x faster than MLX

Share on:
Listing image

BaseRT is the fastest LLM runtime on Apple Silicon. Install it with one command and run local models on your own device.

Get BaseRT · Base Compute Runtime Apple NVIDIA AMD Research Enterprise Careers Blog Contact BaseRT Meet BaseRT, the fastest runtime on Apple Silicon $ curl -LsSf https://basecompute.co/install.sh | sh Copy Technical Reports BaseRT docs GitHub Discord Benchmarks Faster than MLX and llama.cpp On Prefill, up to 6.4x vs llama.cpp and 3.9x vs MLX. Up to 1.33x on Decode. BaseRT MLX Llama.cpp Decode · tg128 Qwen3 0.6B · Q4 531 398 +33% 386 +37% Llama 3.2 1B · Q4 342 298 +15% 267 +28% Llama 3.2 3B · Q4 137 131 +5% 120 +14% Prefill Qwen3 30B-A3B · pp128 · Q4 2,478 639 +288% 1,280 +94% Gemma 4 E2B · pp2

Related listings

AI-powered construction takeoffs, estimates, and invoices

Vapi for FaceTime: AI video agents in a few lines

Build evals and custom benchmarks for real-world tasks

A MagSafe AI Recorder That Acts for You

Spend up to 85% less and run 3× longer coding agent sessions

AI investing, done responsibly