Soup CLI logo

Soup CLI

Fine-tune an 8B LLM on a 4 GB laptop GPU

Share on:
Listing image

Fine-tune Llama-3.1-8B on a 4 GB laptop GPU, and align on the same card. Soup v0.72.4 streams the frozen base from CPU RAM one decoder layer at a time into a small pool of VRAM buffers while only the LoRA adapters stay resident, and quantizing that streamed base to NF4 shrinks it about fourfold, so peak VRAM is bounded by one layer instead of the whole model: measured on a 4 GB RTX 3050 Laptop, Llama-3.1-8B trains at 119.6 tok/s in 3.32 GB and Qwen2.5-3B at 264.2 tok/s in 1.76 GB. v0.72.4 opens that path to the preference losses, DPO, ORPO, SimPO and KTO, taking DPO's reference model from the same streamed base with its adapters switched off so it costs no extra weights: measured 0.914x the supervised peak on a 365M synthetic fixture, where forcing a real second instance cost 730.44 MB, exactly one copy of them. Underneath, it runs across nine architectures (Llama, Qwen, Mistral, Gemma, Phi) with batches above 1, gradient accumulation, resume, a batch- and vocabulary-aware VRAM pre-flight, and an NVMe disk overflow tier. It ships BETA, and every measurement behind it is published, discarded numbers included. Plus soup ship, one SHIP or DON'T-SHIP verdict scored over seven bundled offline suites that you can commit and bind to your weights, the Reward Forge verifier pair, the semantic data moat, the compliance pack, LoRA task arithmetic and Whisper fine-tuning. Lean PyTorch-free install, then [train]. On your own GPU.

Soup CLI — Build, expect, train, x-ray, merge, bisect, ship — one CLI Soup Zero soon Docs PyPI GitHub Discord Support $ pip install "soup-cli[train]" Get Started Free & Open Source Stop tuning. Start training. Soup writes the config for you _ The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config (task, quantization, LR and epochs come from rules, not a search), derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting. And when the model is bigger than the card, i

Related listings

AI-powered construction takeoffs, estimates, and invoices

Vapi for FaceTime: AI video agents in a few lines

Build evals and custom benchmarks for real-world tasks

A MagSafe AI Recorder That Acts for You

Spend up to 85% less and run 3× longer coding agent sessions

AI investing, done responsibly