Trunchbull, run real models against any benchmark in your browser

Hi HN,<p>Today I&#x27;m showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configur

Share on:

Provider-neutral AI model benchmarking. Detect regressions before they reach production.

Trunchbull Skip to Main Content TRUNCHBULL Registry Pricing Products Demos Docs A Tool-Native Benchmarking Platform Run any model against any Benchmark Deploy tools from GitHub, run them against any OpenRouter model, and inspect every decision under controlled budgets. See how it works ↓ BENCHMARK READOUT / RUN 0842 2026-07-19 09:41:02 UTC LIVE Model Score Delta State Claude Opus 4 94.2 base GPT-5 91.8 -2.4% Gemini 2.5 Pro 88.5 -5.7% 2 REGRESSIONS DETECTED 14 TOOLS / $1.84 REMOTE TOOLS Isolated Workers MODELS OpenRouter / BYOK LIMITS Steps · tokens · cost EVIDENCE Full run traces 01 Tools as i

Related listings

design database schemas for your team and agents

Tests for your Claude Code setup that run on every release

China company code vs. US VIN check digit

a 3D imageboard where being there is the write permission

OpenTelemetry-native tracing for LLM apps and agents

retrieval methods visualized