Any model in.
One dashboard out.
Point OmniBench at any LLM — API, aggregator or local — and run a five-pillar battery with every run, task, turn, span, token and event recorded. Inspect, diff, replay and compare with full trust in how the test actually went.
Capability is table stakes. The other four are where models actually differ.
Every pillar writes to the same run store with stable metric IDs, units and direction — so the board can colour every cell and every comparison shows error bars, not just deltas.
Knowledge, math, code, reasoning, instruction-following and tool-calling on an unsaturated battery.
Truthfulness, hallucination, calibration (ECE), consistency across seeds and abstention behaviour.
Attack-success-rate across injection, jailbreak, leakage, toxicity, bias & agentic probes; OWASP mapped.
Fixed ReAct / Plan-and-Execute scaffolds over a tiered task battery: useful vs. useless turns, HHI, recovery.
Tokens, cache hits, cost-per-solved-task, quality-per-dollar, TTFT / TPOT — priced at write time.
See the whole shape of a model, not one number.
Pillar profiles are extruded into stacked glass layers; the score matrix becomes a landscape you can orbit. Every bar links back to the run, task and span that produced it.
Walk the ladder for any number on the board.
Every score drills down to the exact trajectory that produced it: which tool was called with which arguments, which turns were useless, where the harness had to intervene, and what every span cost. Runs are append-only and stamped with an environment fingerprint, so two runs compare cleanly only when they should.
- Stable IDs — metric_id · task_id · content-hashed prompts · env_fingerprint
- Seeded trials — Wilson / t-intervals on every proportion; CI shown, not hidden
- Judge κ gate — every judge-produced cell carries the calibration κ of its run
- Failure taxonomy — 16 frozen labels — a failure map, not a raw score
Agent efficiency metrics no leaderboard reports.
Fixed scaffolds (L1 ReAct, L2 Plan-and-Execute) over a deterministic tool sandbox. Only the model varies — so turns, interventions and cost are comparable.
interventions ÷ task. 0 = fully autonomous. The autonomy-ceiling signal when plotted against tier.
duplicate · retry-without-change · error-no-recovery · off-task · loop — classified per tool turn.
priced at write time from cached / uncached / output / reasoning tokens. Local models use your $/Mtok.
every turn, tool call, error and hint recorded. Diff two models on the same task, turn by turn.
$ omni # interactive TUI
$ omni run --model openai:gpt-4o-mini --local # your machine, your key
│ run 8f2c… capability · reliability · safety · agency · economics
│ ✓ agency ag_t3_002 T3 3t 1.7s $0.00041
│ ✗ capability cap_math_004 T1 arithmetic_error 1t 363ms $0.00001
⠹ running ██████████████░░░░░░░░░░ 38/66 0m41s · eta ~0m30s $0.0123
└ finished 66 tasks · 1m12s · $0.0211
╭─ openai:gpt-4o-mini ───────────────────────────────╮
│ Omni score 84.8 █████████████████████████░░░░░ │
│ bundle omnibench-runs/omnibench-…-8f2c.bundle.json │
╰────────────────────────────────────────────────────╯
$ omni upload omnibench-runs/*.bundle.json --to https://board --as you
✓ imported · verified omni 84.8 ✓ fingerprint ✓ evidence 55/0✗ ✓ trace
$ omni eval 8f2c --threshold 0.8 # CI gate → exit 1 belowRuns from your terminal or your CI. Never requires a browser.
node cli/omni.mjs is a full TUI when run bare, and a scriptable CLI otherwise. Drive a server through its REST API, or run the harness locally with your own keys — no server needed. Every finished run becomes a portable omnibench.run/1 bundle you can omni upload to a community board, where it is re-scored from raw evidence and pooled into per-model averages with confidence intervals.
The dashboard is just a client.
REST reflects the committed run store; the live stream is the single source of truth during a run. Build your own views on top.
Stop trusting leaderboards. Start reading traces.
Three offline reference models are pre-benchmarked so you can explore every view immediately. Add an API key to benchmark your own.