Task set 2026.09-v1 · 66 tasks · 45 frozen metrics · 16 failure labels

Any model in.
One dashboard out.

Point OmniBench at any LLM — API, aggregator or local — and run a five-pillar battery with every run, task, turn, span, token and event recorded. Inspect, diff, replay and compare with full trust in how the test actually went.

OpenAI Anthropic Ollama vLLM OpenRouter Any OpenAI-compatible
5
pillars
45
frozen metrics
66
curated tasks
6
recording levels
Five pillars, breadth-first

Capability is table stakes. The other four are where models actually differ.

Every pillar writes to the same run store with stable metric IDs, units and direction — so the board can colour every cell and every comparison shows error bars, not just deltas.

Capability

Knowledge, math, code, reasoning, instruction-following and tool-calling on an unsaturated battery.

9 metrics
Reliability

Truthfulness, hallucination, calibration (ECE), consistency across seeds and abstention behaviour.

7 metrics
Safety

Attack-success-rate across injection, jailbreak, leakage, toxicity, bias & agentic probes; OWASP mapped.

7 metrics
Agency

Fixed ReAct / Plan-and-Execute scaffolds over a tiered task battery: useful vs. useless turns, HHI, recovery.

11 metrics
Economics

Tokens, cache hits, cost-per-solved-task, quality-per-dollar, TTFT / TPOT — priced at write time.

11 metrics
Live from the run store

See the whole shape of a model, not one number.

Pillar profiles are extruded into stacked glass layers; the score matrix becomes a landscape you can orbit. Every bar links back to the run, task and span that produced it.

Open the board
Rendering 3D view…
Rendering 3D view…
Full-fidelity recording

Walk the ladder for any number on the board.

Every score drills down to the exact trajectory that produced it: which tool was called with which arguments, which turns were useless, where the harness had to intervene, and what every span cost. Runs are append-only and stamped with an environment fingerprint, so two runs compare cleanly only when they should.

  • Stable IDs — metric_id · task_id · content-hashed prompts · env_fingerprint
  • Seeded trials — Wilson / t-intervals on every proportion; CI shown, not hidden
  • Judge κ gate — every judge-produced cell carries the calibration κ of its run
  • Failure taxonomy — 16 frozen labels — a failure map, not a raw score
RUNconfig_hash · env_fingerprint · seeds
TASKtier · domain · scoring · failure_label
TURNmodel | tool | judge | intervention
SPANllm_call | tool_call | judge_call · duration · ttft
TOKEN_USAGEcached · uncached · output · reasoning · $
EVENTerror | retry | loop | permission_denied
+ SCORE and JUDGE_RESULT attached at the right level
The differentiator

Agent efficiency metrics no leaderboard reports.

Fixed scaffolds (L1 ReAct, L2 Plan-and-Execute) over a deterministic tool sandbox. Only the model varies — so turns, interventions and cost are comparable.

Hand-Holding Index

interventions ÷ task. 0 = fully autonomous. The autonomy-ceiling signal when plotted against tier.

Useful vs. useless turns

duplicate · retry-without-change · error-no-recovery · off-task · loop — classified per tool turn.

Cost per solved task

priced at write time from cached / uncached / output / reasoning tokens. Local models use your $/Mtok.

Trajectory replay

every turn, tool call, error and hint recorded. Diff two models on the same task, turn by turn.

omni — CLI
$ omni                              # interactive TUI
$ omni run --model openai:gpt-4o-mini --local     # your machine, your key
  │ run 8f2c…  capability · reliability · safety · agency · economics
  │ ✓ agency      ag_t3_002   T3                3t     1.7s  $0.00041
  │ ✗ capability  cap_math_004 T1 arithmetic_error 1t   363ms  $0.00001
  ⠹ running  ██████████████░░░░░░░░░░ 38/66  0m41s · eta ~0m30s  $0.0123
  └ finished 66 tasks · 1m12s · $0.0211
  ╭─ openai:gpt-4o-mini ───────────────────────────────╮
  │ Omni score  84.8  █████████████████████████░░░░░   │
  │ bundle  omnibench-runs/omnibench-…-8f2c.bundle.json │
  ╰────────────────────────────────────────────────────╯

$ omni upload omnibench-runs/*.bundle.json --to https://board --as you
  ✓ imported · verified  omni 84.8   ✓ fingerprint ✓ evidence 55/0✗ ✓ trace

$ omni eval 8f2c --threshold 0.8      # CI gate → exit 1 below
CLI-first, headless-safe

Runs from your terminal or your CI. Never requires a browser.

node cli/omni.mjs is a full TUI when run bare, and a scriptable CLI otherwise. Drive a server through its REST API, or run the harness locally with your own keys — no server needed. Every finished run becomes a portable omnibench.run/1 bundle you can omni upload to a community board, where it is re-scored from raw evidence and pooled into per-model averages with confidence intervals.

omni run --local
omni upload
omni push
omni pull
omni inspect
omni diff
omni eval
omni community
Frozen API contract

The dashboard is just a client.

REST reflects the committed run store; the live stream is the single source of truth during a run. Build your own views on top.

GET
/api/models
models with aggregate scores
GET
/api/models/{id}
one model: pillar scores + CI
GET
/api/analytics
tiers · difficulty · failures · latency · turns · spans · judge
GET
/api/runs
runs (id, model, created, fingerprint)
POST
/api/runs
start a run (202 Accepted)
GET
/api/runs/{id}/tasks/{task}
full trajectory: turns, spans, tokens, events
GET
/api/runs/{id}/trace
span tree for the waterfall
SSE
/api/runs/{id}/stream
span_finished · task_finished · run_finished
GET
/api/compare?models=a,b
aligned metric matrix w/ significance
PostgreSQL run store · Drizzle ORM export JSON · JSONL · CSV · JUnit · self-contained HTML WebGL visualisations via three.js

Stop trusting leaderboards. Start reading traces.

Three offline reference models are pre-benchmarked so you can explore every view immediately. Add an API key to benchmark your own.