Overview

Benchmark control room

Every number here is a view over the run store. Click through any model, run or task to walk the recording ladder down to individual spans.

Models
3
3 benchmarked
Runs
5
all settled
Task results
330
151 passed · 47 failed
Spans recorded
759
463 events · 0 judge calls
Total spend
$0.5851
106.7k tokens · priced at write time

Score landscape

Latest finished run per model · pillar composites 0–100 · drag to orbit, hover for values

Rendering 3D view…

Leader · Swift Budget (simulated)

Omni score and pillar gauges

82
omni
Elo 1510
68
Capa
72
Reli
90
Safe
94
Agen
86
Econ
Open model →

Omni score across runs

Every finished run, chronological

Spend per run

USD, priced at write time

Task outcomes

Latest run per model, pooled

76%
pass rate

Recent runs

All runs →
RunModelPillarsStatusOmniWhen
f2f54103sim:local-llama
CapabilityReliabilitySafetyAgencyEconomics
finished58.422s ago
f5e098e1sim:local-llama
CapabilityReliabilitySafetyAgencyEconomics
finished58.422s ago
1504bf5fsim:swift-budget
CapabilityReliabilitySafetyAgencyEconomics
finished82.023s ago
1d732c0dsim:swift-budget
CapabilityReliabilitySafetyAgencyEconomics
finished82.023s ago
422ad29bsim:atlas-frontier
CapabilityReliabilitySafetyAgencyEconomics
finished79.524s ago

Pillars

What each score composes

  • Capability
    Knowledge, math, code, reasoning, instruction-following and tool-calling on an unsaturated battery.
  • Reliability
    Truthfulness, hallucination, calibration (ECE), consistency across seeds and abstention behaviour.
  • Safety
    Attack-success-rate across injection, jailbreak, leakage, toxicity, bias & agentic probes; OWASP mapped.
  • Agency
    Fixed ReAct / Plan-and-Execute scaffolds over a tiered task battery: useful vs. useless turns, HHI, recovery.
  • Economics
    Tokens, cache hits, cost-per-solved-task, quality-per-dollar, TTFT / TPOT — priced at write time.