Board

Multi-model matrix

Latest finished run per model · cells are heat-coloured by direction-aware quality · click a header to sort. Fairness: same harness, tasks, tools and judge; costs normalized; judge κ reported.

Board as a landscape

Rows = models (sorted as the table) · columns = Omni + five pillar composites · orbit, hover for exact values

Rendering 3D view…
#ModelEloOmni score ↓Capability score Reliability score Safety score Agency score Economics score Battery pass rate Task success Hand-Holding Index OWASP pass rate ECE Cost / solved task TTFT κ
1Swift Budget (simulated)
sim:swift-budget
151082.067.972.290.094.285.667.9%100.0%0.2590.0%0.333$0.00008310ms0.83
2Atlas Frontier (simulated)
sim:atlas-frontier
157079.596.491.7100.077.531.896.4%100.0%0.08100.0%0.080$0.0107665ms0.83
3Local Llama 8B (simulated)
sim:local-llama · local
142058.446.462.340.053.289.946.4%75.0%1.1740.0%0.347$0.00005907ms0.83

Pairwise Elo

Task-level win/loss/tie across shared tasks (trial 0), K=32, 4 passes

ModelEloWLT
sim:atlas-frontier157043287
sim:swift-budget1510232188
sim:local-llama142085173
sim:local-llama vs sim:swift-budget: 7–22 (37 ties)
sim:local-llama vs sim:atlas-frontier: 1–29 (36 ties)
sim:swift-budget vs sim:atlas-frontier: 1–14 (51 ties)

Efficiency frontier

Capability vs. cost per solved task — Pareto-optimal models highlighted