Community

Crowd-sourced leaderboard

Every finished run of a model — whether executed on this server or uploaded from someone's own machine with their own API key — is pooled here. Scores are recomputed from the raw task evidence on import; we show the mean with a 95% confidence interval so you can see how stable each model really is.

Models
3
with ≥1 finished run
Pooled runs
5
server + uploads
Uploads
2
community bundles
Contributors
3
distinct submitters

Aggregate board

Ranked by mean Omni score across server runs + verified uploads. Unverified uploads are listed but never move the average. Wider bars = less agreement between runs.

#ModelOmni (mean)95% CICapRelSafAgeEcoRunsPeopleCost/runTrendLast
1Swift Budget (simulated)
sim:swift-budget 1 verified
82.0
σ 0.0
82.0 – 82.0
68729094862 (1↑)2$0.0031822s ago
2Atlas Frontier (simulated)
sim:atlas-frontier
79.5
79.5 – 79.5
9692100783211$0.5756single run23s ago
3Local Llama 8B (simulated)
sim:local-llama 1 verified
58.4
σ 0.0
58.4 – 58.4
46624053902 (1↑)2$0.0016222s ago

Upload a run bundle

Anyone can contribute. Bundles are validated, de-duplicated, and re-scored from raw task results — claimed scores are never trusted.

Bundle upload is disabled in this demo

This GitHub Pages build is a static snapshot of the OmniBench dashboard, seeded with the three offline reference models. Starting runs, live traces and community uploads need the full server build — run npm install && npx drizzle-kit push && npm run dev locally, or point the CLI at any hosted instance.

From the CLI
# run locally with your own key, then share the result
export OPENAI_API_KEY=sk-...
node cli/omni.mjs run --model openai:gpt-4o-mini --local        # → omnibench-runs/<id>.bundle.json
node cli/omni.mjs upload omnibench-runs/<id>.bundle.json --to https://<this-host> --as "you"

# or push a run straight from your own OmniBench server
node cli/omni.mjs push <run-id> --to https://<this-host> --as "you" --note "M3 Max, ollama q4"

Recent submissions

2 most recent