Benchmark from your terminal
The same harness that powers this dashboard runs from a single Node script. Use it against this server, or fully offline with your own API keys — then share the result so it counts toward the community averages.
Benchmark any model — step by step
From clone to community-verified result in five commands
- 1
Get the CLI
One zero-dependency Node script — no global install. It ships inside this repo.
git clone https://github.com/AnshRoshan/llm-benchmark.git cd llm-benchmark && npm install node cli/omni.mjs --help
- 2
Pick where the harness runs
Remote = runs execute on the server and appear in this dashboard instantly. Local = your machine, your API keys, no server needed.
node cli/omni.mjs doctor --url http://localhost:3000 # check a server node cli/omni.mjs doctor --local # or check your own setup
- 3
Run the 5-pillar battery
66 frozen tasks per full pass — capability, reliability, safety, agency, economics — with live per-pillar progress. Deterministic: same seed → same run.
node cli/omni.mjs run --model sim:swift-budget --url http://localhost:3000 node cli/omni.mjs run --model openai:gpt-4o-mini --local node cli/omni.mjs run --model ollama:llama3.1 --local --pillar capability --pillar agency --trials 3 --seed 42
- 4
Walk the recording
Every run is a full-fidelity recording: task results → turns → spans → tokens → events. Drop from run summary to a single tool call.
node cli/omni.mjs runs node cli/omni.mjs inspect <run-id> --task ag_t3_002 node cli/omni.mjs diff sim:atlas-frontier sim:swift-budget
- 5
Share it with everybody
Publish a bundle to any OmniBench board. The server validates it and re-scores it from raw evidence before it moves the community averages.
node cli/omni.mjs push <run-id> --to https://<board> --as you --note "RTX 4090 · q4_K_M"
- 6
Gate your CI on it
Export reports in five formats and fail the pipeline when quality regresses.
node cli/omni.mjs export <run-id> -o report.html -o results.junit.xml node cli/omni.mjs eval <run-id> --threshold 0.8 # exit 1 below 0.8
Remote mode
Drive this server through its REST API
export OMNI_URL=https://<this-host> node cli/omni.mjs run --model sim:atlas-frontier node cli/omni.mjs runs node cli/omni.mjs inspect <run-id>
Runs execute on the server and stream back over SSE. They appear in the dashboard immediately.
Local mode
No server — your machine, your keys
export DATABASE_URL=postgres://… export OPENAI_API_KEY=sk-… node cli/omni.mjs run --model openai:gpt-4o-mini --local # → omnibench-runs/omnibench-openai_gpt-4o-mini-<id>.bundle.json
The harness runs in-process via tsx and writes a portable bundle when it finishes. Ollama / vLLM / OpenRouter work through ollama: and openai-compat:.
Share mode
Publish to a community board
node cli/omni.mjs upload omnibench-runs/*.bundle.json \ --to https://<this-host> --as "you" --note "M3 Max, q4" # or push a run living on your own server node cli/omni.mjs push <run-id> --to https://<this-host>
Bundles are validated, de-duplicated by content hash, and re-scored from raw task evidence before they count toward the pooled averages.
Command reference
omni v0.4.0
| omni | interactive TUI — guided run wizard, run browser, boards, upload |
| omni run --model <p:m> [--local] | run the 5-pillar battery; live per-pillar progress |
| omni run -c omni.yaml | declarative, version-controlled config (multiple models) |
| omni runs [--local] | list runs (server, or your local DB) |
| omni inspect <run> [--task id] | metrics with CIs, per-task verdicts, turns and spans |
| omni diff <a> <b> | stat-sig metric deltas between two models |
| omni pull <run> [-o file] | download a portable bundle from a server |
| omni bundle <run> [-o dir] | build a bundle straight from DATABASE_URL |
| omni upload <file> --to <host> | publish a bundle to any community board |
| omni push <run> --to <host> | pull + upload in one step |
| omni export <run> -o x.html -o x.junit.xml | reports: json · jsonl · csv · junit · html · bundle |
| omni eval <run> --threshold 0.8 | CI gate — exit 1 when below threshold |
| omni board | community | models | tasks | read-only views of the store |
| omni doctor | check server, database, API keys |
| --json / --url <base> | machine-readable output · target another server |
The bundle format
omnibench.run/1
A single self-describing JSON file with the whole recording ladder, so anyone can audit exactly how a test went:
{
"format": "omnibench.run/1",
"run": { modelId, config, configHash, envFingerprint, harnessVersion… },
"model": { id, provider, name },
"taskResults":[ { taskId, status, score, failureLabel, finalOutput, costUsd… } ],
"turns": [ … ], "spans": [ … ],
"tokenUsage": [ … ], "events": [ … ],
"judgeResults":[ … ], "scores": [ … ] // informational only
}Verification on import
- • harness / task-set match — same battery version as this server
- • env fingerprint match — judge, sandbox and pricing versions agree
- • scores agree — server recomputes Omni from raw task results; claimed value must be within 0.5
- • evidence re-scored — every deterministic task's raw output is re-graded against the battery; any verdict that doesn't match is a mismatch
- • trace present — turns and spans included so anyone can inspect the trajectory
- • content hash — identical evidence is never counted twice
Unverified uploads are still stored and browsable, but flagged on every board.
Providers
Configured on this server
- sim:sim:atlas-frontierready
- echo:echoready
- openai:openai:gpt-4o-minineeds OPENAI_API_KEY
- anthropic:anthropic:claude-sonnet-4-5needs ANTHROPIC_API_KEY
- ollama:ollama:llama3.1needs OLLAMA_BASE_URL
- openai-compat:openai-compat:meta-llama/llama-3.3-70b-instructneeds OPENAI_COMPAT_BASE_URL