CLI / TUI

Benchmark from your terminal

The same harness that powers this dashboard runs from a single Node script. Use it against this server, or fully offline with your own API keys — then share the result so it counts toward the community averages.

Benchmark any model — step by step

From clone to community-verified result in five commands

  1. 1

    Get the CLI

    One zero-dependency Node script — no global install. It ships inside this repo.

    git clone https://github.com/AnshRoshan/llm-benchmark.git
    cd llm-benchmark && npm install
    node cli/omni.mjs --help
  2. 2

    Pick where the harness runs

    Remote = runs execute on the server and appear in this dashboard instantly. Local = your machine, your API keys, no server needed.

    node cli/omni.mjs doctor --url http://localhost:3000   # check a server
    node cli/omni.mjs doctor --local                        # or check your own setup
  3. 3

    Run the 5-pillar battery

    66 frozen tasks per full pass — capability, reliability, safety, agency, economics — with live per-pillar progress. Deterministic: same seed → same run.

    node cli/omni.mjs run --model sim:swift-budget --url http://localhost:3000
    node cli/omni.mjs run --model openai:gpt-4o-mini --local
    node cli/omni.mjs run --model ollama:llama3.1 --local   --pillar capability --pillar agency --trials 3 --seed 42
  4. 4

    Walk the recording

    Every run is a full-fidelity recording: task results → turns → spans → tokens → events. Drop from run summary to a single tool call.

    node cli/omni.mjs runs
    node cli/omni.mjs inspect <run-id> --task ag_t3_002
    node cli/omni.mjs diff sim:atlas-frontier sim:swift-budget
  5. 5

    Share it with everybody

    Publish a bundle to any OmniBench board. The server validates it and re-scores it from raw evidence before it moves the community averages.

    node cli/omni.mjs push <run-id> --to https://<board>   --as you --note "RTX 4090 · q4_K_M"
  6. 6

    Gate your CI on it

    Export reports in five formats and fail the pipeline when quality regresses.

    node cli/omni.mjs export <run-id> -o report.html -o results.junit.xml
    node cli/omni.mjs eval <run-id> --threshold 0.8        # exit 1 below 0.8
Environment: OMNI_URL target server (default http://localhost:3000) · OMNI_BOARD_URL default --to for uploads · OMNI_SUBMITTER default --as · DATABASE_URL local-mode store (embedded Postgres when unset) · provider keys OPENAI_API_KEY · ANTHROPIC_API_KEY · OLLAMA_BASE_URL · OPENAI_COMPAT_*. Every command accepts --json for machine-readable output and --url to override the target.

Remote mode

Drive this server through its REST API

export OMNI_URL=https://<this-host>
node cli/omni.mjs run --model sim:atlas-frontier
node cli/omni.mjs runs
node cli/omni.mjs inspect <run-id>

Runs execute on the server and stream back over SSE. They appear in the dashboard immediately.

Local mode

No server — your machine, your keys

export DATABASE_URL=postgres://…
export OPENAI_API_KEY=sk-…
node cli/omni.mjs run --model openai:gpt-4o-mini --local
# → omnibench-runs/omnibench-openai_gpt-4o-mini-<id>.bundle.json

The harness runs in-process via tsx and writes a portable bundle when it finishes. Ollama / vLLM / OpenRouter work through ollama: and openai-compat:.

Share mode

Publish to a community board

node cli/omni.mjs upload omnibench-runs/*.bundle.json \
  --to https://<this-host> --as "you" --note "M3 Max, q4"

# or push a run living on your own server
node cli/omni.mjs push <run-id> --to https://<this-host>

Bundles are validated, de-duplicated by content hash, and re-scored from raw task evidence before they count toward the pooled averages.

Command reference

omni v0.4.0

omniinteractive TUI — guided run wizard, run browser, boards, upload
omni run --model <p:m> [--local]run the 5-pillar battery; live per-pillar progress
omni run -c omni.yamldeclarative, version-controlled config (multiple models)
omni runs [--local]list runs (server, or your local DB)
omni inspect <run> [--task id]metrics with CIs, per-task verdicts, turns and spans
omni diff <a> <b>stat-sig metric deltas between two models
omni pull <run> [-o file]download a portable bundle from a server
omni bundle <run> [-o dir]build a bundle straight from DATABASE_URL
omni upload <file> --to <host>publish a bundle to any community board
omni push <run> --to <host>pull + upload in one step
omni export <run> -o x.html -o x.junit.xmlreports: json · jsonl · csv · junit · html · bundle
omni eval <run> --threshold 0.8CI gate — exit 1 when below threshold
omni board | community | models | tasksread-only views of the store
omni doctorcheck server, database, API keys
--json / --url <base>machine-readable output · target another server

The bundle format

omnibench.run/1

A single self-describing JSON file with the whole recording ladder, so anyone can audit exactly how a test went:

{
  "format": "omnibench.run/1",
  "run":        { modelId, config, configHash, envFingerprint, harnessVersion… },
  "model":      { id, provider, name },
  "taskResults":[ { taskId, status, score, failureLabel, finalOutput, costUsd… } ],
  "turns":      [ … ],      "spans":       [ … ],
  "tokenUsage": [ … ],      "events":      [ … ],
  "judgeResults":[ … ],     "scores":      [ … ]   // informational only
}
harness 0.4.0task set 2026.09-v1json · ≤ 40 MB

Verification on import

  • • harness / task-set match — same battery version as this server
  • • env fingerprint match — judge, sandbox and pricing versions agree
  • • scores agree — server recomputes Omni from raw task results; claimed value must be within 0.5
  • • evidence re-scored — every deterministic task's raw output is re-graded against the battery; any verdict that doesn't match is a mismatch
  • • trace present — turns and spans included so anyone can inspect the trajectory
  • • content hash — identical evidence is never counted twice

Unverified uploads are still stored and browsable, but flagged on every board.

Providers

Configured on this server

  • sim:sim:atlas-frontierready
  • echo:echoready
  • openai:openai:gpt-4o-minineeds OPENAI_API_KEY
  • anthropic:anthropic:claude-sonnet-4-5needs ANTHROPIC_API_KEY
  • ollama:ollama:llama3.1needs OLLAMA_BASE_URL
  • openai-compat:openai-compat:meta-llama/llama-3.3-70b-instructneeds OPENAI_COMPAT_BASE_URL