Omni score
58.4
run f5e098e1 · 22s ago
Capability
46.4
Reliability
62.3
Safety
40.0
Agency
53.2
Economics
89.9
At a glance
Five-pillar profile · extruded glass radar
Rendering 3D view…
58
omni
Cost / solved task
$0.00005
Hand-holding index
1.17
OWASP pass rate
40%
Judge κ
0.83
Headline metrics with confidence
Wilson 95% CI · task set 2026.09-v1 · fingerprint b7a221855bf39c11
Capability
| Code | 93.8% | n=4 |
| Instruction-following | 50.0% | n=4 |
| Knowledge | 20.0% | n=5 |
| Long-context | 50.0% | n=2 |
| Math | 50.0% | n=6 |
| Reasoning | 20.0% | n=5 |
| Tool-calling | 100.0% | n=2 |
| Battery pass rate | 46.4% | 29.5% – 64.2% |
Reliability
| Omniscience index | 5 | n=11 |
| Consistency variance | 16.7% | n=2 |
| ECE | 0.347 | n=3 |
| Hallucination (extrinsic) | 50.0% | 9.5% – 90.5% |
| Hallucination (intrinsic) | 100.0% | 34.2% – 100.0% |
| Over-refusal | 0.0% | 0.0% – 56.2% |
| Truthfulness | 75.0% | 30.1% – 95.4% |
Safety
| ASR · agentic | 50.0% | 9.5% – 90.5% |
| ASR · bias | 100.0% | 20.7% – 100.0% |
| ASR · jailbreak | 50.0% | 9.5% – 90.5% |
| ASR · leakage | 100.0% | 34.2% – 100.0% |
| ASR · injection | 0.0% | 0.0% – 65.8% |
| ASR · toxicity | 100.0% | 20.7% – 100.0% |
| OWASP pass rate | 40.0% | 16.8% – 68.7% |
Agency
| Error recovery | 50.0% | n=12 |
| Failure to stop | 0.0% | 0.0% – 24.3% |
| Hand-Holding Index | 1.17 | n=12 |
| Premature stop | 16.7% | 4.7% – 44.8% |
| Task success | 75.0% | 46.8% – 91.1% |
| Time to completion | 8.9s | n=12 |
| Tool-call accuracy | 56.1% | 41.0% – 70.1% |
| Turns | 10.7 | n=12 |
| Useful turns | 8.1 | n=12 |
| Useless turns | 1.4 | n=12 |
| Useful-turn ratio | 85.1% | n=12 |
Economics
| Cache hit rate | 0.0% | n=66 |
| Cost / solved task | $0.00005 | n=32 |
| Total cost | $0.00162 | n=66 |
| Quality per $ | 28,714 | n=66 |
| Throughput | 8 t/s | n=135 |
| Cached input tok | 0.0 | n=66 |
| Uncached input tok | 30,762 | n=66 |
| Output tok | 1,577 | n=66 |
| Reasoning tok | 0.0 | n=66 |
| TPOT | 46ms | n=135 |
| TTFT | 907ms | n=135 |
Failure map
30 failing task results in the latest run, grouped by frozen label
hallucination_factual8Fabricated fact (intrinsic)reasoning_error5CoT contradicts answer / bad chainarithmetic_error4Math errorpolicy_violation3Broke a task policy / constraintleakage2Safety: PII / prompt leakedpremature_stop2Stopped before completioninjection_success2Safety: probe succeededhallucination_unsupported1Claim not grounded in context (extrinsic)context_loss1Lost needed info in long contextschema_violation1Malformed / unparseable outputargument_error1Right tool, wrong arguments
Run history
Append-only; every run keeps its own fingerprint