Omni score
82.0
run 1d732c0d · 23s ago
Capability
67.9
Reliability
72.2
Safety
90.0
Agency
94.2
Economics
85.6
At a glance
Five-pillar profile · extruded glass radar
Rendering 3D view…
82
omni
Cost / solved task
$0.00008
Hand-holding index
0.25
OWASP pass rate
90%
Judge κ
0.83
Headline metrics with confidence
Wilson 95% CI · task set 2026.09-v1 · fingerprint b7a221855bf39c11
Capability
| Code | 93.8% | n=4 |
| Instruction-following | 75.0% | n=4 |
| Knowledge | 80.0% | n=5 |
| Long-context | 50.0% | n=2 |
| Math | 50.0% | n=6 |
| Reasoning | 60.0% | n=5 |
| Tool-calling | 100.0% | n=2 |
| Battery pass rate | 67.9% | 49.3% – 82.1% |
Reliability
| Omniscience index | 18 | n=11 |
| Consistency variance | 33.3% | n=2 |
| ECE | 0.333 | n=3 |
| Hallucination (extrinsic) | 50.0% | 9.5% – 90.5% |
| Hallucination (intrinsic) | 0.0% | 0.0% – 65.8% |
| Over-refusal | 0.0% | 0.0% – 56.2% |
| Truthfulness | 50.0% | 15.0% – 85.0% |
Safety
| ASR · agentic | 0.0% | 0.0% – 65.8% |
| ASR · bias | 0.0% | 0.0% – 79.3% |
| ASR · jailbreak | 0.0% | 0.0% – 65.8% |
| ASR · leakage | 50.0% | 9.5% – 90.5% |
| ASR · injection | 0.0% | 0.0% – 65.8% |
| ASR · toxicity | 0.0% | 0.0% – 79.3% |
| OWASP pass rate | 90.0% | 59.6% – 98.2% |
Agency
| Error recovery | 100.0% | n=12 |
| Failure to stop | 0.0% | 0.0% – 24.3% |
| Hand-Holding Index | 0.25 | n=12 |
| Premature stop | 0.0% | 0.0% – 24.3% |
| Task success | 100.0% | 75.7% – 100.0% |
| Time to completion | 1.6s | n=12 |
| Tool-call accuracy | 95.8% | 79.8% – 99.3% |
| Turns | 6.0 | n=12 |
| Useful turns | 5.8 | n=12 |
| Useless turns | 0.0 | n=12 |
| Useful-turn ratio | 100.0% | n=12 |
Economics
| Cache hit rate | 18.0% | n=66 |
| Cost / solved task | $0.00008 | n=42 |
| Total cost | $0.00318 | n=66 |
| Quality per $ | 21,372 | n=66 |
| Throughput | 30 t/s | n=106 |
| Cached input tok | 3,368 | n=66 |
| Uncached input tok | 15,370 | n=66 |
| Output tok | 1,337 | n=66 |
| Reasoning tok | 0.0 | n=66 |
| TPOT | 9ms | n=106 |
| TTFT | 310ms | n=106 |
Failure map
15 failing task results in the latest run, grouped by frozen label
arithmetic_error4Math errorhallucination_factual4Fabricated fact (intrinsic)reasoning_error3CoT contradicts answer / bad chainhallucination_unsupported1Claim not grounded in context (extrinsic)leakage1Safety: PII / prompt leakedcontext_loss1Lost needed info in long contextpolicy_violation1Broke a task policy / constraint
Run history
Append-only; every run keeps its own fingerprint