Everything the ladder records
Per-model aggregates over the latest finished run: tier & difficulty curves, domain coverage, the full failure taxonomy, latency percentiles, turn classification, tool accuracy, span and event breakdowns, judge dimensions and trial consistency. Also exposed at /api/analytics.
Difficulty & tier curves
Pass rate on scored tasks · latest run per modelTier landscape
T1 single-step → T6 parallel/graph · pass-rate %
Pass rate by difficulty
Where each model's curve breaks
Pass rate by tier
Autonomy ceiling: the tier where success collapses
Domain coverage
22 domains across the batteryDomain heat-matrix
Pass-rate % per domain · colour = quality (rose → emerald)
Scaffold & pillar pass
L1 ReAct vs L2 Plan-and-Execute · per pillar
Failure taxonomy
16 frozen labels · a failure map, not a raw scoreFailure treemap
Pooled across models · area = count
Failure matrix
Count per label per model · taller = more failures
Agency mechanics
Only the model varies — scaffold, tools and judge are fixedTurn classification
useful vs useless vs harness interventions
Why turns were useless
duplicate · retry-without-change · error-no-recovery · off-task · loop
Tool accuracy & recovery
correct tool + args ÷ calls · errors recovered without intervention
Latency & token anatomy
Per task result · latest run per modelLatency bands
median bar · p95 band · dashed max
Token mix
Share of uncached / cached / output / reasoning
Instrumentation
SPAN · EVENT · JUDGE_RESULT levels of the ladderSpan breakdown
count · mean duration · errors · mean TTFT
| Model | Type | n | mean | err | TTFT |
|---|---|---|---|---|---|
| Swift Budget (simulated) | tool_call | 27 | 6ms | 1 | — |
| Swift Budget (simulated) | llm_call | 106 | 423ms | 0 | 310ms |
| Atlas Frontier (simulated) | tool_call | 28 | 7ms | 2 | — |
| Atlas Frontier (simulated) | llm_call | 105 | 6.76s | 0 | 665ms |
| Local Llama 8B (simulated) | tool_call | 45 | 6ms | 2 | — |
| Local Llama 8B (simulated) | llm_call | 135 | 1.44s | 0 | 907ms |
Event stream
Pooled counts by event type
Judge calibration
κ, mean score, rubric dimensions and modes
Consistency & profiles
Seeded trials · status disagreement across trials of the same taskPillar profile overlay
All benchmarked models
Trial consistency
tasks with >1 trial · disagreeing outcomes