Deep dive

Everything the ladder records

Per-model aggregates over the latest finished run: tier & difficulty curves, domain coverage, the full failure taxonomy, latency percentiles, turn classification, tool accuracy, span and event breakdowns, judge dimensions and trial consistency. Also exposed at /api/analytics.

Task results
198
76% pass rate
Turns
464
model · tool · judge · intervention
Spans
446
llm_call · tool_call · judge_call
Events
463
error · retry · loop · intervention…
Judge calls
0
pointwise · rubric · pass/fail
Tokens
106.7k
$0.5803

Difficulty & tier curves

Pass rate on scored tasks · latest run per model

Tier landscape

T1 single-step → T6 parallel/graph · pass-rate %

Rendering 3D view…

Pass rate by difficulty

Where each model's curve breaks

Pass rate by tier

Autonomy ceiling: the tier where success collapses

Domain coverage

22 domains across the battery

Domain heat-matrix

Pass-rate % per domain · colour = quality (rose → emerald)

Rendering 3D view…

Scaffold & pillar pass

L1 ReAct vs L2 Plan-and-Execute · per pillar

Swift Budget (simulated)omni 82.0
L1100% n=10
L2100% n=6
Capability68%
Reliability69%
Safety90%
Agency100%
Economics—
Atlas Frontier (simulated)omni 79.5
L1100% n=10
L2100% n=6
Capability96%
Reliability94%
Safety100%
Agency100%
Economics—
Local Llama 8B (simulated)omni 58.4
L190% n=10
L250% n=6
Capability46%
Reliability63%
Safety40%
Agency75%
Economics—

Failure taxonomy

16 frozen labels · a failure map, not a raw score

Failure treemap

Pooled across models · area = count

Failure matrix

Count per label per model · taller = more failures

Rendering 3D view…

Agency mechanics

Only the model varies — scaffold, tools and judge are fixed

Turn classification

useful vs useless vs harness interventions

Why turns were useless

duplicate · retry-without-change · error-no-recovery · off-task · loop

17
useless turns
duplicate · 5loop · 12

Tool accuracy & recovery

correct tool + args ÷ calls · errors recovered without intervention

Swift Budget (simulated)
96
tool
100
recov
26/27 calls · 1/1 errors
Atlas Frontier (simulated)
96
tool
0
recov
27/28 calls · 0/2 errors
Local Llama 8B (simulated)
58
tool
50
recov
26/45 calls · 1/2 errors

Latency & token anatomy

Per task result · latest run per model

Latency bands

median bar · p95 band · dashed max

Token mix

Share of uncached / cached / output / reasoning

Instrumentation

SPAN · EVENT · JUDGE_RESULT levels of the ladder

Span breakdown

count · mean duration · errors · mean TTFT

ModelTypenmeanerrTTFT
Swift Budget (simulated)tool_call276ms1—
Swift Budget (simulated)llm_call106423ms0310ms
Atlas Frontier (simulated)tool_call287ms2—
Atlas Frontier (simulated)llm_call1056.76s0665ms
Local Llama 8B (simulated)tool_call456ms2—
Local Llama 8B (simulated)llm_call1351.44s0907ms

Event stream

Pooled counts by event type

463
events
task_finished198
task_started198
replan18
intervention18
loop12
error5
retry5
run_started3
judge_calibrated3
run_finished3

Judge calibration

κ, mean score, rubric dimensions and modes

Swift Budget (simulated)
κ 0.830 calls
mean score—
Atlas Frontier (simulated)
κ 0.830 calls
mean score—
Local Llama 8B (simulated)
κ 0.830 calls
mean score—

Consistency & profiles

Seeded trials · status disagreement across trials of the same task

Pillar profile overlay

All benchmarked models

Rendering 3D view…

Trial consistency

tasks with >1 trial · disagreeing outcomes

Swift Budget (simulated)1 trial
0/0 multi-trial tasks · statuses: failed 15 · passed 51
open run 1d732c0d →
Atlas Frontier (simulated)1 trial
0/0 multi-trial tasks · statuses: passed 64 · failed 2
open run 422ad29b →
Local Llama 8B (simulated)1 trial
0/0 multi-trial tasks · statuses: passed 36 · failed 30
open run f5e098e1 →