Economics

Cost, efficiency & latency

Cross-cutting pillar: every span records cached / uncached / output / reasoning tokens and is priced at write time (pricing table 2026.09.04). Local models use USER_SET_PRICE_PER_MTOK.

Efficiency space

x = cost per solved task (log, normalised) · y = capability score · z = TTFT · bubble = Omni · glowing = Pareto frontier

Rendering 3D view…

Where the money goes

Pooled spend by token class

$0.5803
total
Local Llama 8B (simulated)
$0.00162 · 0% cached
Swift Budget (simulated)
$0.00318 · 18% cached
Atlas Frontier (simulated)
$0.5756 · 33% cached

Cost breakdown

Stacked by token class · latest run per model

Efficiency frontier

Capability score vs. cost per solved task (log) — frontier models outlined

Quality per dollar

capability composite ÷ total cost

Time to first token

mean across LLM spans

Cache hit rate

cached ÷ total input tokens

Cost table

ModelTotalCost / solvedQuality / $Input (unc / cached)OutputReasoningTTFTTPOTThroughputList $/Mtok (in / cached / out)
Local Llama 8B (simulated) pareto local$0.00162$0.0000528,71430,762 / 01,5770907ms45.8ms8 t/s$0.05 / $0.05 / $0.05
Swift Budget (simulated) pareto $0.00318$0.0000821,37215,370 / 3,3681,3370310ms9.0ms30 t/s$0.15 / $0.02 / $0.6
Atlas Frontier (simulated) pareto $0.5756$0.010716812,472 / 6,0531,36234,393665ms17.8ms50 t/s$3 / $0.3 / $15