Economics
Cost, efficiency & latency
Cross-cutting pillar: every span records cached / uncached / output / reasoning tokens and is priced at write time (pricing table 2026.09.04). Local models use USER_SET_PRICE_PER_MTOK.
Efficiency space
x = cost per solved task (log, normalised) · y = capability score · z = TTFT · bubble = Omni · glowing = Pareto frontier
Rendering 3D view…
Where the money goes
Pooled spend by token class
$0.5803
total
Local Llama 8B (simulated)
$0.00162 · 0% cached
Swift Budget (simulated)
$0.00318 · 18% cached
Atlas Frontier (simulated)
$0.5756 · 33% cached
Cost breakdown
Stacked by token class · latest run per model
Efficiency frontier
Capability score vs. cost per solved task (log) — frontier models outlined
Quality per dollar
capability composite ÷ total cost
Time to first token
mean across LLM spans
Cache hit rate
cached ÷ total input tokens
Cost table
| Model | Total | Cost / solved | Quality / $ | Input (unc / cached) | Output | Reasoning | TTFT | TPOT | Throughput | List $/Mtok (in / cached / out) |
|---|---|---|---|---|---|---|---|---|---|---|
| Local Llama 8B (simulated) pareto local | $0.00162 | $0.00005 | 28,714 | 30,762 / 0 | 1,577 | 0 | 907ms | 45.8ms | 8 t/s | $0.05 / $0.05 / $0.05 |
| Swift Budget (simulated) pareto | $0.00318 | $0.00008 | 21,372 | 15,370 / 3,368 | 1,337 | 0 | 310ms | 9.0ms | 30 t/s | $0.15 / $0.02 / $0.6 |
| Atlas Frontier (simulated) pareto | $0.5756 | $0.0107 | 168 | 12,472 / 6,053 | 1,362 | 34,393 | 665ms | 17.8ms | 50 t/s | $3 / $0.3 / $15 |