Skip to content
Routing intelligence

Every turn has a tradeoff.

One surface for cache behavior, route decisions, model quality, tokens, latency, spend, savings and benchmark evidence — so “auto” never means opaque.

Cache awareModel scoringCost attributionDecision traceBenchmark evidence
Window
Objective
Routed turns
14,286
2,041 / day
Cache hit
38.4%
+4.7 pts
Model score
92.1
quality floor held
Actual cost
$18.42
$1.29 / 1K turns
Saved
$31.67
63.2% vs baseline
P95 latency
1.18s
−22.8%

Turns over cost.

The line that matters: cumulative spend as work compounds. Compare the routed path against an explicit “always frontier” baseline.

Cumulative cost / turns

7 day sample
Arcana routedFrontier baseline
$50$35$20$5 03.5K7K10.5K14K turns

Why that model?

Expose the decision path instead of returning a model name without context. Cache, task fit, candidates, score and fallback policy remain inspectable.

Decision trace

sample
01
Cache lookupsemantic + exact request key
MISS
02
Task profilecoding · repo context · tool use
0.87 HARD
03
Candidate set4 healthy models satisfy policy
4 / 9
04
Score + selectquality floor first, then cost
OPUS 5 · 94.7
05
Fallback planGemini 3.7 Flash → GPT-5.6 Sol
ARMED
Candidate
Score
Quality
Task fit
Reliability
Latency
Est. cost
Claude Opus 5Anthropic · selected
94.7
96strong
98coding
99.1%7d
1.21sp50
$0.021turn
GPT-5.6 SolOpenAI · frontier
93.8
97strong
96coding
98.7%7d
1.34sp50
$0.027turn
Gemini 3.7 FlashGoogle · fast alternate
91.5
91good
94coding
99.4%7d
0.62sp50
$0.006turn
DeepSeek V4 ProDeepSeek · cost alternate
88.9
89good
90coding
98.9%7d
0.71sp50
$0.004turn

Cache is a first-class metric.

Separate hits, misses, writes and cached tokens. Savings should be attributable, not buried inside a blended cost number.

Cache effectiveness

all routes
38.4%hit rate
Hits / misses5,486 / 8,800
Cached tokens4.20M
Cache write tokens812K
Cost avoided$8.91

Tokens by model

input · output · cached · reasoning
Claude Opus 5
4.82M
Gemini 3.7
3.94M
GPT-5.6 Sol
2.71M
DeepSeek V4
1.86M
Qwen 3.8
1.22M
InputOutputCachedReasoning

Model economics, side by side.

Route share alone is not enough. Put score, tokens, cache, latency, actual spend and counterfactual savings on the same row.

Model
Score
Route share
Tokens
Cache
Cost
Saved
Claude Opus 5Anthropic
94.7
31.8%4,542 turns
4.82M33.1%
41.2%1.99M tok
$8.7447.4%
$9.0650.9%
Gemini 3.7 FlashGoogle
91.5
28.4%4,057 turns
3.94M27.1%
35.4%1.39M tok
$3.1216.9%
$8.7773.8%
GPT-5.6 SolOpenAI
93.8
18.7%2,671 turns
2.71M18.6%
37.9%1.03M tok
$4.6225.1%
$5.1152.5%
DeepSeek V4 ProDeepSeek
88.9
13.9%1,986 turns
1.86M12.8%
29.8%554K tok
$1.216.6%
$6.5284.3%
Qwen 3.8 MaxAlibaba
89.6
7.2%1,030 turns
1.22M8.4%
20.7%252K tok
$0.734.0%
$2.2175.2%

Benchmarks belong next to routing.

Quality and cost are a joint decision. Keep external evidence and production evidence visible beside the model score instead of hiding methodology elsewhere.

Evidence surface
Opus 5
GPT-5.6
Gemini 3.7
DeepSeek V4
Qwen 3.8
Coding indexnormalized evidence
96
95
91
90
89
Reasoning indexnormalized evidence
97
98
92
93
94
Agentic indextools + long horizon
95
94
91
88
87
Production successexample 7d telemetry
99.1%
98.7%
99.4%
98.9%
98.8%
Blended cost efficiencyquality-adjusted $/turn
72
68
96
94
88
Data status: benchmark and production values on this initial static page are illustrative UI fixtures, not verified Arcana benchmark claims. Replace them only from versioned benchmark artifacts or production telemetry.

The routing ledger.

A compact decision history for debugging, cost review and governance: what happened, why, how much it cost, and what the counterfactual would have been.

Time
Task
Cache
Model
Score
Cost
Saved
13:42:08
Refactor auth middleware
MISS
Opus 5
94.7
$0.021
$0.046
13:41:54
Explain test failure
HIT
cache
$0.000
$0.012
13:41:31
Generate migration query
MISS
Gemini 3.7
91.2
$0.004
$0.019
13:40:58
Review security boundary
MISS
GPT-5.6
95.1
$0.025
$0.031
13:40:12
Summarize build output
HIT
cache
$0.000
$0.006
13:39:44
Classify issue labels
MISS
DeepSeek V4
88.8
$0.002
$0.011

Make “auto” inspectable.

Connect real telemetry later without redesigning the page: the visual contract is already there.