Agent routing intelligence
The task is the benchmark.
Inspect the path, the failures and the full cost of getting work done.
Published real evaluations
0No paid benchmarks executedExample routing paths
3Synthetic, clearly separatedTask categories
3Bugfix · Feature · RefactorValidated recommendations
0Evidence gates not metExplore the example suite
SYNTHETIC SUITE| Execution path | Success / attempts | Cost / success | Evidence |
|---|---|---|---|
| Balanced path Example model A Example endpoint A | 33/36 (91.7%) 12 example tasks × 3 | $0.276Includes failed attempts | Example only Inspect → |
| Economy path Example model B Example endpoint B | 29/36 (80.6%) 12 example tasks × 3 | $0.107Includes failed attempts | Example only Inspect → |
| Fast path Example model C Unknown upstream | 25/36 (69.4%) 12 example tasks × 3 | UnknownMissing usage | Example only Inspect → |
Metrics here are arithmetic on synthetic records. No inferential confidence interval or real model ranking is implied.
A complete cost picture
Count the unsuccessful work, too.
A low token price does not tell you what a completed task costs. Include failed attempts, transport retries and paid tool calls.
Try the success cost calculator →ILLUSTRATIVE CALCULATION
$2.00 ÷ 8 = $0.25
10 attempts · 8 successes · all attempt costs included
What makes a result usable?
01
Exact execution path
Harness, router, model and actual provider.
02
Task-level outcomes
Failures and retries stay in the denominator.
03
Traceable cost
Pricing date, usage and billing basis.
04
Read the protocol →Enough evidence
Independent tasks, uncertainty and fresh versions.
Your imported results stay private. Importing a file does not turn it into a verified public benchmark.