Research lab / Routing evidence
Real recommendations: 0
Agent routing intelligence

The task is the benchmark.

Inspect the path, the failures and the full cost of getting work done.

Import results
Published real evaluations
0No paid benchmarks executed
Example routing paths
3Synthetic, clearly separated
Task categories
3Bugfix · Feature · Refactor
Validated recommendations
0Evidence gates not met
Example data. All example models, endpoints and results are synthetic. They demonstrate the workflow and are not recommendations or measured benchmarks.

Explore the example suite

SYNTHETIC SUITE
Execution pathSuccess / attemptsCost / successEvidence
Balanced path Example model A
Example endpoint A
33/36 (91.7%)
12 example tasks × 3
$0.276Includes failed attemptsExample only
Inspect →
Economy path Example model B
Example endpoint B
29/36 (80.6%)
12 example tasks × 3
$0.107Includes failed attemptsExample only
Inspect →
Fast path Example model C
Unknown upstream
25/36 (69.4%)
12 example tasks × 3
UnknownMissing usageExample only
Inspect →
Metrics here are arithmetic on synthetic records. No inferential confidence interval or real model ranking is implied.

A complete cost picture

Count the unsuccessful work, too.

A low token price does not tell you what a completed task costs. Include failed attempts, transport retries and paid tool calls.

Try the success cost calculator →

ILLUSTRATIVE CALCULATION

$2.00 ÷ 8 = $0.25

10 attempts · 8 successes · all attempt costs included

What makes a result usable?

01

Exact execution path

Harness, router, model and actual provider.

02

Task-level outcomes

Failures and retries stay in the denominator.

03

Traceable cost

Pricing date, usage and billing basis.

04

Enough evidence

Independent tasks, uncertainty and fresh versions.

Read the protocol →
Your imported results stay private. Importing a file does not turn it into a verified public benchmark.