How ModelYield measures work
Task-level outcomes, explicit identity and honest cost.
Measurement protocol
Harness × model × provider × task
The router and adapter are additional versioned parts of the path. Claude Code Router is not itself the coding agent harness. An unknown upstream cannot support a provider-specific conclusion.
Success and retries
An attempt is one complete task run. Transport retries stay inside that attempt, with all time and costs. Three attempts with one success mean 1/3 success, not 100% pass@1.
Cost per success
Total cost of every evaluated attempt divided by successful attempts. Zero successes produces no finite cost per success. Missing cost produces unknown. Reconciled invoices, provider reports and estimates remain separate.
Evidence gates
At least 12 distinct tasks per category, three repetitions, fixed configuration and adequate coverage are planned minimums for real comparisons. Formal confidence intervals must resample by task, not treat repeated attempts as independent. Selection also requires paired comparisons and independent holdout validation.
This release
The example suite is synthetic. User imports are unverified observations. Neither produces a statistically supported recommendation. No exact model performance claims are made.
AgentFit and AgentCostLedger
Compatibility probes and request-level accounting are separate future capabilities. Their current pages show the data contract and untested states rather than pretend they are already running.
Existing coding-agent benchmark methodology