Platform

Replay strategies over traffic you already ran

Compare routing strategies against the same graded cases, including baselines designed to beat you. If a simple baseline wins, that is the finding.

Baselines that can win

The harness ships always-premium, always-balanced, always-economy and a task-tier heuristic. Always-balanced is the strongest simple baseline and is reported alongside every comparison — omitting it is how routing products flatter themselves.

Sweep the floor, not just the average

With no quality floor a cost-minimising router correctly reduces to 'cheapest', and any comparison is vacuous. Results are reported across a range of floors so the trade is visible rather than cherry-picked at its best point.

Repeat before believing

A single run makes a one-case difference look like a three-point gap. Strategy comparisons pool repeat runs, because a benchmark you ran once is an anecdote with a table around it.

Measurements are still sparse

One to two trials per endpoint and task type means the posterior stays prior-dominated, so a cheap model cannot yet earn a high enough bound to be selected for easy work even where it is in fact reliable. More trials, and a more heterogeneous pool, would change the verdict.

Elsewhere in the platform