Platform

Quality you measured, not a tier you were told

Vendor tiers are a claim. Planverity treats them as a prior with deliberately wide uncertainty, then lets graded trials move the estimate — including against the vendor's own claim.

32
benchmark cases across three difficulty bands
7
trials per cell, pooled
$0.000350
cost per successful task — best of any strategy

Lower bounds, not point estimates

Routing compares a lower-confidence bound on quality. Wide uncertainty pushes that bound down, so an unproven route cannot satisfy a high floor on optimism alone — the router escalates or refuses instead of routing cheap on unearned confidence.

Evidence can contradict the vendor

Measured outcomes update a Beta-Binomial posterior seeded by the tier prior. A hard 'use the raw success rate once n is large enough' cutoff makes the estimate lurch, and twenty successes in a row would otherwise assert perfection. Evidence must be able to move the estimate against the tier, or the tier is an unfalsifiable claim.

Three exclusions that keep noise out

Mock runs are excluded — a mock echoes the prompt, so its grades describe the harness. Infrastructure errors are excluded — a rate limit is not a quality failure, and counting it as one makes an unavailable model look incompetent. And no model grades another: these labels calibrate the quality model, so grading with a model would make it circular.

It does not dominate a good single model

Against always-Sonnet, Planverity offers a dial — 95.1% success at half the cost, or equal quality at about 5% more — not a free lunch. It beats always-premium and a hand-rolled task-tier heuristic, and that heuristic is itself dominated by simply using one balanced model. When one model is near Pareto-optimal alone, routing has little room to add value.

Elsewhere in the platform