JevBench · head to head
typesafe-ai%2Fjev vs trustedrouter%2Fgev-1.0
Same eval, same problems, one draw per problem per model. The problems where exactly one of the two was right carry all of the evidence, and an exact paired test over just those has real power at sample sizes where comparing two rates is hopeless. Pick the same model twice to compare two providers serving the same weights.
No comparison
typesafe-ai%2Fjev has no scored runs on this eval.
A refusal here is the feature working: the page only renders a paired test when the pairs are real — same problem, same scaffold, same token limit, same judge, and a single provider on each side.