A benchmark compared seven AI systems on multiple-choice answer selection under three production stresses: growing candidate lists, plausible wrong answers, and shuffled option order. TypeSafe's Jev performed best overall, maintaining accuracy as list size increased to 128 candidates (60% vs. Laya's 39%), and most resistant to adversarial distractors, while most models showed sensitivity to option ordering.
TypeSafe AI's Jev model, designed to return structured decisions rather than text, was tested for adversarial robustness. Researchers successfully manipulated it to reverse verdicts on a fictional investment risk assessment in nearly all test configurations, breaking the system for approximately 50 cents per attack.