OpenAI's Decisions API, powered by Luna (a version of GPT-6), struggles with confidence calibration and multi-step reasoning compared to competitor Jev. In testing on logical inference problems, Luna achieved only 68% accuracy when claiming 99% confidence, and performance dropped to coin-flip levels (45-46%) on five-step reasoning tasks versus Jev's 81-89%.