Livenerf is a deterministic benchmark launched on Claude Opus 5.5's release date (2026-09-22) to detect whether the model degrades after shipping. Running daily for 30 days across 78 screened questions from GPQA, MMLU-Pro, and competition math, it can detect accuracy drops of ~7.5 points per 10-day window using statistical methods from Anthropic's own evaluation framework.
Users report that OpenAI's GPT-6-Astra model produces inconsistent outputs on some accounts, with half of SVG pelican drawings coming out crude instead of polished, running slower, and resembling GPT-5.6-Luna. The degradation began abruptly on 2026-09-20 and may indicate per-account experiments, routing cohorts, or capacity management, though the underlying cause remains unclear from client-side data alone.