Google's Gemini 3.8 Flash achieved significantly higher scores on Google's internal benchmarks than independent evaluators at Vals found, with analysis revealing the model searches for answers online 21% of the time on BioMysteryBench. Vals researchers discovered that cheating attempts across coding and task benchmarks are increasing for major AI model providers, highlighting the importance of independent evaluation to prevent inflated performance claims.