Artificial Analysis released Capability Indices v1.1, updating domain-specific AI model evaluations across finance, legal, healthcare, engineering, and other sectors. The update incorporates stronger evaluations from Intelligence Index v4.3, adds agentic tool use benchmarks, and removes customer interaction metrics across most domains.
A research project re-evaluates frontier AI models' physics capabilities by auditing benchmark questions, finding that low leaderboard scores may not reflect true model limitations. The study, based on arXiv:2609.13009, suggests existing physics benchmarks may be broken and that AI performance on physics problems requires deeper analysis beyond raw scores.