A September 2026 research report tested 18 popular AI models including ChatGPT, Gemini, Claude, and CoPilot on 121 financial questions across pensions, tax, debt, and savings. The models made mistakes 57% of the time on average, raising concerns about AI-delivered financial advice and whether regulatory guardrails on large language models could better protect consumers.