A detailed benchmark compared Claude and ChatGPT's performance on data analysis tasks using live Zendesk support ticket data. Both models were tested identically across seven stages from discovery to self-audit, with results evaluated by a third Claude instance for factual accuracy against actual returned metrics.