An article on building personal AI benchmarks to evaluate models based on your actual work needs rather than standardized tests like MMLU-Pro. The author, now head of evaluations at Every, describes how testing models on real tasks—writing, dashboards, presentations—helps determine which model works best for specific jobs and whether cheaper alternatives suffice.
Good Start Labs, a startup spun out of Every with $3.6 million in funding, uses games like Diplomacy and 1830 as training environments to teach AI models strategic reasoning and real-world skills. The company found that training a 30B model on the railroad strategy game 1830 improved its performance on financial research tasks, demonstrating that gaming mechanics can transfer to practical work applications.