An article on building personal AI benchmarks to evaluate models based on your actual work needs rather than standardized tests like MMLU-Pro. The author, now head of evaluations at Every, describes how testing models on real tasks—writing, dashboards, presentations—helps determine which model works best for specific jobs and whether cheaper alternatives suffice.
An independent first-person 3D adaptation of the classic text adventure Zork I, playable in a browser or as a downloadable Windows app. Players explore the white house and underground world, solve puzzles, and collect nineteen treasures while encountering enemies like trolls and thieves, with support for mouse, keyboard, and touch controls.