A freeCodeCamp tutorial teaches how to build, pre-train, and fine-tune a 25-million parameter language model on a CPU, covering modern LLM architecture, pre-training from scratch, reinforcement learning post-training, and research methodology. The smaller-scale model enables rapid local iteration and experimentation without expensive cloud compute, using techniques like hybrid attention, Mixture of Experts, and multimodal inputs.
Seahaven is an open-source Python framework for building synthetic RL environments that isolate agent runs with reproducibility and state tracking. It handles parallel instances, change logging, and serves environments via OpenEnv, with features like fixtures for known starting states, concurrent serving, and a web console for inspection.
A blog post introducing reinforcement learning's Reinforce algorithm by building intuition through a deterministic approach to discrete random variables, avoiding pseudorandom sampling to clarify core concepts like variance reduction and expectation estimation.
Researchers at ETH Zurich have transformed a commercial robotic hand into a self-contained walking robot that crawls on its fingertips while manipulating objects. Using reinforcement learning and novel training techniques to account for the hand's asymmetrical finger structure, the team created an AI controller that enables the hand to traverse diverse surfaces and perform tasks like keyboard pressing and object manipulation. The approach could enable robots to reach confined spaces in maintenance and search-and-rescue operations.
AI labs are scaling agent testing infrastructure requiring massive CPU capacity for sandbox environments where models execute code and tools before release. AMD CPUs are positioned to benefit from this shift, with labs like OpenAI, Anthropic, and Google running tens of thousands of isolated agent instances simultaneously for safety evaluations and reinforcement learning.
ProVer is a framework that improves credit assignment in agentic reinforcement learning by using a model judge to identify pivotal decisions in trajectories, then verifying these segments through empirical outcome comparison rather than exhaustively evaluating every step. Tested on ALFWorld, WebShop, and SearchQA, ProVer achieves 9.91% and 7.12% improvements over GRPO for Qwen models at different scales.
The author experiments with TypeSafe's Jev, a decision-making AI model, to play the classic game Manic Miner. Jev shows mixed results compared to simpler algorithmic approaches, struggling to understand spatial relationships from text descriptions of game maps despite receiving ASCII representations and factual information about objectives.
This paper investigates language drift—unusual, non-standard language in LLM reasoning chains—that emerges during reinforcement learning with verifiable reward (RLVR) post-training. The authors prove theoretically that RLVR permits unbounded language drift while supervised fine-tuning does not, and show empirically that drift occurs on novel reasoning tasks. They demonstrate that constraining language drift necessarily harms performance, presenting a fundamental trade-off in frontier LLM post-training.
Researchers developed superhuman AI for Stratego, a board wargame with hidden information, using self-play reinforcement learning and test-time search. The achievement surpasses previous failed attempts and requires only thousands of dollars rather than millions, establishing new benchmarks for AI performance on classical games.
Context Language Models (CLMs) enable language models to natively manage their own context by treating it as an editable file, allowing models to learn what information to maintain. CLMs outperform existing context management strategies across multiple tasks with significant efficiency gains, and support both in-context and parametric learning of context strategies through natural-language steering and reinforcement learning.
Cloudflare released Clef and Clef-flash, open-source decision models hosted on Workers AI that classify inputs and produce bounded structured outputs for agentic workflows. The models outperform competitors like Jev on benchmarks, support both text and images, and offer 2x faster latency than LLMs for decision-making tasks. Cloudflare also introduced a reinforcement learning product to fine-tune Clef for specific use cases.