source&pool
A daily wire of long-form journalism, video, and discourse — filed, tagged, and laid out flat.
VOL. I·NO. 01
THURSDAY, OCTOBER 8, 2026
  1. 001Hacker NewsOCT · 08English

    Build and Train a 25M Parameter LLM from Scratch on Your CPU

    A freeCodeCamp tutorial teaches how to build, pre-train, and fine-tune a 25-million parameter language model on a CPU, covering modern LLM architecture, pre-training from scratch, reinforcement learning post-training, and research methodology. The smaller-scale model enables rapid local iteration and experimentation without expensive cloud compute, using techniques like hybrid attention, Mixture of Experts, and multimodal inputs.

    By Beau Carnes
  2. 002Hacker NewsOCT · 08English

    Show HN: Seahaven – Open-source framework for building RL environments

    Seahaven is an open-source Python framework for building synthetic RL environments that isolate agent runs with reproducibility and state tracking. It handles parallel instances, change logging, and serves environments via OpenEnv, with features like fixtures for known starting states, concurrent serving, and a web console for inspection.

    By scosman
  3. 003Hacker NewsOCT · 08English

    From Sampling to Reinforce

    A blog post introducing reinforcement learning's Reinforce algorithm by building intuition through a deterministic approach to discrete random variables, avoiding pseudorandom sampling to clarify core concepts like variance reduction and expectation estimation.

    By jxmorris12
  4. 004Hacker NewsOCT · 08English

    Walking Robotic Hand Uses AI to Learn New Tricks

    Researchers at ETH Zurich have transformed a commercial robotic hand into a self-contained walking robot that crawls on its fingertips while manipulating objects. Using reinforcement learning and novel training techniques to account for the hand's asymmetrical finger structure, the team created an AI controller that enables the hand to traverse diverse surfaces and perform tasks like keyboard pressing and object manipulation. The approach could enable robots to reach confined spaces in maintenance and search-and-rescue operations.

    By Edd Gent
  5. 005X 主题热门OCT · 02English

    HBM demand · X 热门 · 2026-10-02 15:53 UTC

    AI labs are scaling agent testing infrastructure requiring massive CPU capacity for sandbox environments where models execute code and tools before release. AMD CPUs are positioned to benefit from this shift, with labs like OpenAI, Anthropic, and Google running tens of thousands of isolated agent instances simultaneously for safety evaluations and reinforcement learning.

  6. 006Hacker NewsOCT · 02English

    Fixing GRPO's credit assignment problem without evaluating every step

    ProVer is a framework that improves credit assignment in agentic reinforcement learning by using a model judge to identify pivotal decisions in trajectories, then verifying these segments through empirical outcome comparison rather than exhaustively evaluating every step. Tested on ALFWorld, WebShop, and SearchQA, ProVer achieves 9.91% and 7.12% improvements over GRPO for Qwen models at different scales.

    By Jung; Dongwon; Ramesh; Hemanth Neelgund; Wang; Yifan; Li; Xiaomin; Hao; Yuexing; Hu; Chen; Muhao; Chandrasekaran; Varun; Banburski-Fahey; Andrzej; Lanier; Jaron
  7. 007Hacker NewsOCT · 02English

    Jev Plays Manic Miner

    The author experiments with TypeSafe's Jev, a decision-making AI model, to play the classic game Manic Miner. Jev shows mixed results compared to simpler algorithmic approaches, struggling to understand spatial relationships from text descriptions of game maps despite receiving ASCII representations and factual information about objectives.

    By Chris Greening
  8. 008Hacker NewsOCT · 02English

    Language Drift During RLVR Post-Training

    This paper investigates language drift—unusual, non-standard language in LLM reasoning chains—that emerges during reinforcement learning with verifiable reward (RLVR) post-training. The authors prove theoretically that RLVR permits unbounded language drift while supervised fine-tuning does not, and show empirically that drift occurs on novel reasoning tasks. They demonstrate that constraining language drift necessarily harms performance, presenting a fundamental trade-off in frontier LLM post-training.

    By Sullivan; Michael; Koller; Alexander
  9. 009Hacker NewsOCT · 02English

    Superhuman AI for Stratego

    Researchers developed superhuman AI for Stratego, a board wargame with hidden information, using self-play reinforcement learning and test-time search. The achievement surpasses previous failed attempts and requires only thousands of dollars rather than millions, establishing new benchmarks for AI performance on classical games.

    By Sokota; Samuel; Vinitsky; Eugene; Hu; Hengyuan; Kolter; J Zico; Farina; Gabriele
  10. 010Hacker NewsOCT · 01English

    Context Language Models

    Context Language Models (CLMs) enable language models to natively manage their own context by treating it as an editable file, allowing models to learn what information to maintain. CLMs outperform existing context management strategies across multiple tasks with significant efficiency gains, and support both in-context and parametric learning of context strategies through natural-language steering and reinforcement learning.

    By Shao; Rulin; Shen; Shannon Zejiang; Yin; Junjie Oscar; Li; Yuetai; Wang; Minheng; Ivison; Hamish; Poovendran; Radha; Lambert; Nathan; Xiao; Teng; Lewis; Mike; Yih; Wen-tau; Zettlemoyer; Luke; Koh; Pang Wei
  11. 011Hacker NewsOCT · 01English

    Clef: our open-source decision models

    Cloudflare released Clef and Clef-flash, open-source decision models hosted on Workers AI that classify inputs and produce bounded structured outputs for agentic workflows. The models outperform competitors like Jev on benchmarks, support both text and images, and offer 2x faster latency than LLMs for decision-making tasks. Cloudflare also introduced a reinforcement learning product to fine-tune Clef for specific use cases.

    By jasondavies