Science or Slop? is a tool that detects scientific reasoning failures in AI-generated papers by analyzing structural and logical patterns across sections, rather than surface-level prose. Users can upload papers to receive a private report measuring six indicators of scientific slop across three analytical planes, with results ranked on a 0-100 index.
ArXiv announced changes to its rate limiting policy in response to exponential growth in submissions, restricting each submitter to a maximum of two submissions per month.
arXiv implemented a new rate-limiting policy effective October 1st, 2026, restricting submitters to two papers per month and three active submissions simultaneously. The policy responds to a doubling of submissions in two years, driven largely by AI tools enabling authors to flood the repository with low-quality papers, overwhelming volunteer moderators and delaying legitimate research.
arXiv announced submission caps of two papers per month and three active submissions to manage exponential growth in submissions, which have nearly doubled in two years and overwhelmed volunteer moderators. The policy aims to maintain quality and ensure the platform's infrastructure can sustain human review processes amid AI-accelerated research productivity.
Context Language Models (CLMs) enable language models to natively manage their own context by treating it as an editable file, allowing models to learn what information to maintain. CLMs outperform existing context management strategies across multiple tasks with significant efficiency gains, and support both in-context and parametric learning of context strategies through natural-language steering and reinforcement learning.
Researchers present a framework for bounded loops in agent systems that guarantees termination, prevents drift, and enforces spending limits through independent gates and repair relations. They introduce a measurement instrument to detect vacuous and self-attesting gates in agent harnesses, finding 47 problematic gates in production code across a 69-loop catalog.
This paper introduces agentic meta-reasoning, an inference-time framework where a controller makes explicit decisions about task execution while workers perform computations. Tested on multiple benchmarks including program reconstruction, the approach achieves significant improvements over direct control baselines, with gains of 3.6–4.2 points across frontier models and better performance at larger compute budgets.
Context Language Models (CLMs) enable language models to manage their own context by treating it as an editable file, allowing them to learn what information to maintain. This approach outperforms existing context management strategies across multiple benchmarks with significant efficiency gains, and supports multi-agent systems and reinforcement learning for improved performance.
RRSI is a method for evolving AI agent harnesses that avoids overfitting to training benchmarks through regularized search constraints. The approach improves performance on held-out benchmarks by controlling edit magnitude, requiring measured gains to exceed variance, and eliminating benchmark-specific logic, with all candidates logged for transparency.
JEV-as-a-Judge proposes a decision-only LLM judge that returns confidence scores instead of reasoning, accepting verdicts when confident and escalating uncertain cases to reasoning judges. The approach matches GPT-4's accuracy at 41% of the cost with 0.15-second latency but struggles on math, code, and logic tasks where reasoning is essential.
A research paper argues that incremental increases in AI capabilities pose existential risks through gradual human disempowerment, as AI systems become competitive alternatives across labor, governance, and culture. Without human participation being necessary for societal functions, institutions will lack incentives to ensure human flourishing, while economic and political feedback loops make coordinated human resistance increasingly difficult.
WorldCrafter is a video world model that uses implicit 3D-aware memory to maintain consistent scenes during camera-controlled exploration. Starting from a single image or text prompt, it generates minute-scale videos with coherent geometry and revisit consistency, supporting real-time interaction through distillation across static and dynamic environments.
A researcher created an interactive terrain visualization mapping 3,466 AI safety papers from arXiv and LessWrong, organized by citation count and semantic clustering across 18 sub-fields. The visualization reveals that alignment training and scalable oversight dominate citations through recent high-impact papers like InstructGPT and Constitutional AI, while adversarial robustness and interpretability have more total works but lower individual citation counts.
A research paper demonstrates that language model capabilities can transfer to unrelated tasks through post-training artifacts. The release includes reproducible code and frozen training data for experiments using the Qwen2.5-1.5B model on HumanEval+ benchmarks, with detailed instructions for verification and replication across CPU and GPU environments.
ParkourNote is a research workspace built solo in one month using Claude. It allows users to import papers, books, videos, and files into an Inbox where an AI suggests project organization, with features for deduplication and cross-project linking.
A 2018 paper introduces functional decision theory (FDT), a novel approach to instrumental rationality that treats decisions as outputs of mathematical functions optimizing for best outcomes. FDT outperforms causal decision theory and evidential decision theory across classic decision problems including Newcomb's problem, the smoking lesion problem, and Parfit's hitchhiker.