Building frontier AI models is no longer just a research endeavor. It now requires a complex, industrial-scale R&D system. At leading AI companies, these systems rely on large teams of experts. Architecture, infrastructure, training, inference, deployment, and evaluation are only part of it. Behind every frontier AI model is a continuous cycle of experimentation, iteration, and refinement.

At NaiveAI, we broke with human-centered R&D from day one. We put AI models to work on AI R&D itself. They write code, run experiments, monitor progress, analyze results, and iterate. Human researchers set direction, define constraints and criteria, and make critical decisions. The human edge lies in experience, insight, and judgment.

NaiveAI built production-scale infrastructure for AI-centered R&D, giving AI models secure work environments and access to GPU compute. The system serves close to ten million sandboxes each week, with 100,000 active concurrently at peak. A unified control plane manages compute, environments, tools, permissions, and security at scale.

Naive-N0.5-Flash was built this way. AI explored and designed its hybrid attention architecture while optimizing its training, inference, and deployment systems. Human researchers provided guidance and made key decisions. Naive-N0.5-Flash is also trained for AI R&D, allowing it to participate directly in the R&D process and opening a path toward recursive self-improvement (RSI).

Model Overview

Naive-N0.5-Flash is a 309B MoE model with 15.5B active parameters, built for coding and AI R&D.

- Native 1M context, without full attention. Naive-N0.5-Flash combines Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA) with GQA4 at a predominantly 5:1 SWA–DSA layout. The entire network remains local or sparse, with no full-attention layers.

- AI-optimized inference up to 2,000 tokens/s. NaiveRT, our inference system for Naive-N0.5-Flash, was built and optimized through AI-centered R&D. It combines mega-kernel fusion, Programmatic Dependent Launch (PDL), and speculative decoding, delivering 50 tokens/s per user in Standard mode and up to 2,000 tokens/s in Ultrafast mode.

- Open weights and API. Model weights and inference code are released under the MIT license. API access will also be provided, with pricing set at $0.10 / $0.40 / $0.01 per million tokens for input, output, and cache reads, respectively.

Evaluation Results

Evaluation setup. Unless otherwise noted, our evaluations of Naive-N0.5-Flash use Claude Code 2.1.207 with a 1M-token context window, temperature 1.0, and top-p 0.95. The harness exposes only basic file I/O and Bash tools.

Sources for reported benchmark scores

- GLM-5.3 and GLM-5.3-Flash: GLM 5.3 blog and GLM 5.3 Flash blog, respectively.

- Kimi-K3: Kimi-K3 model page.

- Qwen-3.8-Max: Qwen 3.8 Max blog.

- Hy4-preview: Hy4-preview model page.

- DeepSeek-V4.1-Flash: DeepSeek-V4.1-Flash model page.

- Step-5-preview: Step-5-Preview-BF16 model page.

- Fable-5 (w/ fallback): GLM 5.3 blog.

- SWE-Bench Pro: GPT-5.6-Sol, Opus-5, and Opus-5.5 scores are drawn from the GPT-5.6 blog and the Claude Opus 5.5 System Card.

- DeepSWE v1.1: The Muse-Spark-1.3 score comes from its Muse Spark 1.3 blog. GPT-5.6-Sol and Opus-5 scores come from the DeepSWE v1.1 leaderboard. The Opus-5.5 score comes from the Claude Opus 5.5 System Card.

- Terminal-Bench 2.1: Muse-Spark-1.3, GPT-5.6-Sol, and Opus-5 scores come from the Muse Spark 1.3 blog. The GPT-6-Astra score comes from the Terminal-Bench 2.1 leaderboard.

- ALE-CLI: GPT-5.6-Sol, GPT-6-Astra, Muse-Spark-1.3, Opus-5, and Opus-5.5 scores come from the ALE-CLI leaderboard.

- FrontierSWE v1: We calculate the Dominance score using the competing systems’ results as of August 23, 2026.

- ProgramBench: We report the Almost@1 score. GPT-5.6-Sol and Opus-5 scores come from the ProgramBench leaderboard.

- MLE-bench-30: Gemini-3.5-Flash, Gemini-3.6-Flash, Grok-4.5, and GPT-5.6-Luna scores come from the Gemini 3.6 Flash model card. We follow the Gemini 3.6 Flash model card and report the average position score of Naive-N0.5-Flash.

- PaperBench: MiniMax M3, Opus-4.7, GPT-5.5, and Gemini-3.1-Pro scores come from the MiniMax M3 model page.

- Sol-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun: Naive-N0.5-Flash scores were obtained using our inhouse AutoResearch harness. Recursive Superintelligence Inc. scores come from its research article. Note that, following the Sol-ExecBench update, we use Recursive Superintelligence Inc.’s updated score on the leaderboard.

AI-Centered R&D in Practice

Naive-N0.5-Flash helps researchers with open-ended work across AI research and systems engineering.

Technical Details

Model Architecture

Naive-N0.5-Flash builds on the open-weight MiMo-V2.5 base model, which has a simple architecture with strong capabilities in world knowledge and deep research. Most layers use Sliding-Window Attention (SWA), whose per-token decoding cost does not grow with context length, while a small number of global-attention layers preserve long-range information. At million-token context lengths, however, these global-attention layers account for much of the decoding overhead.

Naive-N0.5-Flash replaces the global-attention layers with DeepSeek Sparse Attention (DSA). A lightweight indexer scores the full history, while the backbone computes attention only over a selected subset of tokens. Although the indexer still scans the full history and the full KV cache is retained, sparse attention substantially reduces attention computation and memory access. Adapting the model to this new attention structure was one objective of continued pretraining.

NaiveAI researchers and AI models jointly explored the architecture. Researchers defined the objective and evaluation protocol: improve decoding efficiency at million-token context lengths while preserving model quality. AI models implemented candidate architectures and training approaches, ran ablations, and summarized the results. Among the candidates that met these objectives, researchers selected one of the simplest to implement.

Hybrid SWA–DSA Architecture

The network consists of eight six-layer modules. A standard module contains five SWA layers followed by one DSA layer, with the first layer of the first module also replaced by DSA. SWA uses a 128-token window, while DSA selects the top 2,048 tokens for backbone attention. Both attention types incorporate sink bias.

Unlike the original MLA-based DSA implementation, Naive-N0.5-Flash replaces MLA with grouped-query attention (GQA) using four KV groups. Through infrastructure–algorithm co-design, we developed a lightweight indexer with 16 query heads, reducing index-selection wall time by 44% relative to the original DSA implementation while preserving performance on agent tasks at 1M context.

Training Procedure

Following the architectural changes, Naive-N0.5-Flash completed 3.25T tokens of multi-stage training with a native 1M-token context window: 50B tokens of Indexer Warmup, 3T tokens of Sparse Attention Training, and 200B tokens of Learning Rate Decay. This process adapted the model to its new sparse attention architecture while substantially improving its AI R&D and coding capabilities.

- Indexer Warmup — 50B tokens. Only the newly introduced DSA indexer is trained, while all other model parameters remain frozen. The layers being converted to DSA retain full attention for forward computation at 1M context, while the SWA layers remain unchanged. The full-attention distributions in those layers provide the supervision signal, with a KL-divergence loss aligning the indexer and backbone attention distributions and establishing the indexing behavior required for the subsequent transition to sparse attention.

- Sparse Attention Training — 3T tokens. After warmup, the model switches to sparse attention and enters continued pretraining (CPT) with a language modeling (LM) loss at a fixed learning rate, focusing primarily on AI R&D and coding. The model remains at 1M context throughout. Continued pretraining adapts the model to the new information-selection and aggregation mechanisms while strengthening long-context modeling for tasks that require retaining task history, connecting dispersed information, and sustaining extended interactions.

- Decay Stage — 200B tokens. The model then enters supervised fine-tuning (SFT), remaining at 1M context with the sparse-attention execution path and LM objective, while the learning rate is gradually reduced over the final 200B tokens.

Training System

Full-to-Sparse Transition

During indexer warmup, the layers being converted to DSA retained full attention, while the SWA layers remained unchanged. The indexer was aligned to the full-attention reference through a KL-divergence objective before the transition to sparse attention.

To validate the transition to sparse attention, AI models automatically analyzed top-k recall for the indexer’s top-2,048 selections against the full-attention reference. They identified and fixed numerical-stability issues in indexer top-k selection and independently verified the fixes. Based on these results, researchers determined the transition point and accuracy thresholds, and aligned validation standards across training and deployment.

Hybrid Sequence Parallelism

While analyzing attention computation at 1M context, AI models identified a key parallelism property of the hybrid architecture: DSA requires access to the full history, whereas SWA attends only to a 128-token window to the left. They therefore proposed and validated a hybrid sequence-parallel scheme using different execution paths for the two attention types.

DSA layers use Ulysses Sequence Parallelism, with all-to-all redistribution providing access to the full sequence. SWA layers use SWA Halo Sequence Parallel, where overlapping shards exchange only the required ghost region with their left-hand neighbor, with the overlap determined by sample boundaries. This preserves DSA’s global modeling capability while reducing SWA communication complexity from O(L) to O(w).

The same principle extends to inference: long-context prefill, row-sharded computation, and index retrieval use their respective context-parallel schemes rather than sharing a single cross-GPU communication pattern.

GPU Memory Optimization

AI models also analyzed and optimized GPU memory use. They extended the training framework’s activation-memory management to handle intermediate activations from sparse indexing and attention alignment. By analyzing recomputation paths, they refined offloading granularity down to individual operator outputs, with independently configurable offloading for each intermediate activation. Top-k indices are explicitly retained to prevent recomputation from changing the selected positions.

At 1M sequence lengths, correctness checks and performance profiling also uncovered precision issues in positional encoding and potential out-of-bounds sequence indexing.

Across numerical stability, hybrid sequence parallelism, tensor-layout transformations, long-sequence support, and memory optimization, AI models participated directly in the engineering loop of analysis, discovery, repair, and verification. Researchers defined the objectives, constraints, and decision boundaries, while AI models carried out automated analysis, problem discovery, and independent verification — a concrete example of using models to build models.

With a 1M-token context configuration, the resulting training system can process 1T training tokens in approximately four days on 512 GPUs.

License

Naive-N0.5-Flash model weights and inference code are released under the MIT License.

Citation

If you find Naive-N0.5-Flash useful in your research or work, please cite:

@misc{naiveai2026naiven05flash,

title = {Naive-N0.5-Flash: Building Frontier AI with AI},

author = {{NaiveAI Team}},

year = {2026},

}Acknowledgments

Naive-N0.5-Flash builds on the work of the open-source community and gives back to it. We thank the Xiaomi MiMo team for making their MiMo-V2.5 base model publicly available, the DeepSeek team for their work on DeepSeek Sparse Attention (DSA), and the SGLang team and community for their open-source inference infrastructure.

Contact

For questions, feedback, or collaboration, please contact us at contact@naive.ai or follow us on X at @naiveailab. You can also find our open-source projects and model releases on GitHub and Hugging Face.