PyroWave is a GPU-accelerated intra-only video codec optimized for ultra-low-latency game streaming over local networks, achieving sub-0.1ms encode/decode times at 1080p using Vulkan compute shaders and wavelet transforms similar to JPEG2000.
Grant Sanderson frames intelligence as compression—the ability to distill information and connect concepts. The author explores parallels between how AI systems generalize rather than memorize and how humans who avoid rote memorization develop deeper understanding, suggesting that constraints on memorization can paradoxically enhance learning.
This article explains LLM inference optimization techniques for faster, cheaper production deployments. It covers the two-phase inference process (prefill and decode), memory management strategies like KV caching and PagedAttention, and methods such as model compression and speculative decoding to reduce cost and improve throughput.
Grant Sanderson's "compression is intelligence" framework suggests that true intelligence is the ability to distill and abstract information rather than memorize it. The author draws parallels between how AI systems must be prevented from overfitting to avoid memorization, and how humans who avoid rote memorization often develop stronger generalization and learning abilities. This perspective reframes the advantage of not studying as a feature that forces deeper conceptual understanding.
TriFlow is a generative method for creating compact 3D meshes with artist-like topology from signed distance fields. It represents mesh topology as a nearest-vertex vector field, trained using flow matching and optimized with topology-aware mesh simplification. The approach achieves 90% lower Chamfer Distance and 8× speedup compared to existing learning-based methods.
DeltaTensors is a storage optimization tool that compresses fine-tuned model checkpoints by storing only the weight differences from a base model instead of full copies, reducing storage by ~69% on tested models while maintaining minimal performance impact. It works with existing fine-tunes post-training, supports version history through chained deltas, and enables streaming operations without loading multiple full models into memory.
David Andersen explores how to compactly represent English grammar rules for article usage (a vs. an) using succinct data structures. Starting from a 1252-byte program and progressing through compression techniques, he demonstrates how a trie-based representation reduces the data to 200 bytes, then discusses information-theoretic approaches like LOUDS encoding for optimal compression.