A discussion on X critiques HBM capacity as an overblown bottleneck in AI inference, arguing that recent optimizations like V4.1 Flash have reduced memory requirements per token from 3.5KB to 890 bytes through techniques such as CSA2 and SWA Bounded Replay. The author contends that bandwidth, not capacity, is the true limiting factor in decode operations, and that moving from 12-Hi to 8-Hi HBM stacks reflects this reality rather than supply constraints.
A user discusses how DeepSeek V4.1-Flash's reduction in KV Cache usage validates their thesis that memory optimization, not raw compute, will be the key competitive advantage for AI agents. They argue that as AI becomes cheaper, demand increases and total inference grows, making efficient memory hierarchies and infrastructure the next frontier for investment.