A systems characterization study examines executing the ~375 GB Kimi K2.5 model on a single 128 GB AMD Ryzen AI MAX+ 395 PC using storage-backed expert caching. The bounded locality cache achieved 7.7% hit rate, reducing expert traffic by ~70 GB and decreasing package energy from 117.50 J to 113.47 J per generated token while maintaining ~0.438 tokens/s throughput.
Researchers developed accurate MATLAB-based models of AMD GPU matrix multipliers across three architectures (CDNA 1-3) to characterize their numerical behavior, since these units don't conform to IEEE 754 standards. The models were validated for bit-level reproducibility against hardware using 10 million test vectors and applied to compare accuracy differences between AMD and NVIDIA matrix cores.
Citi projects memory undersupply through 2031 driven by AI demand for HBM and DRAM, with HBM bit demand rising 62% in 2027 and 69% in 2028. Morgan Stanley estimates Nvidia will consume 37.3% of total HBM demand, followed by Google at 36% and AMD at 12.1%. Industry analysts debate whether a 2028 memory downcycle could occur, citing factors like HBM despec and new fab capacity.