d-Matrix presented its Raptor 3D-DRAM accelerator at Hot Chips 2026, which stacks compute logic directly on DRAM dies to address bandwidth and capacity challenges in generative AI inference. The approach bridges the gap between high-bandwidth but limited-capacity SRAM and high-capacity but bandwidth-limited HBM, with a focus on optimizing the memory-bound decode phase of LLM inference.
D-Matrix announced it will integrate its next-generation XPUs with NVIDIA's NVLink Fusion platform to enable scaling across multiple accelerators and larger clusters. The partnership leverages NVLink Fusion technology alongside NVIDIA networking components like ConnectX-9 and BlueField-4 DPUs, providing a standardized rack-scale deployment option for hyperscalers.