# HBM demand — X 热门讨论 (2026-09-26 09:41 UTC)

## @hsu_steve (steve hsu) · 09-25 20:14 · ♥35 ↻8 💬3 $15B revenue, ~$6B profits? Love engram

Grok: DeepSeek V4.1 Flash is sold at $0.30 / $0.006 / $1.20 per million tokens (input / cached input / output). SemiAnalysis models an owner running it on NVIDIA B200 + vLLM at P90 125 tok/s/user and 60% utilization, with no license fee.

At that point one utility gigawatt of B200s bills about $15.2B/year. Fully loaded compute TCO is about $8.9B ($1.73 per chip-hour), leaving $6.3B profit (~42% margin). The tweet’s “$15B profits” is the revenue bar, not profit.

Power is a small slice of that TCO. One GW running all year is 8.76 TWh; at $0.05–$0.08/kWh electricity is roughly $0.4–$0.7B. Most of the $8.9B is GPU, server, network, and building depreciation.

Engram DRAM offload moves the model’s large n-gram lookup tables off HBM onto host memory. That frees GPU memory for KV cache and can shrink tensor-parallel width, which SemiAnalysis says lifts revenue per GW by ~50% if demand and list prices hold. The four-bar chart is the baseline without that extra lift. > 引用 @SemiAnalysis_: MONEY PRINTER ALERT🚨 NVIDIA vLLM B200 CAN GENERATE UP TO💰️$15 BILLION💰️OF ANNUAL PROFITS PER GIGAWATT serving the open DeepSeekv4.1 Flash model at the official interactivity & official selling prices.

Using Engram DRAM offloading on NVIDIA results in a 50% increase in revenue per GigaWatt. https://x.com/hsu_steve/status/2103578949695467926