# HBM demand — X 热门讨论 (2026-09-28 09:51 UTC)

## @FundaAI (FUNDA) · 09-28 03:09 · ♥34 ↻6 💬4 Happy to be part of the Peer Review Group! @FredaDuan’s framework makes sense: separate the sandbox from inference, and size physical resources around concurrent usage.

But the most interesting part, as we discussed in our earlier tweet on Muse’s ROI, is how the workload itself could evolve. Based on today’s workload, we estimate Muse costs $20 per DAU per month, and expect that to fall to around $10 soon. But today’s workload could also change quickly as users delegate more work.

The good news is that if personal agents have that much work to do, they’re likely delivering a lot of value.

Today, most tasks are fairly self-contained: book a flight, negotiate a bill. Over time, an agent could manage your inbox, follow a project for weeks, plan a family trip, or keep track of several purchasing and investment decisions. Each task could involve more websites, more files, more tool calls, and several subagents working in parallel. A user might check in a few times a day while the agent continues working in the background.

That matters for CPU demand. User count is only one input; how much work each user delegates matters just as much. Heavy users could have several agents running in parallel throughout the day, accounting for a disproportionate share of CPU usage. As tasks become more complex and agents run more frequently and for longer, we think CPU demand for heavy users could reach 3–5x Freda’s baseline over the medium term, while the increase for the average user could be smaller. That range assumes a more mature product and deeper usage, and will depend in part on how many users become heavy users. It will need to be tested against actual usage data.

The same applies to DRAM. A VM using 3GB today may need more as it handles more browser tabs, larger files and concurrent tasks. Longer-running tasks also leave more state to retain. The increase will depend on architecture and resource sharing, but we wouldn’t hold memory per user constant as the product evolves.

We’re particularly bullish on NAND. A personal agent used over months or years accumulates files, conversation history, preferences and unfinished work. That data stays when the user logs off. When they return, the agent needs to pick up where it left off.

CPU capacity can be reassigned once a task finishes. The user’s stored data still occupies space. Someone who has used an agent for two years could have a very different storage footprint from someone who signed up two weeks ago. Frequently retrieved history, personal knowledge indexes and working files are all natural uses for NAND. You can’t apply a two-hour daily activity assumption to that storage requirement.

Jensen’s Inference Context Memory Storage announcement at CES earlier this year addressed a related need. It adds a flash tier for reusable KV cache, extending capacity beyond GPU and host memory so agents can reuse context across interactions. Persistent personal data and reusable inference context are distinct workloads, but both add demand for NAND. We think both become more significant as people entrust agents with more of their daily work.

There is also memory and storage demand beyond the sandbox. Freda’s VM DRAM estimate covers one part of the system. Training and inference clusters need GPU HBM, host DRAM, and local and shared SSD storage.

For a sense of scale, a DGX B200 system has eight GPUs with 1.44TB of HBM, another 2–4TB of host memory, and roughly 30.7TB of local data SSDs, before external storage. Meta’s configuration will differ, but the point is that GPU deployments bring substantial memory and storage requirements of their own. Counting VM memory alone misses that demand.

On GPUs, Freda’s estimate of roughly 1GW for 100 million DAU is a useful starting point. Cache reuse, model improvements and custom silicon should keep bringing down the cost of a given task. At the same time, users may ask for more tasks, longer tasks and harder tasks. Total demand depends on how quickly usage grows relative to those efficiency gains. At the same time, Meta still has plenty of room to optimize. It could eventually get costs down to ChatGPT Plus levels.

We’ve seen this repeatedly over the past few years: existing workloads get cheaper, and new workloads emerge. Extrapolating from usage at a single point in time tends to miss what comes next. Muse is still very early. Booking flights and negotiating bills may tell us relatively little about how people will use a mature personal agent.

The question we care about is how much users will delegate a year from now. Moving from occasional questions to several agents doing useful work every day increases more than token consumption. It adds CPU work, working memory, and a growing stock of personal data and context. Infrastructure demand per user could rise substantially even as each individual operation gets cheaper.

That is also how we think about Meta’s compute spending. If Muse, Business Agent and glasses drive repeat usage and more demanding workloads, Meta has a reason to keep expanding inference capacity. It also has a reason to invest more in pretraining and RSI to keep its models competitive.

We’d use Freda’s numbers as a baseline and underwrite the investment case against how usage could develop from here. CPU demand could rise above that baseline as users delegate more work, alongside NAND demand that builds as users accumulate data over time. If personal agents retain users and take on a growing share of their work, we think Meta is more likely to be short of compute in 2027 than sitting on excess capacity. > 引用 @FredaDuan: A humble attempt to est. the infra required to serve 100M DAU @Muse

Rough conclusion is:

1 GW of power to serve 100M DAU in the base case, of which only ~0.1 GW comes from the CPU/VM layer. Depending on the # of reasoning-equivalent model calls one Muse DAU generates per day, 3-4GW is entirely plausible. Maybe that’s why @Meta is rumored to be adding 7-10GW of compute next year.

The sandbox layer = sub $1B of CPU content and ~$2B of DRAM content, which is much smaller than many expected.

Lot of moving assumptions. Welcome all feedbacks/ pushbacks.

------ Two very different pieces of infrastructure behind Muse.

1. Muse VM / sandbox infrastructure

2 vCPUs, ~8 GB of RAM and ~100 GB of persistent logical storage per user. https://t.co/7rP5vDKnT7

2. Muse Spark inference

Model inference goes out through Meta's external inference infrastructure. https://t.co/e203XUp6jq ------

1/ Sandbox infrastructure

A. CPU The first mistake is assuming that 100M DAU means 100M VMs are actively consuming compute at the same time.

Suppose the average Muse DAU has an agent actively working for two hours per day.

100M users * 2 hours / 24 hours = ~8M average simultaneous active VMs

Meta obviously cannot provision only for the daily average. Usage will be concentrated during waking hours and bursty.

Assume a 2.5x peak-to-average ratio:

8M * 2.5 = ~20M peak active VMs

Then add roughly 20% capacity headroom: ~25M provisioned live VMs. So the base assumption is effectively that Meta needs enough infrastructure to support roughly 25% of DAU being live simultaneously.

The next important distinction is between virtual CPU allocation and physical CPU demand. Agent sandboxes are particularly well suited to CPU oversubscription. They spend a lot of time waiting. During those periods, the VM may still be alive, but it is barely using CPU.

DeepSeek’s recently published DSec infrastructure provides a useful benchmark. Its production agent sandbox platform runs approximately 30,000 physical CPU cores and 250TB of DRAM across ~160 nodes, with peak concurrency above 380,000 sandboxes. https://t.co/PFlgc6nnUs DSec also demonstrates stable operation at around: 800 microVMs per node. With roughly 188 physical cores per node: 188 physical cores / 800 microVMs = ~0.23 physical cores per live VM.

DeepSeek is obviously the King of efficiency. The number for Muse might be at 0.3-0.75 physical cores per live VM, or assume 0.5 physical cores per live VM as the base case. That is equivalent to roughly two simultaneously live Muse VMs per physical CPU core.

Using the base assumptions: 25M live VMs * 0.5 physical cores per VM = 12.5M physical CPU cores.

On a 256-core CPU: 12.5M cores / 256 cores per CPU = ~50K CPUs; Or on a 192-core CPU that would be 65K CPUs.

At the current public pricing, that is ~$800M.

B. DRAM CPU can be aggressively oversubscribed because a VM that is waiting may consume almost no CPU. Memory is harder to oversubscribe because a live VM still needs to retain its working state.

Muse exposes roughly 8GB of RAM to the user environment, but one observed instance was actually using only around 3GB at the time of measurement.

25M live VMs * 3GB = 75PB of physical DRAM, call it ~75-100PB of physical DRAM feels like a reasonable base range.

At the current public pricing, that is ~$2B.

C. Sandbox power ~0.1 GW for the entire Muse sandbox / VM layer at 100M DAU.

------

2/ Inference Muse’s personal computer executes tools and stores state locally, but the actual model runs on separate inference infrastructure. Meta’s Muse architecture

Energy per inference event Microsoft’s 2026 study estimates that optimized frontier-scale inference consumes a median of approximately: 0.31Wh per normal query

But a long reasoning query with roughly 15x the token count consumes approximately 13x as much energy, or around: 4Wh per long reasoning query

The study specifically highlights reasoning and agentic workloads as significantly more energy intensive. https://t.co/z1YEsvVeMo

Sensitivity analysis on # reasoning-equivalent events per DAU per day

Suppose each active @Muse user generates the equivalent of 50 heavy inference events per day.

At 5Wh each:

100M users * 50 events/day * 5Wh = 25GWh/day

25GWh/day / 24 hours = ~1.0GW average power

So inference alone could require: ~1-2GW of average power

A 3-4GW Muse is entirely plausible. Maybe that’s why @Meta is rumored to be adding 7-10GW of compute next year.

------ The popular framing around Muse is that giving every user 2 vCPUs and 8GB of RAM creates an enormous CPU requirement. But the naive calculation materially exaggerates the CPU requirement because it treats logical VM allocation as dedicated physical infrastructure.

The more interesting conclusion is: Consumer agents may be a meaningful new demand driver for CPUs and conventional DRAM, but inference remains the real compute bottleneck. And as agents do more work, run longer trajectories and increasingly spawn other agents, inference demand can scale much faster than the number of users itself.

+++ Calling my peer review group: @bubbleboi @damnang2 @Midnight_Captl @FundaAI @fi56622380 . Feedback/ Pushbacks pls :).

++ Better formatted: https://t.co/qVwYKNr1ZA https://x.com/FundaAI/status/2104408164154454466