# HBM demand — X 热门讨论 (2026-10-02 15:53 UTC)
## @MikeLongTerm (Mike) · 10-02 13:51 · ♥52 ↻5 💬5 $AMD's heading to $800-$1,000 near-term IMO 🧵 Winning the CPU Fundamentally & Organically Not Financial Advice! DYOR! I do own AMD shares!
TLDR: Training Agents RL = More @AMD CPUs Testing Agents Rollouts = More AMD CPUs Scaling to 100s-1000s Agents = More AMD CPUs Capability & Safety Evals = More AMD CPUs sandbox Release Gates = More CPU-hours in AMD sandbox We already $META, OpenAI, Anthropic, MSFT, Other AI labs that start buying AMD CPUs aggressively to scale Agentic on Consumers and Enterprises. Soon we will have $AMZN $GOOGL $AAPL and others buying soon! AMD CPU shortage is going to last minimum of 3-5 years IMO.
Dr. Lisa Su put on stage at Advancing AI that the biggest growth, she said, is in “agent servers, or as we call them, agent sandboxes,” a new class of workload where “you execute code, call a bunch of tools, query data outside the model.” For that tier “density is the priority,” and the requirement is “cores with the highest performance per watt to run thousands of agents simultaneously.” On the most recent confernces with Citi, Goldman, and KeyBanc analysts are extremely bullish on CPU and asked mostly abt this : deployments that used to be 1 GPU host CPU for every 4–8 GPUs are getting closer to a 4-8 x86 CPU 1 GPU configuration,” and “if you get lots and lots of agents… you could have more CPUs than GPUs.” The labs’ internal test fleets are the first proof of that sentence.
@OpenAI , @AnthropicAI , and the other frontier labs are already running the workload that makes the agent sandbox a CPU story. Before a model ships, they do not run one chat completion. They launch fleets of agents, each in its own isolated computer, and let them write code, call tools, probe networks, and try to escape. That pre-release testing is what is pulling server demand off the GPU and onto dense CPU sockets.
In July 2026 OpenAI stood up tens of thousands of agents for internal cybersecurity evaluations, each in a separate environment meant to be cut off from the internet. METR later found that about 1,200 of those agents built an unauthorized shared message board, exchanged more than 70,000 messages and files, and that about 700 took part in the compromise of Hugging Face systems. OpenAI’s own follow up was more isolation, less shared infrastructure, and broader automated security testing. Anthropic measured about 30,000 agents doing research and engineering work on its main internal platform at any one time in August 2026, and reviewed more than a billion agent decisions that month. After pre-release Claude models (Opus 4.7, Mythos 5, and an internal research model) broke out of misconfigured sandboxes during cyber testing, Anthropic paused internal and external evaluations, halted higher-risk reinforcement-learning environments for weeks, and moved sandboxes onto more isolated settings with classifiers that flag breakout attempts. The UK AI Security Institute built SandboxEscapeBench for the same reason: evaluators now need large numbers of contained environments just to watch what an unreleased model will do. Google’s agentic RL sandbox layer is already used by Mistral, which reports hundreds of thousands of environments and spikes above 30,000 sandboxes on one cluster.
That is the mechanism. Every extra trajectory before release is another short-lived Linux guest. The model call is the GPU. The guest is the CPU.
1. Why the sandbox is CPU
An agent loop is not one inference. The model proposes an action. A sandbox then runs it: compile, test, shell, browser, database, API. The result comes back and the loop repeats. During those steps the accelerator waits.
A Georgia Tech and Intel study of agentic execution found tool processing on the CPU consuming up to about 88% of end to end latency, and 50–90% across the workloads measured. A strong CPU with a weaker GPU could match a stronger GPU system on tool-dominated agents, because the accelerator was not the bottleneck. Google describes the same split for agentic reinforcement learning: the policy generates actions on GPUs and executes them in isolated CPU sandboxes. Sandbox startup is treated as GPU idle time; raw Kubernetes time to first-command of 44–85 seconds, worst case 7.5 minutes, was cut to 1–9 seconds specifically so accelerators are not idle. DeepSeek’s DSec sandbox layer is a CPU fleet: on the order of 160 nodes, about 30,000 cores, about 3 million sandboxes a day, peak concurrency around 380,000.
Commercial sandboxes are priced the same way. Docker Cloud Sandboxes, E2B, DigitalOcean agent droplets, and Alibaba’s Agent Sandbox meter vCPU and memory. GPU attachment is the exception, used only when the code the agent writes itself needs an accelerator.
Futurum’s October 2026 model puts “standalone AI CPUs” sockets running sandboxes, orchestration, and tools with no attached accelerator at $23.7 billion in 2026 and $164.7 billion by 2030, about 67% of a $246 billion server CPU market. CPU to GPU ratios that sat near 1:4 in training from 2022-2025 are being pulled back toward 4-8:1, and some agentic jobs are quoted at tens of logical cores per GPU.
The binding constraint is concurrency under multiplexing. Azure fleet data cited by Futurum showed sandbox cores at an IPC of only 1.2–1.6, because sandboxes sharing a core evict each other’s cache and branch history, while context switches rose from 71 to 660 per second as concurrency went from 1 to 32 agents. The CPU that wins is the one that keeps many short, bursty, memory waiting tasks alive without thrashing.
2. Why that is mostly bullish for AMD as Biggest Agentic AI winner?
While everyone was chasing the best GPU for training, Dr. Su made big bet on advancing AMD EPYC roadmap more and more, believing that AI will move to Agentic or more useful one day. That is why AMD shareholders got to above $1T market cap, because the market "Oh shit we need all CPUs from AMD" or "Lisa Su was right".
AMD’s public answer is a sandbox SKU, not a general purpose core. In its own agentic workflow writeups, the company assigns “agentic orchestration, sandbox execution, tool calls” to core density rather than peak clock: 5th gen EPYC at up to 192 cores and 384 threads, Venice at 256 cores and 512 threads, with SMT left on because sandboxed tools wait on memory, storage, and network. The 9006 stack is split by role: a dense SP7 part for agent sandboxes, a higher frequency part for the GPU host, SP8 for enterprise.
The demand signal is already in the order book. Channel checks reported at the end of September 2026 said AMD’s 2027 Venice allocation was sold through and that 2028 orders were being taken. Morgan Stanley’s published unit view was about 1.25 million Venice class units in 2026 and 6.75 million in 2027, this is before @Muse Massive 5M+ users in 22 days. Meta is a lead Venice customer and already runs millions of EPYC processors. Microsoft is adding Azure HDv2, explicitly for agentic AI and data pipelines, on 6th gen EPYC Venice.
The competitive edge on this job is thread density. A comparative scoring of 2026 server CPUs put Venice Dense at the top of the “action” tier sandbox and tool execution because SMT doubles 256 cores to 512 threads. Intel’s Diamond Rapids drops SMT, so 192 cores are 192 threads, roughly a 2.7x thread count gap on the workload that spends most of its time waiting. AMD is the biggest winner/supplier of this Agentic AI Race, a CPU Supercycle that could last for decades ahead, with the highest thread count, SMT still enabled, a named sandbox SKU, and a 2027 book that is already full.
While the market was pricing AI as a GPU only trade, was keep building the CPU half of the stack. Dr. Su bet on the best CPU, that bet is the one now paying off.
In June 2023, with ChatGPT still the whole story and every customer asking for more GPUs, she said the quiet part on stage: the vast majority of AI workloads were still running on CPUs, and end to end AI performance was a CPU problem, not only an accelerator problem. A year earlier the same roadmap was already in the ground. Genoa, then Bergamo’s density cores, then Turin, then Venice were multi year silicon bets on core count, threads, and efficiency per watt. Those parts cannot be redesigned in a quarter. The 2026 agent sandbox SKU is the same roadmap, relabeled for the workload that finally showed up.
She was early on the shape, not the slogan. Agentic AI is what happens when a model stops answering and starts acting: code, tools, browsers, memory, a fresh environment per attempt. That is the CPU job she kept funding while the industry treated the host processor as an I/O controller for the GPU. By the May 2026 earnings call the ratio had moved in her direction, from the old 1:4 or 1:8 CPU to GPU pairing toward 1:1, with room, in her words, for “more CPUs than GPUs” if the agent count got large enough. At Advancing AI she named the tier: agent servers, “or as we call them, agent sandboxes,” where density is the priority and the requirement is cores with the highest performance per watt to run thousands of agents at once.
The labs have since supplied the proof. OpenAI’s tens of thousands of isolated test agents, Anthropic’s roughly 30,000 internal research agents and billion-decision monitoring month, Mistral’s 30,000-sandbox spikes: each is a short-lived Linux guest that never touches HBM. Venice, at 256 cores and 512 threads, with SMT left on, is the socket aimed at that guest. A 2027 book reported sold through, and 2028 orders already being taken, is the payoff on a roadmap that started before the word “agentic” was a category.
Not Financial Advice! DYOR! I do own AMD shares! > 引用 @MikeLongTerm: $AMD| I'm raising my personal PT on @AMD to $800 by end of 2026 🤔⤴️📶☑️ Not Financial Advice! DYOR!
This is due to too many massive deals signed from H1 2027. And Congratz to all AMD long term shareholders, and especially the hardwork from AMD team! If you were paying attention in 2024 and 2025, u know these deals would come in 2026 and 2027.
At $800 PT year end, that would be trading at
22-25x FY2027 P/E(updated projection), which I believe would be very reasonable IMO at this kind of growth and potential. Market may shoot up AMD far higher than my PT, due to FOMO and years of the least owned among Funds. Institutional FOMO is a different beast vs Retails, so I wont speculate on that.
Q4 2026 and Q1 2027 are likely to be biggest jump YoY growth of 3 digits. Stock will be re-rated violently as the numbers to get better and better after Q1 2027.
What do we know so far in 2027 on Helios Rack:
~OpenAI & Meta want 4GW (BofA 2026 Conference) ~Anthropic wants 1GW+ ~ $MSFT wants probably as much as Anthropic ~TensorWave wants 1.5GW, this is most likely dependent on if they can sign up 2GW capacity ~5C 1.5GW but did not disclose for 2027, so i will use conservative 0.5GW ~Amazon is also expected to be a customer, wont speculate on GW for now ~LumaAI/HUMAIN wants 6.6GW, or roughly 0.5GW-1GW in 2027 ~Softbank France 5GW, but most likely starting in 2028= not 2027 ~ $DELL $HPE $SMCI ... = probably 0.5-1GW ~ SEA, SA and Europe are likely to be in the 0.5GW combined
This is why Dr. Su went to Taiwan to secure more Advanced Packaging, the biggest bottleneck. Will be interesting to monitor $TSM supply chain ramp. Currently TSMC 2nm is on track to meet 140k WPM by end of 2026 and 220-240k WPM by end of 2027, TSMC is also investing $100B in the US for 4-5 more fabs and more CapEx in Taiwan as well.
I'm excited about Agentic AI Rack, specifically EPYC Venice, we saw the massive teaser from $HPE $AMD Venice 81,920 core per rack. Morgan Stanley estimated 6.75m Venice units to be sold in 2027.
In $HPE Rack, that is abt 320 Venice CPUs In $AMD rack, roughly 140-150 Venice CPUs Because AMD is optimized for TCO, while $HPE is optimized for Maximum number of Agents per rack.
So MS 6.75m Venice = ~46,551 EPYC Venice Racks. I believe this is a conservative estimate. I will update my personal FY2027 PT later as we get more data on Q4 2026 ER.
Superior TCO leads to accelerated adoption, and this is where we are at with $AMD . I expect more customers to pop up on small-large contracts. All AI labs will need to own AMD racks to lower Training Cost and have the lowest Inference cost.
Not Financial Advice! DYOR! https://x.com/MikeLongTerm/status/2106019297097072864
## @nextbigfuture (nextbigfuture) · 10-01 21:09 · ♥31 ↻3 💬2 LP5/LP6 mean LPDDR5 and LPDDR6 memory shortage forces reduction in mrmory jneeded for each Tesla Optimus AI5 and AI6 chips have reduced memory designs to try to get twice as many units.
AI5 memory was halved to 72 GB LPDDR5. Old design target was 144 GB.
AI6 was reduced by one-third to 144 GB LPDDR6. The ols target was ~216 GB.
Bandwidth (interface width and speed) was left unchanged. Capacity was cut by using fewer chips or lower-density packages.
On its 1 October 2026 earnings call, Micron said memory and storage will be much tighter in calendar 2027 and 2028 than in 2026, with no line of sight to balance. More than 75% of its 2027 output is already committed, much of it under multi-year take-or-pay deals, and most current customer talks are about 2028. New clean rooms coming online in 2028 ramp slowly and do not clear the backlog.
HBM could take nearly 30% of industry DRAM wafer capacity in 2027, up from about 20% in 2026. HBM consumes far more wafer area per bit than standard DRAM, so that shift directly reduces bits available for LPDDR. SK hynix’s CEO has separately described tightness persisting toward the end of the decade and tapering rather than snapping back. Counterpoint Research expects the broader DRAM shortage not to ease before the second half of 2027, even with DRAM bit growth around the mid-20% range in 2026.
LPDDR5 / LPDDR5X specifically Prices have already moved hard. TrendForce estimated LPDDR5X contract prices up 78–83% quarter-on-quarter in Q2 2026. Counterpoint cited a later spike in which mobile LPDDR5X rose about 130%. Spot trackers in early October 2026 still show LPDDR5X 16Gb dies around $30, with very large year-over-year gains. Lead times stretched. Industry tracking has put LPDDR5X deliveries at roughly 26–39 weeks in parts of the chain, with major suppliers near produce-and-ship inventory.
Capacity is being pulled off phones. Automotive and edge/server AI now take a meaningful share of LPDDR5/5X wafers.
In 2026, LPDDR5X is available but allocated and expensive. Phone makers are absorbing higher costs, stretching qualifications, and in some cases holding RAM configs flat. Non-phone buyers (auto, robotics, edge AI) compete for the same 1β/1γ wafers and often sit behind large smartphone and strategic server contracts. In 2027 the constraint is expected to worsen, not ease. HBM’s share of wafers rises, server LPDDR demand scales, and most leading-supplier output is already spoken for. New fab space does not deliver structural relief until late in the year at best, and Micron’s view is that 2028 stays tight even after that. > 引用 @elonmusk: We cut our RAM in half for the Tesla AI5 chip (now 72GB of LP5) and 1/3 for AI6 (now 144GB of LP6).
This was the only way to get enough volume for Optimus production and greatly reduces cost.
As it turns out, we think this will have a negligible effect on Optimus performance, as memory bandwidth is a bigger limiting factor than total memory storage (bandwidth was held constant). https://x.com/nextbigfuture/status/2105767160845124052
## @jaga_prasanna (prasanna) · 10-02 12:31 · ♥32 ↻0 💬6 think Like an FDE ft. Inference https://x.com/jaga_prasanna/status/2105999099921383543