# AI capex — X 热门讨论 (2026-09-11 19:18 UTC)
## @porquettfin (Tendencias Finanzas) · 09-11 15:03 · ♥30 ↻8 💬0 "Oracle"
Porque la empresa acaba de mostrar números de boom de IA: su infraestructura cloud creció 121% y el backlog llegó a US$664.000 millones.
La acción abrió con una suba de más del 6%… y después devolvió casi todo. La pregunta es por qué.
Los números fueron muy fuertes: ingresos de US$19.300 millones (+30%), cloud de US$11.600 millones (+62%) y EPS ajustado de US$1,92 vs. US$1,74 esperado.
Además, $ORCL sumó más de US$30.000 millones en nuevos contratos de AI Cloud durante el trimestre. La demanda claramente no parece ser el problema.
El problema está del otro lado del balance.
Oracle gastó US$28.500 millones en CapEx en apenas un trimestre y generó un flujo de caja libre negativo de aproximadamente US$5.400 millones. Para todo el año espera invertir entre US$90.000 y US$95.000 millones. https://x.com/porquettfin/status/2098427223510294719
## @rzayev7895 (Yunis Rzayev) · 09-11 17:16 · ♥36 ↻1 💬6 $SNDK $MU DeepSeek V4.1-Flash neden bellek hisselerini baskıladı ve neden bundan korkmuyorum?
Not: Bu analizi DeepSeek'in paylaştığı rakamların doğru olduğunu varsayarak yaptık. Ancak DeepSeek'in geçmişte yaptığı bazı açıklamalara ve sonradan ortaya çıkan verilere baktığımızda, şirketin teknolojik gelişmelerini zaman zaman oldukça iddialı şekilde sunduğunu da unutmamak gerekiyor. Dolayısıyla burada verilen %75 HBM ve %87,5 depolama tasarrufu gibi rakamları şimdilik şirketin kendi verileri olarak değerlendirmek, bunların gerçek dünya ölçeğindeki etkisini ise bağımsız veriler ve hyperscaler sonuçları üzerinden takip etmek daha doğru olacaktır.
Bugün bellek hisselerinde gördüğümüz zayıflığın önemli nedenlerinden biri DeepSeek'in V4.1-Flash modelini duyurması. İlk bakışta haber gerçekten negatif görünüyor. DeepSeek, yeni modelde KV cache için gereken HBM miktarını yaklaşık %75, kalıcı SSD depolama ihtiyacını ise %87,5 azaltmış durumda. Global KV cache gereksinimi token başına 890 byte'a kadar düşürülmüş. V4-Flash'ta bu rakam 3.514 byte, V3.2'de ise 48.068 byte seviyesindeydi.
Piyasanın ilk tepkisi de oldukça basit: Eğer AI modelleri aynı işi daha az bellekle yapabiliyorsa, gelecekte HBM talebi azalabilir. Dolayısıyla SK Hynix, Micron ve Samsung gibi bellek üreticilerinin büyüme beklentileri sorgulanmaya başlıyor.
Ancak burada bence çok önemli bir ayrım var. DeepSeek'in açıkladığı %75'lik düşüş, toplam GPU belleğinin %75 azalması anlamına gelmiyor. Bu rakam KV cache'in HBM gereksinimiyle ilgili. Model ağırlıkları, hesaplama için gereken bellek, GPU'lar, networking, güç altyapısı ve veri merkezinin diğer bileşenleri bundan bağımsız olarak varlığını sürdürüyor. Benzer şekilde %87,5'lik rakam da toplam veri merkezi depolamasının %87,5 azalacağı anlamına gelmiyor.
Daha da önemlisi, DeepSeek'in amacı AI kullanımını azaltmak değil. Tam tersine, modeli daha düşük maliyetle çalıştırabilmek. V4.1-Flash toplam 552 milyar parametreye sahip ve Causal Encoder–Decoder mimarisi sayesinde girişlerde yaklaşık 8 milyar, çıktılarda ise 16 milyar parametre aktive ediyor. Yani burada gördüğümüz şey AI altyapısının ortadan kalkması değil, aynı altyapıyla daha fazla iş yapılabilmesi.
Bence asıl bakılması gereken nokta da bu.
AI'da birim başına kullanılan bellek miktarı düşebilir ama AI kullanımı bundan çok daha hızlı büyüyebilir. Bir modelin inference maliyeti düştükçe şirketlerin daha fazla agent kullanması, daha uzun context'lerle çalışması ve daha fazla inference gerçekleştirmesi ekonomik hale geliyor. Özellikle agentic AI tarafında modeller sürekli olarak kod, doküman, geçmiş konuşmalar ve araç sonuçlarıyla çalışacağı için context yönetimi çok daha önemli hale gelecek.
Dolayısıyla denklem sadece "bir milyon token için ne kadar HBM gerekiyor?" şeklinde kurulamaz. Asıl soru, önümüzdeki yıllarda kaç milyar token işleneceği. Token başına bellek ihtiyacının düşmesi olumlu bir gelişme olabilir; çünkü AI kullanımını daha ucuz hale getirir. Kullanım arttığında ise toplam compute, networking, storage ve bellek talebi yine büyüyebilir.
SSD tarafındaki gelişmeyi de benzer şekilde değerlendiriyorum. DeepSeek, kısa ömürlü sliding-window attention verilerini kalıcı SSD üzerinde tutmak yerine sunucunun normal belleğindeki paylaşılan bir havuza taşıyor. Bu veriler dakikalar sonra zaten değerini kaybediyor. Global context ise daha uzun süre tutuluyor. Eksik kısa süreli cache gerektiğinde de sınırlı yeniden hesaplama yapılabiliyor. Burada yapılan şey AI'ın depolama ihtiyacını ortadan kaldırmak değil, hangi verinin nerede ve ne kadar süre tutulacağını daha verimli hale getirmek.
Bence piyasadaki en büyük hata, "daha verimli model = daha az toplam bellek talebi" şeklinde doğrudan bir bağlantı kurmak. Teknoloji tarihinde bunun tam tersini çok defa gördük. Bir teknolojinin birim maliyeti düştüğünde kullanım genellikle artıyor. AI için de benzer bir durum yaşanabilir. Inference ucuzladıkça daha fazla şirket AI agent kullanmaya başlayabilir, mevcut kullanıcılar daha fazla inference gerçekleştirebilir ve AI uygulamalarının kullanım alanı genişleyebilir.
Bu nedenle benim için asıl takip edilmesi gereken DeepSeek'in 890 byte'lık KV cache rakamı değil. Hyperscaler'ların 2027 ve sonrasındaki AI capex planları. NVIDIA ve AMD accelerator siparişleri, HBM kontratları, HBM fiyatları ve bellek üreticilerinin ileriye dönük kapasite planlarında bir bozulma görüyor muyuz? Eğer bunlarda ciddi bir değişiklik yoksa, tek başına bir modelin daha verimli hale gelmesi AI memory tezinin bozulduğu anlamına gelmez.
Hatta uzun vadede bunun olumlu bir tarafı bile var. AI'ın çalışma maliyeti düştükçe daha önce ekonomik olmayan kullanım alanları mümkün hale geliyor. Bugün bir şirketin 100 AI agent çalıştırması pahalı olabilir; ancak inference maliyetleri ciddi şekilde düştüğünde bu sayı çok daha yukarı çıkabilir. Dolayısıyla model başına daha az bellek kullanılırken toplam AI iş yükü çok daha hızlı büyüyebilir.
Bu yüzden bugün MU, SK Hynix veya SNDK gibi hisselerde yaşanan zayıflığı şu aşamada temel tez kırılması olarak görmüyorum. Elbette DeepSeek ve benzeri modellerin bellek verimliliğini sürekli artırması uzun vadede izlenmesi gereken bir konu. Fakat bunun gerçekleşmesi için önce AI capex'in, HBM kontratlarının ve gerçek bellek talebinin aşağı geldiğini görmemiz gerekiyor.
Şimdilik gördüğümüz şey bana daha çok piyasanın yeni bir teknoloji gelişmesini çok hızlı şekilde fiyatlaması gibi geliyor. DeepSeek, AI'ın bellek ihtiyacını ortadan kaldırmadı; AI'ın aynı işi daha az kaynakla yapabilmesini sağladı.
Aradaki fark oldukça önemli.
Kısacası, korkmamız gereken şey daha verimli AI modelleri değil. Korkmamız gereken şey AI kullanımının ve hyperscaler capex'inin gerçekten yavaşlaması. Şu anda ikincisini gösteren bir tablo olmadığı sürece, bugünkü hareketi bellek hisselerinde uzun vadeli tezin bozulması olarak okumak için erken olduğunu düşünüyorum.
Yatırım tavsiyesi değildir. Kendi araştırmanızı yapın. > 引用 @wallstengine: DeepSeek just launched V4.1-Flash, cutting KV-cache HBM requirements by roughly 75% and persistent SSD storage by 87.5% compared with V4-Flash.
Its global KV cache is down to 890 bytes per token, versus 3,514 for V4-Flash and 48,068 for V3.2. That is roughly 54x smaller than last December’s model.
V4.1-Flash has 552B total parameters, but its new Causal Encoder–Decoder architecture activates just 8B when processing inputs and 16B when generating outputs.
It also adds native image understanding, alongside new pretraining methods and larger-scale reinforcement learning.
The memory savings matter because AI agents repeatedly reuse context from conversations, documents, code and tool results.
The KV cache preserves earlier calculations so the model doesn’t have to redo all that work.
At 890 bytes per token, a million tokens would occupy about 890MB of global KV cache. That excludes model weights and other memory needed to run the system.
The SSD reduction comes from changing both what gets stored and where.
Previously, DeepSeek kept global attention data and short-lived sliding-window attention data in its persistent SSD cache. Both typically stayed there for more than 72 hours, even though the sliding-window data was mainly useful for minutes during an active session.
V4.1 moves that short-lived data into a shared pool using 10% of each server’s regular memory. Entries expire after minutes, while global context remains in persistent storage for at least 72 hours.
When short-lived data is missing, a new technique called “bounded replay” reconstructs it using a limited window of tokens instead of the much larger computation previously required. Removing that data from SSDs, combined with compressing the remaining global cache, produces the reported 8x storage reduction.
Reported benchmark scores include 63.9 on Humanity’s Last Exam with tools, 31.2 on Terminal-Bench 4.0 and 54.8 on Automation-Bench.
V4.1-Flash is live through the API as deepseek-flash, with off-peak pricing at half the peak rate. DeepSeek is also replacing V4-Pro with V4.1-Flash starting Sept. 14, ahead of a future V4.1-Pro release.
Note: These are reductions in cache requirements, not a 75% cut in total GPU memory or an 87.5% cut in all storage. They make retaining and reusing context substantially cheaper without eliminating the other costs of running an agent. https://x.com/rzayev7895/status/2098460796116480286
## @perspez (Perspez) · 09-11 05:25 · ♥31 ↻3 💬1 Microsoft is planning one of the largest data center expansions ever, taking Azure capacity from ~12 GW today to more than 38 GW by 2032.
$MSFT $SLNH
• Microsoft plans to add roughly 26 GW of new data center capacity over the next six years, more than tripling its current footprint.
• The 38+ GW includes Microsoft-owned and leased facilities, but excludes compute rented from neoclouds such as CoreWeave, Nebius and Nscale. Microsoft’s effective compute footprint could therefore be even larger.
• AI-specific infrastructure is expected to rise from roughly 2 GW today to ~12–13 GW by 2032, more than a sixfold increase.
• Microsoft has already faced capacity shortages severe enough to turn away some customers seeking AI compute, while also balancing external Azure demand against capacity needed for its own Copilot and AI products.
• The company has increasingly relied on external providers including CoreWeave, Nebius and Nscale while its own infrastructure buildout catches up.
• CFO Amy Hood said very little additional capacity can realistically be built and brought online over the next 12 months, pushing Microsoft to focus further ahead on securing the right land and power.
• Microsoft expects roughly $175B of capex in calendar 2026, including around $50B in fiscal Q1 2027 alone, as infrastructure spending continues to accelerate.
• The company is also extending some long-term data center leases from 15 years to 25 years, reflecting how far ahead hyperscalers are now planning their infrastructure needs.
• While new campuses take years to deliver, Microsoft is also squeezing more capacity from what already exists: its teams have reduced dock-to-live times by more than 50%, shortening the time between hardware arriving and becoming revenue-ready.
• Microsoft has also signed long-term agreements across the infrastructure supply chain, recognizing that securing GPUs alone is not enough: power, electrical equipment, networking, cooling and the rest of the physical stack must all arrive together.
The scale is difficult to overstate:
~12 GW today → 38+ GW by 2032 ~26 GW of additional capacity ~2 GW AI today → ~12–13 GW AI Neocloud capacity on top
And yet Microsoft is still capacity constrained today. > 引用 @perspez: Big hire for $SLNH, @jbelizaireCEO adds another Microsoft Fairwater veteran to its AI infrastructure team
• 6+ years at Microsoft Neil most recently served as Director, AI Construction within Microsoft’s Datacenter Delivery organization.
• He worked on Fairwater His own profile says he was accountable for large-scale AI data center programs across the Milwaukee metro, including MKE03/04 and MKE07–14. MKE03/04 are part of Microsoft’s Fairwater / Mount Pleasant AI buildout.
• That directly overlaps with Ryan Carver’s Microsoft work Ryan Carver led Microsoft’s Fairwater construction program in Wisconsin before joining Soluna as Chief Development Officer.
So we now have two former Microsoft AI infrastructure leaders at Soluna with overlap inside the same hyperscale AI construction program.
• Neil brings exactly the execution skillset Soluna now needs His Microsoft work included: - advanced electrical distribution - liquid-to-chip cooling - centralized utility plants - integrated commissioning - energy management - supply-chain and construction risk - hyperscale AI delivery
• And the transition is notable Microsoft: through Sep. 2026 Soluna: Sep. 2026
No gap. He moved directly from Microsoft into Soluna.
• At Soluna, Neil says he is now leading the construction organization responsible for delivering hyperscale AI infrastructure across a rapidly expanding national portfolio.
Ryan joined Soluna in July to build the development and execution platform.
Now, just weeks later, another Microsoft AI construction executive who worked within the Fairwater ecosystem follows him to Soluna.
This is a very strong signal that Soluna is building a serious hyperscale execution team around its AI pipeline.
The Microsoft connection around SLNH keeps getting stronger. For anyone who missed my 12-part research series on the Soluna–Microsoft connection, you can catch up here: 👇
https://t.co/7K6QKPLxit https://x.com/perspez/status/2098281868663599523