Citi forecasts severe memory undersupply through 2031, with HBM demand surging 62% in 2027 and 69% in 2028, while DRAM and NAND supply growth will lag demand, creating deficits of 8-10%. A commentator argues that AI labs underestimated efficiency improvements in model architectures and training techniques, which could significantly reduce semiconductor demand growth projections.
Fusion is a new dual-model architecture for Devin Desktop and CLI that pairs a frontier model for planning and review with a cost-effective model for execution, achieving up to 39% better efficiency on coding benchmarks. The system runs two parallel agents with separate contexts, allowing the lead model to maintain control while the sidekick handles implementation, avoiding the pitfalls of traditional model routing. Devin reports that using more expensive, token-efficient models can reduce overall costs by delegating effectively and maintaining prompt caches.
GLiClass is an open-source zero-shot sequence classification model inspired by GLiNER that achieves comparable performance to cross-encoder models while being 10 times faster through single forward pass classification. It supports hierarchical labels, in-context examples, custom prompts, and long document chunking for improved accuracy and flexibility.
Occamy-1.0 is a 35B parameter language model optimized for multi-step workflows combining information gathering, tool use, and coding. Trained on execution-grounded data, it achieves competitive performance with larger frontier models while maintaining cost efficiency on agentic benchmarks. The model weights and training data are released to support research on practical co-work agents.
TypeSafe AI announced Jev, a new frontier model optimized for structured decision-making and automation rather than text generation. Jev achieves comparable intelligence to existing large language models while being 40-400x cheaper and 20-200x faster, with parallel sampling that eliminates hallucinations and type errors.
Carnot engines achieve 68-72% thermal efficiency compared to conventional engines' 25-35% by eliminating cooling systems and heat losses, while supporting multiple fuel types including hydrogen, diesel, and biofuels. This doubling of efficiency reduces fuel consumption and emissions by approximately 50%, with applications targeting hard-to-abate sectors like marine, heavy-duty vehicles, and off-grid power generation.
Real estate agents struggle with manually rewriting property descriptions for different buyer types while risking Fair Housing violations. Automated solutions like PropDesc-AI can tailor descriptions for investors, homeowners, commercial tenants, and renters while maintaining compliance and saving time.
Volkswagen's Mission Efficiency prototype electric car set three world records for efficiency, achieving a 0.158 Cd drag coefficient, consuming 6.48 kWh/100 km in controlled tests and 7.51 kWh/100 km in real-world conditions over a 1,278 km journey. The vehicle uses production-ready technology based on the MEB+ platform and ID. Polo drivetrain, with aerodynamic design being its key strength.
Meta is selectively rehiring managers after spending the past year flattening its organizational structure to become more AI-driven, according to internal reorganization efforts within its Applied AI division. The move represents a partial reversal of the company's efficiency push and highlights tensions between maintaining lean operations and coordinating rapid AI development across its workforce.
Researchers propose intelligence per watt (IPW) as a metric to measure how efficiently local AI models can answer real-world queries on power-constrained devices. Evaluating 20+ local language models across 1M queries, they find local models successfully answer 88.7% of queries with IPW improving 5.3x from 2023-2025, demonstrating that local inference can redistribute significant demand from centralized cloud infrastructure.
An engineer recounts how Orange Portails' network team, initially highly efficient, became dysfunctional after management introduced KPIs based on ticket count. To meet targets, the team fragmented single tasks into multiple tickets, creating bureaucratic overhead that slowed service delivery from weeks to months, exemplifying Goodhart's law where optimizing for metrics undermines actual performance.
MoBA (Mixture of Block Attention) is a novel attention mechanism for long-context LLMs that applies Mixture of Experts principles to reduce computational complexity while allowing models to autonomously determine attention patterns. The approach enables seamless transitions between full and sparse attention and has been deployed in Kimi's long-context system.
Level-5 CEO Akihiro Hino apologized after fans discovered undisclosed AI use in the studio's digital showcase, confirming generative AI was used to make presentation footage "more spectacular." Hino stated that core creative elements like scenarios and character designs remain human-made, with AI primarily used for efficiency in converting artwork to polygons and digitizing creative work.
Samsung's Galaxy S27 Pro and S27 Ultra will feature M16 OLED panels, marking the first major OLED upgrade in three years, offering higher brightness, power efficiency, and color accuracy. The S27 and S27+ will use M14 panels. Samsung's foldable Z series is also expected to adopt M16 panels in 2027.
Recurrent Looped Transformer (RLT) combines a causal encoder with a recurrent decoder that grows temporal depth with each token, traversing tL_D decoder blocks after t tokens while maintaining fixed per-token computation. The architecture integrates global key–value memory, layerwise sliding-window attention caches, and continuous latent computation across prompts and responses, with co-design considerations for hardware efficiency and reinforcement learning scaling.
The article uses Gaudí's Sagrada Família as a starting point to explore 'enshittification'—the gradual degradation of quality across architecture, technology, and institutions when profit and efficiency override craftsmanship and human-centered design. It traces this decline from modernism through digital platforms and into healthcare and education, where metrics-driven systems replace genuine care and intentional work.
A researcher at AI2 describes their transition from quantization research to coding agents, detailing how a small team of five researchers and 32 GPUs developed Sera, a method to finetune large language models on private codebases for efficient coding agent deployment. The work eventually scaled to 96 GPUs and enables cheap specialization of models rivaling larger teacher models on private data.
Nitro Skills is a set of techniques designed to reduce token consumption in Claude agents by optimizing file edits, caching outputs, inspecting APIs, compressing repeated information, and tracking token usage across various development workflows.
Europe's economy has proven resilient despite energy supply disruptions, with growth forecasts down only marginally despite Brent crude surging 58% year-on-year and the Strait of Hormuz effectively closed since March. The EU's two-decade investment in energy efficiency—running on 44% less energy per euro of output since 1995 and cutting emissions 40% since 1990—is shielding it from expected shocks, though efficiency gains appear as GDP decline rather than growth.
X users discuss stablecoin infrastructure development, highlighting Spark Finance's shared liquidity FX layer that has processed $10B+ in volume since June 2026, enabling multiple stablecoins (RLUSD, USDC, USDT, PYUSD) to compete at the asset layer while cooperating on liquidity. Ethena's USDe expands to TRON network with cross-chain bridging via Stargate Finance, reflecting broader adoption of stablecoin ecosystems.