Showing posts sorted by date for query semianalysis. Sort by relevance Show all posts
Showing posts sorted by date for query semianalysis. Sort by relevance Show all posts

Tuesday, September 8, 2026

Google's TPU Chips Will Deliver Up To 50% Better Perfomance Per Dollar On Some Inference Chores Vs. Nvidia (GOOG; NVDA)

Took ya long enough.*

From SemiAnalysis, September 7:

  • TPU Inference Externalization Full Steam Ahead - InferenceX
  • InferenceX, Up to 50% Better Performance per Dollar, Rapid Externalization of TPU stack, Growing Customer Base, Ironwood, TPUv8i, Reducing CUDA Moat 

For more than a decade, the industry has watched Google build an empire on its own silicon. Search, Ads, YouTube, and every generation of Gemini run on TPUs. Few accelerators have attracted as much architectural scrutiny or as much debate about what their performance and economics would look like outside the company that designed them. Anthropic being the biggest user of TPUs, surpassing Deepmind’s own use by 2029.

Google’s internal success was never the question. The question was how much of that advantage the rest of the industry could actually get. Could you take an open-weight model, serve it through a familiar inference engine, and beat NVIDIA on the economics that matter to your business?

Today, we are publishing the first third-party inference results for TPUv7 Ironwood on InferenceX Official Preview. In our apples-to-apples comparisons against B200/B300, Ironwood delivers up to 50% better performance per dollar. Its advantage extends across much of the Pareto curve, and we examine the economics from both sides: Google’s internal total cost of ownership and the external TCO an actual customer pays.

Ironwood (TPUv7) is the first generation in which Google is competing for others’ inference workloads with chips that can be purchased outright or rented through its own cloud. In November 2025, we already said that Anthropic loves TPUs and committed to over one million of them (around 400k+ in direct purchases and 600k+ rented through GCP), used mainly for training but also for inference. Our Accelerator Model has the latest figures for Anthropic’s TPU shipments and Google’s overall TPU shipments by quarter, plus estimates for TPUv8i, v8t, and various v9 / v10, and more

We are excited by how quickly the new TorchTPU stack is developing, the external stack for TPUs. Later in the article, we will discuss the upcoming work needed for TPU software externalization, including optimizing speculative decoding, disaggregated prefill, KV-cache offloading, multi-turn agentic workloads, and more. Even so, we at SemiAnalysis strongly believe that TPU externalization is heading in the right direction and moving full steam ahead. Furthermore, unlike AMD, which is still learning how to build a test-first software culture, Google has decades of software engineering experience and an extremely well established quality-driven culture, so we expect external TPU software to mature rapidly.

In this article, we will cover all the optimizations that went into TPU kernels and the serving stack for open-weight models, including DP attention optimization, MoE routing and kernel optimization, and reducing padding in GDN kernels. We will also take a deep dive into the TPU system and discuss the next steps the amazing TPU performance engineers are pursuing to make the stack widely available.

Google has spent more than a decade demonstrating what it can build with TPUs. Now we get to measure what the rest of the industry can do with them.

Shoutout to the Google (Chris Chan, Jahangir Hasan, Wangyuan Zhang, Anne Stern, Puneith Kaul, Ruizi Dong, Sangam Jindal, Qi Zhou, Madhan Jaganathan, Gang Ji, Jun Wan, Devanshu Jain, Jiaxin Cao, Srinath Mandalapu, Haowen Ning) and Inferact teams and RedHat Teams (Michael Goin) for this amazing TPU foundation and performance! Furthermore, shoutout to the RadixArk team that is also working on TorchTPU SGLang.

InferenceX Official Preview: TPUv7 Ironwood vs. Blackwell and Blackwell Ultra 

We are already seeing strong results from the upcoming native TorchTPU vLLM stack in apples-to-apples comparisons against Nvidia GPUs. Google is using Qwen3.5 397B in FP8 as the initial bring-up model. Later sections take a deep dive into why the new TorchTPU approach is a marked improvement over the previous TorchAX path for external TPU vLLM/SGLang serving.

Once that foundation is in place, Google plans to extend support to other open-weight models, including Kimi K3 and GLM5.3. Once a handful of models are well optimized, we believe adding optimized support for a wide range of popular open models near day 0 becomes far easier. Today, vLLM and SGLang concentrate their day-0 support on Nvidia, with passable day-0 coverage for AMD. We expect TorchTPU to be stable enough in the near future that vLLM and SGLang maintainers may add TPUs to that day-0 list. The stack is expected to leave private beta & be open sourced around October. 

In apples-to-apples comparisons of aggregated serving with FP8 and single-token prediction, we are seeing up to 50% better performance per dollar from TPU than from B200 and B300 running FP8 in aggregated serving. When serving models using FP4 on NVIDIA GPUs, there is quality loss verus FP8. TPUv7 does not have native FP4 computation thus on FP4, NVIDIA GPUs still maintains the lead. This will change with TPUv8i which has native FP4 support thus we strongly believe TPUv8i Boardfly will be competitive to Rubin NVL72.... 

....MUCH MORE
*
April 2017 - Watch Out NVIDIA: "Google Details Tensor Chip Powers" (GOOG; NVDA)
We've said NVIDIA probably has a couple year head start but this bears watching, so to speak....

And many, many more, including November 2025 "CHIPS: Google's Tensor Processing Unit (Finally) A Viable Competitor For Nvidia (GOOG; NVDA)

Sunday, August 9, 2026

SemiAnalysis Looks At SpaceX (SPCX; MSFT; NVDA)

Talk about your "big if true"/"big if, true". The numbers are just so mindbending. 

From SemiAnalysis, August 7:

SpaceX 10GW in 2027 – Why It’s Real, Will Drive $300B ARR for SpaceX, and Why Microsoft Will Be the Largest Offtaker
Inference at 100B/GW/year, SpaceX's stellar pace, Microsoft's 10GW 2026 Awakening, Azure Can Grow Triple-Digits

Elon Musk shocked the world, once again, when he announced on SpaceX’s first earnings his Gigawatt ambitions for next year. He “conservatively” aims to build & deliver an incremental 6-8GW in 2027 alone, with potential for that number to be well above +10GW. At 50B per GW, that’s $300-500B in capex in 2027, on par with what we expect from AWS and Google – an unbelievable number for a company significantly less profitable than rival hyperscalers.

Yet, we believe that the number is real. We see SpaceX on track to build about 10GW by year-end 2027. We’ve evaluated all sites suitable for SpaceX and provided the list to our Datacenter Model subscribers. Our Energy Model subscribers also have the precise list of gas generation equipment available, quarter by quarter, by 30+ turbine, engine, fuel cell suppliers. We provided much of this data, before the market woke up to it. Below, we discuss how Elon bypasses typical datacenter construction constraints.

As explained in our Meta Compute deep dive, large-scale + near-term compute is a remarkably scarce combination, and it’s priced at a huge premium – up to $50B/GW/year. However, AI labs can handle it and make a good living off it.

Our Tokenomics Model and our Inference Simulator demonstrate that at realistic performance levels (e.g. tokens/sec per GPU), both OpenAI and Anthropic can generate over $100B/GW/year of revenue when selling API inference on a GB300 cluster. This is significantly more than the costs of renting a GB300 cluster for a year at current neocloud prices.

Serving inference tokens is unbelievably profitable for the frontier model companies.

 

Source: SemiAnalysis Tokenomics Model, SemiAnalysis Inference Simulator

We assume around $12B/GW/year of cost per year, using a conservative rental pricing rate of $3/GPU-hr, and make a token production estimate using our Inference Simulator with a frontier-class model architecture and our agentic coding benchmark, AgentX (part of InferenceX), which is built by collecting real production coding traces. We blend that token production rate between input, cache-read, cache-write, and output token costs at our real workload ratios, and produce the final estimate, exceeding $100B/GW/year.

For background, our Inference Simulator is built from the ground up with a fundamental understanding of how modern AI accelerators work. We build a roofline and realistic performance model for how frontier models work during inference, with timings for every operation and a real trace output. It is an end-to-end simulation of the actual workload executing on the actual silicon. We have validated the simulators fidelity on a wide range of accelerators and workloads and continue to improve its ability to accurately forecast performance of future accelerators based on design specifications....

***** 

...Please reach out to sales@semianalysis.com for more information on how we apply the Inference Simulator for custom research and analysis.

Beyond OpenAI and Anthropic, there is actually a third company in the world capable of printing such economics per GW: Microsoft. Having full access to OpenAI models, they can generate the exact same revenue and margin per MW, while paying none of the training costs. Satya nailed the negotiations with OpenAI: the deal reworked in April 2026 dropped the old 20% revenue share from the equation. Put simply, Microsoft has a giant incentive to procure as many MWs as possible, as fast as possible. While much of their datacenter capacity currently goes to OpenAI at ~14M/MW/year, they have the opportunity to improve that mix. The potential impact is Microsoft Azure accelerating revenue growth from ~42% to over 100% by next year. A once-in-a-generation opportunity, that SpaceX is incredibly well positioned to serve.

 

Source: SemiAnalysis Tokenomics Model

While Microsoft signing 3GW with SpaceX for 50B/GW/year sounds insane, we view it as possible for two reasons:

  • 1/ Microsoft is already preparing for an epic datacenter ramp. As discussed below, they’ve signed 10GW of contracts year-to-date, for over $300B of total contract value (not including the GPU cost). We expect much more to be signed. Caveat: these contracts contribute to late 2027 and 2028 capacity. There is a near-term gap to fill.

  • 2/ With a 90-day cancellation policy, akin to the SpaceX deals with Anthropic and Google, there is zero balance sheet risk. This is remarkably easy for Amy Hood to sign off, given the revenue opportunity.

For SpaceX, the next natural question is financing. How can Elon afford to pay so much CapEx without the balance sheet of the leading hyperscalers? We expect a combination of the two following items:

  • 1/ Support from Nvidia, in the form of vendor financing to lower the upfront cash cost. This is likely why Elon declared to be Nvidia exclusive on the earnings call! As our Accelerator Model has repeatedly explained, xAI/SpaceX have actively evaluated alternatives like TPU and AMD – so the financial argument likely made them abandon these and focus on Nvidia.

  • 2/ Operating cash-flow financing led by industry-high pricing, enabled by fastest timelines: SpaceX will continue to sell large-scale compute with 3-5 months lead time, an unbeatable offering, and price it accordingly at 30-50M/MW/year. That pays back the capex in less than a year. We dived into this in our Meta Compute article.

The implications of this are a path to $300B of ARR by the end of 2027 for SpaceX. This assumes only 50% of their 2027 incremental compute is monetized, the reminder being for the Grok & Cursor teams for training (no inference revenue modelled)....

....MUCH MORE 

We've chronicled much of the Elon - Jensen frenemy relationship in real-time for over a decade. On August 4 Musk said SPCX would use NVDA's platforms exclusively.

It wasn't always apparent that this is how things would turn out but the two centi-billionaires seem to get along. From a July 2023 post, "CORRECTED—Earnings - Tesla Reports, Stock Slides, Elon's Buying A Supercomputer (TSLA)":

....For some background on Tesla and AI here is our introduction to June 9's "Elon Musk Predicts Nvidia’s Monopoly in A.I. Chips Won’t Last" (NVDA; TSLA)":

Before we get to the headline story, some background. Tesla and Nvidia have a history.

In 2015 - 2016 when everyone thought that autonomous driving was just around the corner, the challenge was seen as both a sensor issue, for example: LIDAR vs cameras, and a machine learning/artificial intelligence problem which boils down to training the AI 'puters with as much data as you can so that out in the real world the autonomous vehicle can say to itself: "Yeah, I've seen this situation before, here's the response that worked best. Both the training and the on-the-road-recall, if they are to be anywhere near efficient, require the fastest chips you can find. Tesla had a whole bunch of data from a few billion miles of actual driving for computers to train on, and, combined with Nvidia's fastest-in-the-world GPU chips, it was a match made in heaven.

Except it wasn't.

The challenge of autonomous driving on open roads alongside non-autonomous vehicles was bigger than anyone in that simple, optimistic time ever envisioned, even in their nightmares. Here's one example about Waymo from a 2017 post:

"When Google was training its self-driving car on the streets of Mountain View, California, the car rounded a corner and  encountered a woman in a wheelchair, waving a broom, chasing a duck. The car hadn’t encountered this before so it stopped and waited."

In May 2015 we were posting " Nvidia Wants to Be the Brains Of Your Autonomous Car (NVDA)" and seven months later the more declarative "Class Act: Nvidia Will Be The Brains Of Your Autonomous Car (NVDA)"

Then in October 2016, what was probably the high-water mark for the relationship "Nvidia Could Make $1B From Tesla's Self-Driving Decree: Analyst (TSLA, NVDA)"

Sadly, the task was just too difficult but Mr. Musk thought it was doable if only he could get even faster chips than Nvidia had on offer:

NVIDIA Partner Tesla Reportedly Developing Chip With AMD (TSLA; NVDA; AMD) 
Today in leveraged WTFs....

"During a talk at a private party, Elon Musk said Tesla is developing specialized AI hardware "'That we think will be the best in the world;" (TSLA)  

"Tesla says it’s dumping Nvidia chips for a homebrew alternative" (TSLA)
The only reason for Tesla to do this is that NVIDIA's chips are general purpose whereas specialized chips are making inroads in stuff like crypto mining (ASICs), Google's Tensor Processing Units (TPUs) for machine learning and Facebook's hardware efforts. 
 
 Watch Out NVIDIA: "Google Details Tensor Chip Powers" (GOOG; NVDA)
We've said NVIDIA probably has a couple year head start but this bears watching, so to speak....

Culminating in August 2018's
"Nvidia CEO is 'more than happy to help' if Tesla's A.I. chip doesn't pan out" (NVDA; TSLA)

And now on to the headliner, from Observer, June 8:
Elon Musk Predicts Nvidia’s Monopoly in A.I. Chips Won’t Last....

And possibly related July 13:
Elon Musk's x.AI Launches
The company was formed in March so it's valuation is probably around a hundred billion or so.

Just kidding. I have no idea what sort of valuation it has been assigned. x.AI is a Nevada corporation which, as our corporate attorney readers well know, is handy as hell for a privately-held stealth company. As part of the company's coming-out I think they dropped the period in the name on the original incorporation papers.

Mr. Musk was one of the founder/funders ($100 million gift not equity) of ChatGPT parent OpenAI when it was a .org (non-profit) and seemed a bit miffed when Sam Alman hooked up with Microsoft to the tune of $10 billion.

So Elon went out and bought a garage-full of GPUs.

Here's a twofer, first up TechCrunch, July 12:....

Tuesday, August 4, 2026

"‘I’d Be Petrified’: Steve Eisman Says Cheap Chinese AI Models Could Wreck OpenAI and Anthropic’s Valuations"

From 24/7 Wall Street, August 4:

“Big Short” investor Steve Eisman said on his own show, Real Eisman Playbook, that “If I was the head of Anthropic or OpenAI, I’d be petrified. That spells to me price war.” The comment lands at an inconvenient moment: both labs have filed confidentially with the SEC and are aiming at public listings near $1 trillion. Eisman literally said “price war.” The valuation-collapse framing in our headline is our inference layered on that quote, since a $1 trillion IPO story assumes pricing power a price war would erode. 

The Moonshot Threat: Kimi K3 and Open Weights 

Eisman’s specific concern is Moonshot AI’s Kimi K3, which he says charges $3 per million input tokens versus $5 for OpenAI’s GPT-5.6 Sol and $10 for Anthropic’s Claude Fable 5. Pricing is only half the story. Moonshot released Kimi K3’s full model weights, so developers can run and customize it independently rather than staying locked to Moonshot’s platform. That undercuts the “stickiness” closed-model economics depend on. If an enterprise buyer can host a comparable model on its own GPUs at a fraction of frontier API pricing, the switching cost justifying premium subscription economics thins every quarter. Eisman made the argument while challenging tech bulls Dan Ives and D.A. Davidson’s Gil Luria on AI moats.

The IPO Stakes

Anthropic filed confidentially with the SEC on June 1, 2026, with OpenAI following shortly after (reporting varies, around early June); both filings remain confidential rather than public S-1s. Anthropic is targeting an October 2026 NASDAQ listing off a $965 billion private valuation, potentially the first company to debut publicly at $1 trillion+. OpenAI has reportedly wavered toward a 2027 listing amid market volatility, with CEO Sam Altman said to have a “hard floor” of a $1 trillion listing price. As of Eisman’s July 29 broadcast, Polymarket traders priced Anthropic’s odds of going public by year-end at ~69%, versus just 19% for OpenAI. Public investors will price the moat directly, which makes Eisman’s price-war framing pointed rather than academic.

China’s Price War Is Already Underway: Baidu

On Bloomberg’s The Asia Trade on August 3, 2026, Bloomberg Intelligence analyst Robert Lee argued the commoditization Eisman fears is already playing out in China. “There’s a high level of commoditization in the AI sector. The sector is overpopulated, flooded with supply. At last count there were 988 large language models officially approved by China,” Lee said. He drew a parallel to solar’s collapse: an oversupplied market where price-cutting is the only lever left. DeepSeek cut API pricing by as much as 50%, and Baidu (NASDAQ:BIDU | BIDU Price Prediction) cut API pricing by 99% earlier in 2026. Baidu’s own numbers show the model shift underneath the price war: AI Cloud Infra revenue rose 79% YoY while Online Marketing Services fell 22% YoY. Lee named Alibaba (NYSE:BABA), Tencent, and Huawei as the best-capitalized survivors. Alibaba backs that up with a Cloud Intelligence Group accelerating 40% externally and Qwen’s open-source family surpassing 1 billion cumulative Hugging Face downloads, per its Q4 FY26 6-K filing.

The Bull Rebuttal: Alphabet and Real Revenue...

....MUCH MORE 

Related:

"Apollo's Sløk: The market faces big risks if hyperscalers' AI profits get delayed"

Here's Apollo, July 9: 

A Slower AI Payoff Would Be Everyone's Problem

This point is key (bolding in original):  

If Chinese models keep gaining and token prices keep falling, the hyperscaler cash flows expected may prove too optimistic.

If interested see also July 7's ""Frontiers of compute: The technologies to reduce AI inference costs"—McKinsey

The cost of inference has dropped by over 99.5% in the last three or four years while the price to the end user definitely has not fallen by that much and in fact all-in costs have actually risen. That gap is the opportunity China is focused on.

More VentureBeat On DeepSeek: "DeepSeek R1’s bold bet on reinforcement learning: How it outpaced OpenAI at 3% of the cost"

 AI: "A brief history of Sam Altman’s hype" (MIT Technology Review's Hype Correction series)

 "OpenAI Considers Drastic Price Cuts, Anticipating War for Users With Anthropic"

 SoftBank Stock Plunges On Possible OpenAI IPO Delay (9984:Tokyo)

 Not Good - "Nvidia in Talks With OpenAI to Guarantee $250 Billion Financing for Data Center"

SemiAnalysis On Moonshot AI's Kimi K3: Probably Good For Nvidia and HBM; Not So Much For Open AI

Sunday, July 19, 2026

SemiAnalysis On Moonshot AI's Kimi K3: Probably Good For Nvidia and HBM; Not So Much For Open AI

Via Xitter: 

Thread at Thread Reader

Ending with a Jevons Paradox Cambrian Explosion, to mix a couple metaphors.

...More efficient attention will further push context lengths from 1M to 5M+, with less context rot. Jevons’ Paradox means that making attention more efficient will lead to wider AI adoption, which will require more networking. 8/8 

In essence the model is so big it will only be optimally run on Nvidia's premier offerings+High Bandwidth Memory+NVLink - Mellenox networks.

Here's their commentary:

[Emergency Episode] Moonshot’s Kimi K3 has Arrived! China has a Frontier Model 

Monday, July 13, 2026

Chips: "TSMC, the world’s largest contract chipmaker, reports 68% surge in June revenue"

From CNBC, July 13:

  • TSMC reported a 6.2% month-on-month and 68% year-on-year jump in June revenue.
  • The company reported a first-half revenue of NT$ 2404.48 billion for 2026, marking a 35.6% increase compared to the same period last year. 
  • The Taiwanese chip giant’s stock rose 1%.

Taiwan Semiconductor Manufacturing Co. reported a 67.9% year-on-year rise in its June sales on Monday, ahead of its second-quarter earnings release later this week.

For the first half of 2026, TSMC’s total revenue reached 2.4 trillion new Taiwan dollars ($74.99 billion), representing a 35.6% increase compared to the same period in 2025. TSMC reported June revenue of NT$ 442.68 billion — a 6.2% increase from the previous month.

The Taiwanese chip giant’s shares rose 1% Monday.

TSMC’s numbers are “quite robust,” said Sravan Kundojjala, an analyst at SemiAnalysis, noting that the chipmaker’s second-quarter revenue exceeded its high-end guidance of $40.2 billion. The result came whereas historically June revenue has declined month-over-month over the past four years, he added. 

“The demand supply situation in AI is still quite tight and TSMC is sold out on N3, which is targeted by all leading AI GPU and CPUs this year,” he added.

The world’s largest contract chipmaker manufactures semiconductors for a wide range of applications, spanning from smartphones to high-performance AI computing systems, with key clients including U.S. technology leaders such as AI darling Nvidia, Apple and Advanced Micro Devices....

....MORE 

And from the company:

TSMC June 2026 Revenue Report 

https://pr.tsmc.com/system/files/news/6a0a7ebdb18210e4d6c9d01676a10f09e0d2281a/June%202026%20%28E%29.jpg 

....MORE 

Wednesday, July 8, 2026

SemiAnalysis: Anthropic May See A 3Q26 Profit Over $1Billion

They grow up so fast. 

From SemiAnalysis, July 7:

Anthropic 3Q26 Profit Over $1B: The Anthropic IPO Financials Sneak Peak
Anthropic’s Opportunity is Theirs to Lose 

Introduction 

When Dario Amodei left OpenAI to start Anthropic in early 2021, the viral release of ChatGPT was over 18 months away and the commercialization of LLMs was practically zero. Just a few short years later, Anthropic and OpenAI combine for ~$100B of ARR and a clear winner emerged in the profitable monetization of AI models in 2026 as Claude Code took the software development world by storm.

Anthropic confidentially filed for IPO on June 1st. Over 1 month later, equity raises from hyperscalers loom, and a reported OpenAI push out of their own IPO until 2027 has led some to question the ability of the labs to raise. However, Anthropic is the clear clubhouse leader in capturing the B2B market today and is doing so in a profitable manner against an unfocused and money-burning competitor.

With this lead, we expect Anthropic to take advantage of their superior business model and margins to invest in further in new models that help extend their lead and monetization over closed and open source competitors. Anthropic has the ability to truly make OpenAI dance and we see Anthropic as the first $6T company as a base-case possibility if they continue to execute. Pricing power, gross margins, business model, and profitability are all reasons for Anthropic to IPO first and put the impetus on OpenAI to open their financials and raise the necessary capital to compete and fund the massive AI buildout still to come.

We’ve already seen 2 AI Labs IPO this year (Zhipu and Minimax from China), but Anthropic would be the first AI lab of this scale to do so. A confidential filing means there are no public numbers disclosed. Fortunately, the Tokenomics team at SemiAnalysis works to build the financials from the bottom-up by SKU, tier, and customer type. Recently, a WSJ article on Anthropic’s financials confirmed the accuracy of the work our Tokenomics team does across labs and hyperscalers to help investors, corporates, and other stakeholders understand the economics and financials of the AI Ecosystem....

   

Source: WSJ, SemiAnalysis Tokenomics Model 

....MORE (paywall) 

HT: Yahoo Finance video, July 8.

Next month we'll be featuring our "Quality of Earnings" module wherein we look at some of the perils that can deceive the unwary junior account.

Monday, July 6, 2026

AI: A Troubling Trend Emerges

“Mr Bond, they have a saying in Chicago: 'Once is happenstance.
Twice is coincidence. The third time it's enemy action'.”

—Auric Goldfinger, Goldfinger

First up, the New York Times, June 5:

SpaceX Has $30 Billion Deal to Provide Google With A.I. Computing Power
Elon Musk’s rocket company said Google would pay it $920 million a month, as it prepared for its initial public offering. 

And from The Register, July 2:

SoftBank enters the rent-a-GPU race as America looks for support for AI training
Japanese giant needs to find some use for that 10 GW US server farm it is building 

SoftBank is set to get into the neocloud business in America, providing resources to hyperscalers and other customers seeking a platform on which to carry out their AI training.

The Japan-based tech investment giant says it will establish a new company called SB Neo, Inc to operate its neocloud business in the US, and expects to start operations in fiscal 2027 (ending March 31, 2028).

In actual fact, ownership of the nascent biz will be split, with 51 percent in the hands of SoftBank Corp, while 49 percent is owned by SoftBank Group Corp. SB Neo will be a consolidated subsidiary of SoftBank Corp, which is itself 40 percent owned by SoftBank Group Corp. All perfectly clear?

And from SemiAnalysis, also July 2:

Meta Compute: Everyone Wants To Be A Neocloud

....With Bloomberg headlines suggesting Meta could become a Neocloud, the market’s reaction was immediate: aggressive sell-off of Neoclouds like Coreweave & Nebius, and debates of “overcapacity” coming back. Let’s set the record straight – we believe that both takes are erroneous and that Meta’s datacenter & compute procurement will accelerate, not slow down. Capex in 2027 will be shockingly high. In just the first six months of the year, Meta has contracted over 5GW of capacity across Cloud & Colo, and that doesn’t even include all their accelerating self-build activity. Everything is computer and everything is a neocloud....

I don't like disagreeing with SemiAnalysis so I'll just say this move among big, big companies to lease out giant chunks of their compute, and axiomatically taking the leases onto their balance sheets as an asset is enough to drive me back to Graham & Dodd, and in the case of SoftBank to Howard Schilit's Financial Shenanigans: How to Detect Accounting Gimmicks & Fraud in Financial Reports.

As we await developments a flashback to June 2019:

"Something is not quite right with SoftBank"
Additionally, SoftBank investee WeWork was out looking for  a few billion dollar line of credit.
Shades of another disruptor, Sam Insull, leverage at the holding company level, leverage at the operating company level, leverage all the way down...
Sam Insull's story is interesting and possibly instructive. He was a Chicago guy.

Friday, July 3, 2026

"US Grid Constraints: Towards 40GW+ of Behind-The-Meter Datacenter by 2028?"

Filed under "Things that could slow the realization of the Master Plan For World Domination."

From SemiAnalysis, June 25:

Why the Grid Can't Keep Up, and Why that Drives Behind-The-Meter 50%+ of DCs/Year By 2028 

Today, the US grid is serving most datacenter load in the US, but we’re reaching a tipping point. As the insatiable demand for power of AI Labs and hyperscalers keeps accelerating, the grid simply can’t add capacity fast enough. That leaves Behind-The-Meter as the only way for the largest players to secure the power they need. Nearly a year ago, our Onsite Gas deep dive was the first to predict the fast rise of new entrants in the BTM gas equipment market. Since then, companies like Bloom Energy, Bergen Engines, Wärtsilä and many others have been remarkably successful. Overcoming GEV and Siemens turbine capacity constraints proved far easier than many had feared.

Today, we go deeper and model US Grid capacity to understand the shortfall that must be filled by Behind-The-Meter solutions for datacenters.

Let’s start with key numbers: first, we continue to see a record datacenter buildout in the US, going from +21GW in 2026 to +84GW by 2030. We explained in detail last week why Datacenter Delays headlines are often overblown.

Our research suggests that BTM will power well over half of new US datacenters in 2028+, and the Total Addressable Market (TAM) for DC BTM equipment to cross 50GW/year by 2029. New Grid Capacity isn’t growing fast enough, and also needs to serve non-datacenter load growth.

 

Source: SemiAnalysis Energy Model, SemiAnalysis Datacenter Model

The chart above shows the three core building blocks of our forecast: Expected Datacenter US Gross Power Demand, available US Grid Capacity, and New Grid Supply. We use the best of SemiAnalysis industry-leading insights to build this forecast.

The first block, datacenter demand, comes from a bottom-up forecast powered by a building-by-building model, supported by chip-by-chip AI demand forecast of the Accelerator Model, and validated by our Tokenomics Model which tracks the economics of the buildout and answers the “bubble” question.

The second building block of our Energy Model, grid headroom, analyzes supply & demand dynamics in each major part of the US grid. Our model follows the methodologies of all ISOs & RTOs and models UCAP/ICAP reserves, supply & demand growth, reliability risks, etc.

The third block forecasts new grid supply, through a bottom-up forecast produced by our new Energy Model. We track 40,000 generation assets in the US and forecast quarter by quarter Commercial Operation Date (COD) for all fuel types. We then estimate the “true” capacity value of power plants via our proprietary ELCC model, adapting to the specifics of each ISO and major non-ISO region.

Our forecast points to barely 15GW of net-new ELCC capacity being added annually, with a rising trend towards 20GW+ by the end of the decade. This is effectively all the firm capacity being added to the system that a grid operator can recognize to serve firm datacenter load - as well as other firm load like industrial plants, semiconductor fabs, etc.

 

 Source: SemiAnalysis Energy Model

Netting that accredited supply against peak demand and required reserve margins is what yields headroom itself, the spare accredited capacity a market has left to host new load after covering its own peak demand and required reserve margin. On this basis, available headroom is already approaching zero and turns negative by 2027, based on our analysis of required reserve margins across the country....

....MUCH MORE 

Thursday, June 4, 2026

Data Centers In Space? Not Until They're Mandated

From SemiAnalysis, June 3:

To Boldly Go: The Case for Space Datacenters
Space DC Total Cost of Ownership Explained. Unpacking constraints from Terrestrial DCs and Chip Production. Space-Earth Parity in the late 2030s, Space DCs could start to be viable even sooner. 

Everyone has been talking about datacenters in space. Interviews given by Elon Musk in the past few months have spent lots of time on orbital compute:

“Five years from now, my prediction is we will launch and be operating every year more AI in space than the cumulative total on Earth... I would expect to be at least, sort of five years from now, a few hundred gigawatts per year of AI in space and rising.”
- Elon Musk on Dwarkesh Podcast, February 2026

Furthering space-based compute was also one of the stated motivations behind the merger of xAI into SpaceX (as a ‘reorganization of entities under common control’), and is a key part of SpaceX’s plans to go public, as stated in their S-1 filing on 20 May 2026.

“Our goal over time is to launch 100 gigawatts of compute to space each year. If operated continuously, the generation resources used to support 100 gigawatts of compute could generate approximately one-fifth of the annual power production in the United States, which was 4.4 thousand terawatt hours in 2025… We expect space‑based compute to massively increase AI compute scale, while also improving token economics.”
- SpaceX, S-1 Filing, May 2026

As expected, many part-time prognosticators in the Substack-verse have emerged from the woodwork to weigh in on the concept. Some articles bring up insightful points, but there are more than a few that are built upon ideas that fly in the face of science.

A few casual arguments made in favor of space datacenters include the following:

  1. Space can provide free solar energy 24 hours a day

  2. Cooling is “free”. Some erroneously point to space being cold as a key positive

  3. Communications latency in space is low as you’re just sending light through a vacuum

  4. There is no need for permitting in space… so far…

Many of these points sound like they hold merit on the surface, but a deeper analysis of each apparent advantage reveals a far more complex story.

While we think that it is possible that space datacenters could scale one day, deploying orbital compute using today’s technology currently costs several times more than deploying terrestrial compute. Achieving Space-Earth cost parity will require significant engineering work, material science breakthroughs and cost scaling progresses and will still take years to achieve. There are also important reliability and servicing obstacles to overcome - for instance - how GPU servers will recover from faults that require human intervention, effectively shielding accelerators from radiation, among many others.

When we deploy compute in space, it won’t be because of the four superficial reasons we have cherry-picked above. Rather, Space-based datacenters make sense in the world where AI demand well exceeds all of the four layers of terrestrial datacenter supply that we will introduce below. For Space datacenters to step up to this call - it is a necessary condition that major space datacenter cost items like radiators, solar arrays and launch costs decline considerably, and that a number of key operational obstacles are overcome.

Users of our AI Space Datacenter TCO Model can see a first-principles, system-level framework for evaluating orbital compute economics, engineering constraints, and supply-demand dynamics across both terrestrial and space-based infrastructure.

The four layers of incremental power supply for terrestrial datacenters include:

  1. Grid-connected supply,

  2. Converted bitcoin miners and powered land,

  3. Behind the meter generation, and finally,

  4. Industrial capacity and manpower to build further power infrastructure.

A necessary condition for AI related IT equipment demand to reach levels exceeding terrestrial datacenter supply is for there to be enough chip fabrication capacity to fulfill this demand in the first place, before we even discuss datacenters! We wrote about this in great detail in our recent article on the Great AI Silicon Shortage, where we concluded that the industry has moved from a power-constrained to an accelerator-constrained regime. Available datacenter capacity and power now exceed AI compute demand, but TSMC’s N3 wafer capacity and HBM supply cannot keep pace with the pace of accelerator deployments. This means that today, and for the next few years, chip manufacturing will be the global constraint before we even worry about supply for these four layers.

The chip constraint forms a separate fifth layer of supply - Semiconductor Production, and it is a “universal” constraint on all chip deployment, whether deployed on Earth or in Space. Users of our AI Space Datacenter TCO Model can see how this constraint applies well into the future, and under what scenarios regarding chip manufacturing capacity addition that Semiconductor Production may not be the constraint.

Elon Musk is clearly well aware of this constraint, and it is the impetus behind his Terafab Initiative. The AI Space Datacenter TCO Model also includes knobs and sliders for users to tune to test out various Terafab scenarios.

Framing the Space Datacenter Debate 
Our various industry models such as the Accelerator Model, the Foundry Industry Model and WFE Models illustrate the aforementioned chip tightness. Meanwhile our AI Datacenter Model forecasts accelerating incremental datacenter additions in 2027 and 2028. Thus, datacenter capacity addition will run ahead of chip constraints in the next few years until fab capacity additions accelerate to catch up. Our suite of industry models will only forecast such wafer fab and datacenter capacity additions once such plans are confirmed.

However, the world in which AI demand is so overwhelming as to exceed the already formidable datacenter capacity additions is a world with no time for half measures. As such, our AI Space Datacenter TCO Model base case departs from our industry models to reflect this world, assuming accelerating incremental datacenter capacity additions and a meaningful step up in the pace of chip fab capacity addition. It is a world where all the stops are pulled out and many obstacles from gas turbine availability to EUV tool production constraints are overcome because clear long-term AI end use ROI justifies enough capital investment to overcome them.

The below chart illustrates what this world could look like - with incremental datacenter capacity additions eventually in the hundreds of GW annually, though adding chip capacity will still be more difficult than adding datacenter capacity....

....MUCH MORE 

Most recently from SemiAnalysis:

May 28 -  Powering Data Centers: "Inside the 800VDC Revolution"

March 18 - Memory: Shortage Could Last Five Years, It's The Wafers

February 21 - "CPUs are Back: The Datacenter CPU Landscape in 2026"

February 10 - "Memory Mania: How a Once-in-Four-Decades Shortage Is Fueling a Memory Boom "

And many more

Tuesday, May 26, 2026

Powering Data Centers: "Inside the 800VDC Revolution"

From SemiAnalysis, May 26:

Four-Phase 800VDC Transition, Power Rack Economics, SST, Equipment Content/MW Build, Supplier Implications 

We’d like to thank DG Matrix, Novos Power, and Aran Industries for their contributions and insights during the preparation of this deep dive.

 Introduction: Welcome to the Power Chain Roller Coaster

Across every major industry conference in the first half of 2026, our research team kept walking past the same scene: a booth ten or fifteen people deep, leaning in to catch every word from another datacenter equipment messiah preaching the gospel of 800VDC. The pitch was the same every time. 800VDC is about to change the electrical infrastructure of the datacenter.

Every architectural shift looked excessive at first. Operators spent decades keeping water and leaks out of the data hall, then GPU thermal density made running coolant right up against the precious silicon unavoidable. Each shift happened anyway, because physics and the economics of compute do not negotiate. 800VDC is next, and the logic is the same. Tokens per watt are what matters.

Source: Nvidia, InferenceX

As GPU clusters become increasingly dense, with Kyber Ultra approaching 660kW per rack, the physics start to break down. Resistive losses scale with current squared, and at these power levels copper mass and thermal envelope exceed what fits inside a rack. Moving to 800VDC eliminates conversion stages, reduces resistive losses, and cuts facility-level power consumption by ~5%. At 1GW of IT load, that is over 50MW of continuous savings, tens of millions in annual electricity costs, or new compute capacity unlocked. For all the inference-king proponents out there, 800VDC is a transition forced by physics and motivated by system economics.

We have been tracking this transition through our InferenceX and Industrials Models, which provide a bottom-up view of where efficiency gains materialize and which equipment categories absorb the disruption. The Industrials Model includes a dedicated 800VDC module, building up from individual accelerator architectures to a top-down view of 800VDC penetration, MW adoption, and market sizing for equipment like the power sidecar and Solid-State Transformers (SSTs).

Source: SemiAnalysis Industrials Model

This deep dive traces the transition phase by phase: from the sidecar retrofit, through faciliy-level DC distribution, to the SST endgame. For each phase, we analyze the BoM and map the changes in equipment content/MW, what survives, what gets redesigned, and what gets eliminated.

The 800VDC revolution is set to dramatically change the revenue trajectory of certain suppliers. We’ve been tracking winners and losers for over a year in Industrials Model, which estimates the BoM for 20+ different datacenter designs broken down into 70+ equipment types and lays out the impact for 500+ suppliers. It is built on our industry-leading Datacenter Model which forecasts quarter-by-quarter MWs for 6000+ datacenters and anticipates design changes.

This has enabled us to successfully call out both winners, and companies inaccurately pictured as losers by the market, before anyone else. If you are wondering whether UPS systems have a place in upcoming 800VDC distribution, what is the market opportunity for SSTs, or which suppliers are leading this transition, stick with us.

Source: SemiAnalysis Industrials Model

Part 1 of this 800VDC Revolution series covers datacenter layout and equipment implications. Part 2 will focus on power electronics and the semiconductor revolution underneath it.

 Understanding The Basics: What is 800VDC and Why It’s Inevitable

At its simplest, 800VDC in this context means distributing power at ~800 volts direct current through the data hall or row and into the rack, then stepping it down near the compute. The number 800 is not arbitrary, but a voltage high enough to materially reduce current (and therefore copper loss and thermal burden) while remaining within the broad regulatory and product-safety classification of “low-voltage DC” in many jurisdictions. For context, EU rules around the Low Voltage Directive scope reference DC equipment ratings up to 1,500 V DC (and AC up to 1,000 V).

Current datacenter electrical architectures generally rely on AC distribution at the facility level. Datacenters today use three-phase AC at 415V or 480V, and the topology relies on conventional UPS architectures before distributing 48-54V DC within the rack.

This works at today’s rack power levels, but starts to fail as rack densities in the next two years approach ~600 kW+, for several reasons:

  • Copper becomes unmanageable at 48–54 V. A 1 MW rack at 48–54 VDC needs ~200 kg of copper busbars. At 1 GW scale, that’s hundreds of tons of copper — brutal on cost, weight, installation complexity, and routing space.

Source: Microsoft

  • Power shelves crowd out compute. Today’s NVL72 racks already use up to 8 power shelves. At Kyber-class rack power, a 48–54V approach would require ~64U-equivalent of power hardware, effectiviely an entire rack, leaving no volume for compute.

  • Current becomes the real limiter. Delivering 600 kW at 48–54 V implies ~12,500A. At 800 V, that drops to ~750 A (~16.7× less), enabling dramatically smaller conductors/busbars and far lower thermal stress. If conductor resistance were held constant, I²R losses fall ~278×, so in practice you shrink copper and “buy” size/weight reductions.

  • Conversion losses compound and hurt reliability. Stacked AC-to-DC and DC-to-DC stages reduce end-to-end efficiency, increase heat, and introduce failure points, raising cooling loads, downtime risk, and maintenance costs.

At the end of the day, 800VDC is the physics enabler for 2,300W TDP chips and 600kW racks, and those 600kW racks are the direct consequence of the push for density, because density is what drives cost per token down. Cost per token is dictated by the size of the scale-up world you can build at full NVLink bandwidth: bigger domains mean wider Expert Parallelism (EP) / Tensor Parallelism (TP), MoE routing on NVLink rather than scale-out, and less serialization across decode. As we laid out in our Vera Rubin Deep Dive and GTC 2026 pieces, Nvidia’s design rule is to pack compute tightly enough that copper reaches everything in the rack. Reiner Pope made this point cleanly on our friend Dwarkesh’s podcast a few weeks ago, indicating that a single rack bounds the size of the expert layer you can build, because the moment an all-to-all crosses a rack boundary, it falls onto a scale-out fabric that is roughly eight times slower than NVLink.

Bigger scale-up worlds mean denser racks, denser racks mean 600kW envelopes, and 800VDC is what makes those envelopes possible.

Source: SemiAnalysis AI Networking Model

 The Four Chapters of the HVDC Transition

The move to 800VDC is a complex metamorphosis that rewrites the entire electrical architecture, introduces new safety standards, requires new regulatory frameworks, and, most importantly, forces operators to make very different strategic choices about when to walk away from their legacy AC distribution.

Source: SemiAnalysis

We frame the 800VDC transition as progressing through four distinct phases. Phases 1 and 2, starting in late 2026 / early 2027, retrofit the existing AC distribution into 800VDC at the rack level via the power rack. Phase 1 is the early-mover stage, driven by hyperscalers willing to pay up for future-proofing and efficiency gains. Phase 2 kicks in once 800VDC-native systems begin shipping at volume. Phase 3 rewrites the electrical architecture itself, taking 800VDC distribution facility-wide. Phase 4 is the end state, built around new pieces of equipment that promise to render much of today’s electrical stack obsolete....

....MUCH MORE 

Previously:

March 27 - Opportunity: "Data Centers Are Transitioning From AC to DC 800-volt DC power delivery: will enable next-gen AI data centers"

Wednesday, March 18, 2026

Memory: Shortage Could Last Five Years, It's The Wafers

One of the first sources to relay the concern that the memory chip shortage would not be over quickly was Tom's Hardware in October 2025, linked here on the blog on January 3:

"AI data centers are swallowing the world's memory and storage supply, setting the stage for a pricing apocalypse that could last a decade" 

Now we return to Tom's to see what the head honcho at #2 memory chip maker SK Hynix has to say, March 18:

SK Group chairman says memory chip shortage will last until 2030 — wafer supply trails demand by 20% 

SK Hynix's CEO is expected to announce price stabilization measures soon. 

SK Group chairman Chey Tae-won told reporters at Nvidia's GTC conference in San Jose on Monday that the global memory chip shortage is likely to persist for another four to five years, with industry-wide wafer supply lagging demand by more than 20%, Bloomberg reported. Chey, whose conglomerate controls SK Hynix, said leading memory makers are expanding capacity but are unlikely to fully meet demand until around 2030 because securing additional wafers takes at least four to five years, according to The Korea Times

SK Hynix holds roughly 57% of the global HBM market and 32% of overall DRAM, and the company is currently building a $13 billion HBM packaging and testing facility at its Cheongju complex in South Korea, with construction scheduled to begin next month and completion targeted for the end of 2027.

Article continues below...

....MUCH MORE  

And it's not just silicon wafers for memory chips.

From SemiAnalysis, March 12:

The Great AI Silicon Shortage
TSMC N3 Wafer Shortages, Memory Constraints, Datacenter Bottlenecks, Supply Chain Wars Winner 

Token demand is skyrocketing and the need for AI compute continues to accelerate. The improvement in model capabilities combined with the rapid emergence of agentic workflows has driven a surge in user adoption and aggregate token demand. Anthropic added a staggering $6B of ARR in the single month of February alone driven by broad adoption of agentic coding platform Claude Code, and if Anthropic had more compute they would have added more. Despite a huge AI infrastructure buildout over the past few years, available compute is scarce. On-demand GPU prices continue to go up even for Hoppers which are almost 2 generations old.

From our own experiences, we have reached out to every neocloud we know asking if they have small clusters available, but everything is already firmly locked up. This tight supply environment explains the sharp reset in hyperscaler capex plans. Consensus estimates have moved materially higher across the board, with Google standing out as the most extreme example, where 2026 capex expectations have roughly doubled versus prior expectations, primarily driven by datacenter and server spend.

 

Source: Company Earnings, Bloomberg

This is a tremendous level of spending, and hyperscalers would deploy even more capital if they could, but they are constrained by one critical factor: silicon supply. There is simply not enough advanced logic and memory fabrication capacity to support the pace of compute deployments. While the AD (After Da launch of ChatGPT) era has been riddled with various constraints such as CoWoS packaging and datacenter power, we are now firmly in the silicon shortage phase.

 

Source: SemiAnalysis Accelerator Model

The TSMC N3 Shortage

One of, if not the, biggest constraints is TSMC’s N3 logic wafer capacity. TSMC’s N3 family started shipping for revenue in 2023, with demand initially driven primarily by smartphones and PCs. N3 got off to a shaky start, with the first variant “N3B” having yield issues and being too expensive relative to the density improvement. Greater adoption came with the refined N3E process, a relaxed variant with far fewer EUV layers and therefore lower cost. Key smartphone and PC customers include Apple, which uses N3 variants for its M3 to M5 Mac chips and A17 to A19 iPhone processors, Qualcomm for its Snapdragon 8 Elite series, MediaTek for its Dimensity smartphone SoCs as well as select automotive and PC chips, and Intel for its Lunar Lake and Arrow Lake client processors.

 

Source: SemiAnalysis Foundry Model

Up until today, N3 demand has been driven primarily by consumer electronics. In 2026, all the main AI accelerator families are transitioning to N3, and AI will account for the majority of N3 demand before transitioning to N2 and beyond.

We can see in the table below the industry-wide convergence toward TSMC’s N3 family as the leading process node for AI accelerators heading into 2026. NVIDIA transitions from 4NP with Blackwell to 3NP with Rubin. AMD, typically the earlier adopter of new nodes, has already adopted N3 for MI350X and will stay on N3 for the AID and MID tiles for MI400 (XCD is N2). Google’s TPU roadmap shifts fully to N3E starting with TPU v7, with TPU seeing a huge upsize in program volumes this year. AWS also transitions to N3P with Trainium3. Meta’s MTIA follows a similar path, though it will be at much lower volumes.

 

Source: SemiAnalysis Accelerator Model

This shift is not limited to XPU silicon. The Vera CPU used in VR racks uses N3P for all its silicon. There is also networking silicon in the form of the NVLink 6 switch, as well as scale out switches like Tomahawk 6 and Spectrum 6. With Rubin offering 1.6T of scale out network per GPU, Rubin kicks off the adoption of 3nm 200G optical DSPs.

This sudden convergence of N3 adoption coupled with the continued growth of AI compute demand has resulted in a huge demand shock for N3 wafer capacity. TSMC has been caught flat-footed, with wafer capacity expansion failing to keep pace with surging AI demand. How did this happen? Although the greatest compute buildout in history began in late 2022, TSMC’s capex only exceeded its previous peak in 2025. This year, TSMC is going to smash through last year’s record Capex, because they have realized how far customer demand is exceeding their capacity....

....MUCH MORE 

It's a pretty big deal. If interested see:

Memory: "The inflation spark that could become a deflation shock?"

From M&G's Bond Vigilantes, February 27:
Memory chips have quietly become the most important commodities in the global economy.

 “'Entry-Level PC Segment Will Disappear by 2028,' Says Gartner, as Soaring Memory Costs Start to Cripple Manufacturers"

"....Micron, SanDisk & Memory Stocks Are Crashing Today"

 Thanks for the Memories: "South Korea’s Kospi plunges 12% amid broader declines in Asia markets as Iran conflict rages"
The index, which has been driven by the memory chip makers, Samsung Electronics Co. and SK Hynix Inc. et al., up over 145% from March 2025 to the February 25, 2026 peak is now down 10% on the day, March 4th....