# The Entire AI Data Center Explained — From Electricity to ChatGPT

- Source: https://www.youtube.com/watch?v=ckoi0RTEgcY (YouTube)
- Creator: Leo Cui, Ph.D., CFA
- Published: 2026-07-23T17:17:04.000Z
- Transcribed by Memora: 2026-08-07T05:01:51.009Z
- Canonical page: https://dailymemora.com/youtube/90

> Transcript and summary produced by Memora from the publicly available
> video linked above. The original video belongs to its creator.

## Summary

### Four hyperscalers spending $725B/year — more than the US interstate highway system

The AI infrastructure boom is led by **Amazon, Microsoft, Google, and Meta**, collectively spending ~$725B annually on data centers, surpassing the cost of the US interstate highway system. This marks a shift where **Scaling Laws** (2020 OpenAI discovery) turned AI from a research problem into a capital expenditure problem, making software 'heavy industry.' The user query process is traced: from phone through fiber, API gateway, tokenization, prefill (parallel prompt processing), decode (serial token generation), and back.

### Power density jumps 120–600 kW per rack, fueling a behind-the-meter gold rush

Massive power density increases (120–600 kW per rack) strain grid capacity, causing **4–5 year interconnection queues** and transformer shortages. Companies bypass the grid with behind-the-meter solutions:
- **Nuclear restart**: Three Mile Island for Microsoft
- **Small modular reactors**: speculative
- **Fuel cells**: Bloom Energy (controversial)
- **Gas turbines**: GE Vernova, Siemens Energy

Electrical equipment makers (Vertiv, Schneider, Eaton) and backup generator firms (Caterpillar, Cummins) benefit, while political risks emerge from residential bill hikes. Heat density forces a shift from air to **liquid cooling** (direct-to-chip, immersion) when racks exceed 30–50 kW.

### Nvidia dominates accelerators, but Broadcom leads custom chips and networking

Nvidia holds **~80-86% AI accelerator market share** with 75% margins, powered by CUDA software lock-in. Competition:
- **AMD Instinct GPUs**: 5-7% share, software gap
- **Intel**: marginal
- **Broadcom**: custom chip designer for hyperscalers, 60%+ market, 106% AI revenue growth

Rack assemblers (Super Micro, Dell, HPE, ODMs) capture only 6-10% margins. Networking is critical: scale-up via Nvidia's NVLink, scale-out war between **InfiniBand and open Ethernet** (now two-thirds of new clusters). Key players: Broadcom (Tomahawk switches), Arista (~6% margins), Cisco, Marvell, Astera Labs (76% margins). Optical transceivers convert electricity to light: Coherent, Lumentum, Innolight, Fabrinet, Corning, Amphenol.

### Memory bottlenecks and the software stack that cuts costs 10×

HBM (high-bandwidth memory) stacks vertically next to the GPU, boosting bandwidth **5–6× at 5–6× the cost**, dominated by SK Hynix, Samsung, Micron. Storage hierarchy from on-chip cache to spinning hard drives, resurrected by AI's endless data output. The software stack: **CUDA** creates a 20-year moat via developer lock-in; serving engines (**vLLM, TensorRT**) cut token costs 3–10× through batching, caching, quantization. RAG extends models with enterprise documents.

### $7 trillion machine funded by vendor debt, resting on a bet that user revenue outgrows costs

Hyperscalers, neoclouds, and AI labs spend **$725B/year**, with vendor financing recycling capital. The only fresh money comes from end users and enterprises, yet the two largest labs’ revenue (**$70B run rate**) remains dwarfed by spending. The entire infrastructure is a bet that user revenue will outgrow costs — no one knows whether it will be history's most important or most expensive endeavor.

## Key points

- Four hyperscalers are spending ~$725B annually on AI data centers, exceeding the cost of the US interstate highway system.
- AI data centers face a power and cooling crisis: racks now draw 120–600 kW, forcing companies to restart nuclear plants or deploy fuel cells and gas turbines to bypass the grid.
- Nvidia holds 80-86% accelerator market share with 75% margins, but Broadcom leads in custom chip design and networking switches.
- HBM memory stacks alongside GPUs deliver 5-6× bandwidth at 5-6× cost, while software serving engines like vLLM and TensorRT cut token costs 3-10×.
- The entire $7 trillion infrastructure bet relies on end-user revenue eventually exceeding the massive capital spending — a question left open.

## Chapters

- [0:00](https://dailymemora.com/youtube/90?t=0) Why AI?
- [6:44](https://dailymemora.com/youtube/90?t=404) The Build
- [21:07](https://dailymemora.com/youtube/90?t=1267) Hardware
- [33:39](https://dailymemora.com/youtube/90?t=2019) Economics

## Transcript

**[0:00]** Last night, sometime around

**[0:01]** 7 PM, You pulled out your

**[0:03]** phone, you typed a question.

**[0:04]** Maybe it was, "What should I make

**[0:06]** for dinner with chicken and rice?"

**[0:08]** And about two seconds later,

**[0:09]** A machine wrote you an answer.

**[0:11]** Two seconds.

**[0:12]** That's what I want to do in this video.

**[0:14]** I want to slow those two

**[0:15]** seconds down, way down.

**[0:16]** Because in those two seconds, your

**[0:18]** questions left your phone, traveled

**[0:20]** hundreds of miles through strands of

**[0:22]** glass thinner than a human hair, and

**[0:24]** arrived at a building the size of several

**[0:25]** football fields, a building that drinks

**[0:27]** as much electricity as a small city.

**[0:29]** Inside that building, your question

**[0:31]** passed through a machine that costs

**[0:33]** as much as a house, got translated

**[0:35]** into pure math, was processed by

**[0:37]** chips running so hot they have to be

**[0:39]** liquid cooled like a race car engine.

**[0:41]** And then the answer came back to

**[0:43]** you letter by letter before you

**[0:44]** had time to lower your thumb.

**[0:46]** and here's the part that should

**[0:47]** get your attention as an investor.

**[0:49]** To make those two seconds possible,

**[0:51]** the largest companies on Earth are

**[0:53]** spending this year alone roughly

**[0:55]** seven hundred and twenty-five

**[0:56]** billion dollars on infrastructure.

**[0:58]** That's just four companies, Amazon,

**[1:00]** Microsoft, Google, and Meta.

**[1:02]** That's more in one year than the

**[1:04]** inflation adjusted cost of the

**[1:05]** entire US interstate highway system.

**[1:08]** Goldman Sachs projects a total build out at

**[1:10]** seven point six trillion dollars between

**[1:12]** twenty twenty-six and twenty thirty-one.

**[1:14]** Jensen Huang, the CEO of Nvidia,

**[1:16]** stood on stage at Davos this

**[1:18]** January and called it his own

**[1:20]** words, "The largest infrastructure

**[1:22]** build out in human history."

**[1:23]** So the question this whole

**[1:24]** video hangs on is simple: where

**[1:26]** does all the money actually go?

**[1:28]** By the end of this video, you're

**[1:29]** going to be able to answer that.

**[1:31]** you'll understand every single step

**[1:32]** your question takes, the power plants,

**[1:34]** the cooling systems, the chips, the

**[1:36]** memory, the fiber optics, and software.

**[1:39]** You'll know which companies sit at every

**[1:40]** step, which companies are printing money,

**[1:43]** which companies are telling stories,

**[1:44]** and where the whole thing could crack

**[1:46]** I'm Leo, a VC investor.

**[1:48]** This is educational purposes,

**[1:50]** not financial advice

**[1:51]** So before we trace your question

**[1:52]** across the country, we have to

**[1:54]** answer something more basic.

**[1:55]** The internet has existed for thirty years.

**[1:58]** Google has answered

**[1:59]** trillions of questions.

**[2:00]** Why did nobody need to spend

**[2:01]** three-quarters of a trillion

**[2:02]** dollars a year until now?

**[2:04]** What changed?

**[2:05]** The answer comes down to a

**[2:06]** difference between two people,

**[2:08]** A librarian and a writer.

**[2:10]** Google is a librarian.

**[2:11]** When you search chicken rice recipe,

**[2:14]** Google doesn't cook anything.

**[2:15]** It walks into a giant library

**[2:17]** it has already organized.

**[2:18]** It indexed the whole internet years

**[2:20]** ago and keeps updating it, and it

**[2:22]** hands you pages that already exist.

**[2:24]** The expensive work happened in advance.

**[2:26]** Answering you is just a lookup.

**[2:28]** Fast, cheap, done.

**[2:30]** A Google search costs

**[2:31]** a fraction of a cent.

**[2:32]** ChatGPT is a writer.

**[2:33]** When you ask it the same question,

**[2:35]** there's no answer sitting on a shelf.

**[2:37]** There's no database entry that says,

**[2:39]** "Here's what to tell this person." The

**[2:40]** model composes your answer from scratch,

**[2:43]** one word at a time, every single time.

**[2:45]** Even if a million people ask the same

**[2:47]** question today, it doesn't retrieve, it

**[2:49]** generates. And generation is expensive.

**[2:52]** A single ChatGPT query can cost 10 to 100

**[2:55]** times more compute than a Google search.

**[2:57]** Now multiply that by 900

**[2:59]** million weekly users.

**[3:00]** That's the entire reason

**[3:01]** this video exists.

**[3:03]** Search retrieves, AI generates, and

**[3:05]** generation is a manufacturing process

**[3:07]** Which brings me to the analogy I'm going

**[3:09]** to use for the rest of this video.

**[3:11]** I want you to think of an AI data center

**[3:13]** as a factory, a very strange factory.

**[3:15]** Raw material goes in

**[3:16]** one side, electricity.

**[3:18]** A product comes out the other end, words.

**[3:20]** And like any factory, it has departments,

**[3:23]** a power plant, a cooling system,

**[3:25]** assembly lines, a shipping department.

**[3:27]** We are going to tour each one.

**[3:29]** But first, three terms you need.

**[3:31]** Term one, the token.

**[3:32]** A token is the product this factory makes.

**[3:35]** Language models don't actually

**[3:36]** read words, they read tokens,

**[3:38]** which are chunks of text, roughly

**[3:40]** three-quarters of a word each.

**[3:42]** Chicken and rice is about four tokens.

**[3:44]** Your question gets chopped

**[3:45]** into tokens on the way in, and

**[3:47]** the answer gets manufactured

**[3:48]** token by token on the way out.

**[3:50]** And here's why investors care.

**[3:52]** Tokens are the unit of

**[3:53]** revenue in the AI economy.

**[3:55]** OpenAI and Anthropic literally price

**[3:57]** their product per million tokens.

**[3:59]** When you hear token, think widget

**[4:01]** comes out the assembly line

**[4:03]** Term two, the flop.

**[4:04]** A flop is one floating point operation.

**[4:07]** one single arithmetic calculation.

**[4:09]** One multiply or one add.

**[4:11]** It's a unit of labor in this factory.

**[4:13]** Manufacturing a single token requires a

**[4:15]** model to do hundreds of billions of these

**[4:17]** calculations, not per answer, per word.

**[4:20]** When people say a chip does a thousand

**[4:22]** trillion flops per second, they

**[4:23]** are telling you how many workers

**[4:25]** that chip has on the factory floor

**[4:27]** Term three, and this is a big

**[4:29]** one, training versus inference.

**[4:31]** Training is building the factory.

**[4:33]** You take a model.

**[4:34]** Think of it as a machine with

**[4:35]** over one trillion adjusting

**[4:36]** knobs called parameters.

**[4:38]** And you show it a huge portion of the

**[4:40]** reading internet, adjusting these knobs

**[4:42]** until it gets good at predicting language.

**[4:44]** This takes months.

**[4:45]** Tens of thousands of chips

**[4:46]** running around the clock, and on

**[4:48]** the order of a hundred million

**[4:49]** dollars or more per frontier model.

**[4:51]** It happens once per model.

**[4:53]** Inference is running the factory.

**[4:55]** Every time you ask ChatGPT

**[4:56]** anything, that's inference.

**[4:57]** The trained model

**[4:58]** manufacturing an answer for you.

**[5:00]** And here's the misconception I

**[5:01]** most want to kill in this video.

**[5:03]** People assume training

**[5:04]** is where the money goes.

**[5:06]** Wrong.

**[5:06]** By twenty twenty-six, roughly two

**[5:08]** thirds of all AI compute is inference.

**[5:11]** Because training happens once, but

**[5:12]** inference happens billions of times a day.

**[5:15]** OpenAI's inference bill alone is projected

**[5:17]** around fourteen billion dollars this year.

**[5:19]** The factory was expensive to build.

**[5:21]** It's even more expensive to run.

**[5:23]** Okay, so why did all of this

**[5:25]** suddenly explode after 2022?

**[5:27]** One discovery.

**[5:28]** The most economically important discovery

**[5:30]** of this decade, and most people have never

**[5:32]** heard of it: scaling law. Around 2020,

**[5:35]** researchers at OpenAI found something

**[5:37]** almost embarrassing in its simplicity.

**[5:39]** If you make the model bigger, give

**[5:41]** it more data, and spend more compute,

**[5:43]** altogether, the model gets smarter.

**[5:45]** Not sometimes. Predictably, on a

**[5:47]** chart, it's nearly a straight line.

**[5:49]** If you spend ten times more,

**[5:51]** you get a reliably better model.

**[5:53]** And stop and think about

**[5:54]** what that means for a CEO.

**[5:55]** For fifty years, better software

**[5:57]** meant hiring smarter programmers.

**[5:59]** Scaling law turned intelligence

**[6:00]** into something you could purchase.

**[6:02]** It converted AI from a research problem

**[6:04]** into a capital expenditure problem.

**[6:06]** And big companies know exactly how

**[6:08]** to compete on capital expenditures.

**[6:10]** Outspend everyone.

**[6:11]** That is the moment software

**[6:12]** stopped being about code and

**[6:14]** started being about concrete.

**[6:16]** That's why we suddenly need factories.

**[6:18]** Because here's the closing

**[6:19]** thought for this act.

**[6:20]** For the entire history of

**[6:21]** Silicon Valley, software was an

**[6:23]** escape from the physical world.

**[6:24]** zero marginal cost, infinite

**[6:26]** copies, no factory needed.

**[6:28]** AI reversed that.

**[6:29]** The frontier of software is now

**[6:31]** poured in concrete, measured in

**[6:33]** megawatts, and cooled with water.

**[6:35]** Every additional smart answer

**[6:36]** requires physical machines, physical

**[6:38]** electricity, physical heat removed.

**[6:41]** Software became heavy industry,

**[6:42]** and that changes who makes money

**[6:44]** Let's un-freeze your question.

**[6:46]** It's 7:00 PM.

**[6:47]** You have typed, "What should I

**[6:48]** make for dinner with chicken and

**[6:49]** rice?" And your thumb hits send.

**[6:51]** Here's the actual complete

**[6:53]** no-step-skips journey

**[6:54]** Step one, the trip.

**[6:55]** Your question leaves your phone as radio

**[6:57]** waves, hitting a cell tower or your

**[6:59]** Wi-Fi router, and within a few miles

**[7:01]** becomes pulses of light inside fiber

**[7:04]** optic cable, glass strands carrying

**[7:06]** data at two-thirds the speed of light.

**[7:08]** and it got routed to the nearest

**[7:09]** entry point of the AI company's

**[7:11]** network, then travel often

**[7:13]** hundreds of miles to a data center.

**[7:15]** Total time so far, a

**[7:17]** few hundred seconds

**[7:18]** Step two, the front door.

**[7:20]** Your question arrive at API gateway.

**[7:22]** think of it as a factory receiving desk.

**[7:24]** It checks who you are, checks you

**[7:26]** are not sending a thousand requests a

**[7:27]** second, runs a safety screen, and staple

**[7:30]** together everything the model needs,

**[7:31]** the system instructions, your past

**[7:33]** conversation, and your new question.

**[7:35]** Step three, tokenization.

**[7:37]** That full text get chopped into tokens.

**[7:39]** The puzzle pieces from act one,

**[7:41]** your dinner question, plus context,

**[7:43]** maybe a few hundred tokens.

**[7:45]** This get converted into numbers, because

**[7:47]** from here on, everything is math.

**[7:49]** Step four, prefill.

**[7:51]** The model reads.

**[7:52]** Here's something almost nobody knows

**[7:54]** The model process your question

**[7:55]** in two totally different phases.

**[7:57]** the first is called prefill.

**[7:59]** The model read your entire prompt

**[8:00]** all at once in parallel and build

**[8:03]** an internal understanding of it.

**[8:04]** This is a burst of raw computation,

**[8:06]** billions of calculations,

**[8:08]** And it produced something

**[8:09]** called the KV cache.

**[8:10]** don't let the name scare you.

**[8:12]** The KV cache is simply the model's

**[8:14]** working memory of your conversation.

**[8:16]** Its notes on everything said

**[8:17]** so far, held in super fast

**[8:19]** memory right next to the chip.

**[8:20]** Ever notice ChatGPT pause for a little

**[8:22]** bit before the first word appears?

**[8:24]** That pause is prefill.

**[8:25]** The factory is reading the

**[8:27]** work order. Step five, decode.

**[8:29]** The model arrives.

**[8:30]** Now the assembly line starts.

**[8:31]** The model generate the

**[8:32]** answer one token at a time.

**[8:34]** It look at your answer plus everything

**[8:36]** it has written so far, runs the

**[8:37]** entire training neural network,

**[8:39]** hundreds of billions of calculations,

**[8:41]** and produces one word, "try."

**[8:43]** Then it does the whole thing again

**[8:44]** for the next word, A. Again, one pass.

**[8:47]** Again, every single word of

**[8:49]** every ChatGPT answer on Earth is

**[8:51]** manufactured this way, one at a time.

**[8:54]** Full network pass each time.

**[8:55]** When you watch the answer type itself onto

**[8:58]** your screen, that's not a design flourish.

**[9:00]** You are literally watching an

**[9:01]** assembly line run in real time.

**[9:03]** Each word appears the

**[9:04]** moment it's manufactured.

**[9:05]** Step six, the trip home.

**[9:07]** Each token streams back through

**[9:09]** the same fiber, and two seconds

**[9:10]** later, after you hit send, you are

**[9:13]** reading the dinner ideas. One more

**[9:14]** thing happening behind the curtain.

**[9:16]** You are not alone in there.

**[9:18]** The factory will batch everything

**[9:20]** you sent with hundreds of other

**[9:21]** people's questions on the same chip

**[9:23]** simultaneously, like a delivery

**[9:25]** driver grouping orders on one route.

**[9:27]** That batching is the difference

**[9:28]** between your question costing

**[9:30]** cents and costing dollars.

**[9:32]** Now zoom all the way out because

**[9:33]** here's the whole factory in layers.

**[9:36]** This is the map for

**[9:37]** the rest of this video.

**[9:38]** 10 layers.

**[9:39]** And here's a one-sentence

**[9:40]** version of this entire video.

**[9:42]** Electricity comes in one

**[9:43]** side, flows through silicon,

**[9:45]** becomes computation and heat.

**[9:46]** The heat gets carried away by water.

**[9:48]** The computation gets coordinated by light,

**[9:51]** and what ships out the door is words.

**[9:53]** Electrons in, tokens out.

**[9:55]** That's the factory.

**[9:56]** So let's start a tour where every

**[9:57]** factory tour starts, the power plant.

**[10:00]** Because, and this surprised me the most

**[10:02]** when I first dug into this ecosystem,

**[10:04]** the story of AI in twenty twenty-six is

**[10:06]** not mainly a story about chips anymore.

**[10:08]** It's a story about electricity.

**[10:10]** Let me give you the number that

**[10:11]** framed this whole industry for me.

**[10:13]** A traditional rack of servers, the kind

**[10:15]** that ran the internet for the last twenty

**[10:17]** years, draws about five to ten kilowatts.

**[10:19]** Think of a kilowatt as ten old-fashioned

**[10:22]** 100-watt light bulbs burning at once.

**[10:24]** Nvidia's flagship AI rack, one rack,

**[10:27]** one refrigerator-sized cabinet, draws

**[10:29]** one hundred and twenty kilowatts.

**[10:31]** And the next generation coming later this

**[10:32]** year, the Vera Rubin racks, are projected

**[10:35]** to approach six hundred kilowatts per rack.

**[10:38]** That's sixty to a hundred times jump

**[10:40]** in power density in under a decade.

**[10:42]** The electrical demand of an entire

**[10:43]** neighborhood packed into a phone booth.

**[10:45]** This is called power density, and it's

**[10:47]** the root cause of nearly everything

**[10:49]** in the next two acts. Now scale up.

**[10:51]** A large AI campus today

**[10:52]** wants a gigawatt or more.

**[10:54]** A gigawatt is a thousand megawatts,

**[10:56]** roughly the output of a full-size

**[10:58]** nuclear reactor, enough electricity

**[11:00]** for about a million homes.

**[11:01]** Individual companies are now planning

**[11:03]** multiple campuses of that size.

**[11:05]** Data centers consume about

**[11:06]** four to five percent of US

**[11:07]** electricity going into this boom.

**[11:10]** The credible projections put it at nine

**[11:12]** to seventeen percent by twenty thirty.

**[11:14]** And here's the collision.

**[11:15]** The US electrical grid was built

**[11:17]** brilliantly decades ago for demand

**[11:20]** that grew one or two percent a year.

**[11:22]** AI showed up asking for

**[11:23]** tens of gigawatts right now.

**[11:25]** The grid physically cannot say yes.

**[11:27]** Two bottleneck numbers, and they are

**[11:29]** the most important numbers in this act

**[11:31]** Number one, the interconnection queue.

**[11:33]** to plug a big new facility into

**[11:35]** the grid, you file a request and

**[11:37]** wait in line while utilities study

**[11:39]** whether the grid can handle you.

**[11:41]** That line is currently

**[11:42]** four to five years long.

**[11:43]** This April, there are about

**[11:45]** four hundred and ten gigawatts of

**[11:46]** large projects waiting to connect.

**[11:49]** Eighty-seven percent of

**[11:49]** They are data centers.

**[11:51]** That's nearly five times the entire

**[11:52]** Texas grid's peak demand waiting in line.

**[11:55]** Number two, the transformer.

**[11:57]** A large power transformer, the giant

**[11:59]** gray box that steps voltage down,

**[12:00]** used to take about a year to order.

**[12:02]** Today, two and a half to four years,

**[12:04]** with prices up nearly eighty percent.

**[12:06]** You can have your chip in six months.

**[12:08]** The gray box that powers them, twenty

**[12:10]** twenty-nine. So what do you do if you're

**[12:12]** Microsoft or Meta and every month

**[12:14]** of waiting costs you the API race?

**[12:16]** You stop waiting for the grid.

**[12:17]** You go around it.

**[12:19]** The industry calls this behind-the-meter

**[12:20]** power, generating electricity

**[12:22]** on site or next door, so you

**[12:24]** never touch the public queue.

**[12:26]** And that decision, thousands of

**[12:27]** companies making it simultaneously,

**[12:29]** is what lit a fire under an entire

**[12:31]** forgotten sector of the stock market,

**[12:33]** boring old industrial power companies.

**[12:35]** Let me introduce the players, from

**[12:37]** most dramatic to most dependable.

**[12:39]** The nuclear resurrection.

**[12:41]** In 2024, Microsoft signed a

**[12:43]** deal that would have sounded

**[12:44]** like satire a decade ago.

**[12:46]** A 20-year agreement with Constellation

**[12:48]** Energy to restart Three Mile Island.

**[12:51]** The undamaged reactor next to

**[12:52]** the one from the 1979 accident.

**[12:54]** Constellation is spending about

**[12:56]** one point six billion, backed by

**[12:58]** a one billion dollar federal loan

**[12:59]** to bridge the 835 megawatt unit

**[13:02]** back in the second half of 2027.

**[13:05]** And Microsoft will buy every

**[13:06]** megawatt it produces for twenty years.

**[13:08]** Constellation operates the

**[13:09]** largest nuclear fleet in America,

**[13:11]** about twenty-two gigawatts.

**[13:12]** And suddenly those aging

**[13:13]** reactors became some of the most

**[13:15]** valuable energy assets on Earth.

**[13:17]** Why?

**[13:18]** Because AI factories run twenty-four

**[13:19]** seven, and nuclear is the only

**[13:21]** carbon-free power source that

**[13:22]** also runs twenty-four seven.

**[13:24]** Constellation stock tells the story.

**[13:26]** It's now roughly a ninety

**[13:27]** billion dollar company.

**[13:28]** Its pair, Vistra, with a thirty-seven

**[13:30]** gigawatt fleet mixing nuclear

**[13:32]** and gas, rode the same wave.

**[13:34]** The moat here is beautiful in simplicity.

**[13:36]** You cannot build a new conventional

**[13:38]** nuclear plant in America this decade.

**[13:40]** Existing reactors are replaceable.

**[13:42]** The small modular reactor lottery tickets.

**[13:44]** You have heard the tickers.

**[13:46]** Oklo, backed by Sam Altman, with

**[13:48]** over 14 gigawatts of signed pipeline.

**[13:50]** NuScale, the only SMR design

**[13:52]** actually certified by US regulators.

**[13:54]** Here's my analytical skeptical

**[13:56]** framing, and I'll be blunt.

**[13:57]** These are pre-revenue companies whose

**[13:59]** first commercial electricity arrives

**[14:01]** around twenty thirty at the earliest.

**[14:02]** Oklo doesn't yet have final

**[14:04]** regulatory approval for its design.

**[14:06]** NuScale booked about thirty-one

**[14:07]** million in revenue against a three

**[14:09]** hundred and fifty-six million loss.

**[14:11]** Both stocks are down sixty-five

**[14:12]** percent to seventy-eight percent from

**[14:14]** their late twenty twenty five peaks.

**[14:16]** That is not a business yet.

**[14:17]** That's an option on the twenty thirties.

**[14:19]** Next, the fastest power in the West.

**[14:21]** If the grid takes four years

**[14:23]** and nuclear takes 10, what can

**[14:25]** you get in 12 to 18 months?

**[14:26]** Fuel cells.

**[14:27]** Bloom Energy makes solid oxide fuel

**[14:29]** cells, boxes that convert natural

**[14:31]** gas into electricity chemically, no

**[14:33]** combustion, and you can park them behind

**[14:35]** the meter next to a data center fast.

**[14:37]** The stock nearly quadrupled

**[14:39]** in 2025, then doubled again in

**[14:41]** the first half of this year.

**[14:42]** And Bloom announced seven point six

**[14:44]** five billion in data center contracts

**[14:46]** in a single nine-day stretch.

**[14:48]** But,

**[14:48]** this July, Hunterbrook, an

**[14:50]** investigative outlet whose affiliated

**[14:53]** funds short a stock it covers.

**[14:55]** So weigh the source accordingly.

**[14:56]** Publish a report

**[14:57]** Saying that Bloom's marketed twenty

**[14:59]** billion backlog is more than forty

**[15:01]** times its binding contract obligations

**[15:03]** versus about two X for typical peers,

**[15:06]** and that scaling to its stated ambitions

**[15:08]** would consume nearly the entire global

**[15:10]** supply of scandium, a metal China

**[15:12]** now requires export licenses for.

**[15:14]** Bloom formally rejected the

**[15:15]** claims as false and misleading

**[15:17]** When a backlog number and the

**[15:19]** SEC filing disagrees by 40X, the

**[15:21]** burden of proof is on a company

**[15:22]** Next, the arms dealer

**[15:24]** setting out through 2030.

**[15:25]** My favorite business in this act is

**[15:27]** the least glamorous, GE Vernova the

**[15:29]** power spin-off of General Electric.

**[15:31]** They make the giant gas turbines

**[15:33]** that are realistically the number

**[15:35]** one near-term power source for AI

**[15:37]** because gas is the only thing you

**[15:38]** can build at scale before 2030.

**[15:40]** GE Vernova's turbine slots are sold

**[15:42]** out through the end of this decade.

**[15:44]** Their backlog is around 163 billion.

**[15:46]** In the first quarter of 2026 alone,

**[15:49]** they booked 2.4 billion in data

**[15:50]** center electrification orders.

**[15:52]** More than all of 2025.

**[15:54]** The stock is up so much it's now

**[15:55]** a nearly $300 billion company.

**[15:57]** Their only real global rival

**[15:59]** at scale, Siemens Energy.

**[16:00]** Between the substation and the

**[16:01]** chips sits a layer of equipment

**[16:04]** most people never think about.

**[16:05]** Switchgear, busways, and interruptible

**[16:08]** power supplies, the UPS, essentially

**[16:09]** a giant battery that catches the load

**[16:11]** instantly if the grid blinks, because even

**[16:14]** a half second outage can cause a training

**[16:16]** run that's been running for a month.

**[16:18]** Three companies own this layer, and

**[16:19]** remember their names because two of

**[16:21]** them show up again in the next act.

**[16:22]** Vertiv, Schneider Electric, and Eaton.

**[16:25]** Eaton's electric backlog grew

**[16:26]** forty-eight percent year over year.

**[16:27]** Vertiv's backlog more than doubled

**[16:29]** to fifteen billion dollars, and the

**[16:31]** company joined S&P 500 in March.

**[16:33]** These are the companies selling shovels

**[16:35]** to every miner, regardless of who wins.

**[16:38]** And the last line of defense,

**[16:39]** Rows of backup generators

**[16:41]** from Caterpillar and Cummins.

**[16:43]** Diesel engines the size of school

**[16:45]** buses idling in wait for the

**[16:47]** one hour a year the grid fails.

**[16:49]** Analysts think Caterpillar's data center

**[16:50]** generator business could triple by

**[16:52]** twenty-thirty. Before we move on, the

**[16:54]** uncomfortable part, because this act has

**[16:56]** one and it's showing up in your inbox.

**[16:58]** Because data centers bid for scarce

**[17:00]** power, they bid against you.

**[17:02]** In the PJM market, the grid covering 13

**[17:04]** states from Illinois to Virginia, data

**[17:07]** center demand added over nine billion

**[17:09]** dollars to the latest capacity auction,

**[17:11]** translating to residential bills rising

**[17:13]** sixteen dollars to eighteen dollars a

**[17:14]** month in parts of Ohio and Maryland.

**[17:17]** Communities are noticing.

**[17:19]** Moratoriums are being proposed.

**[17:21]** This is becoming a genuine political

**[17:22]** risk to the build-out, and any honest

**[17:25]** map of this industry has to include it.

**[17:27]** So the factory has power.

**[17:28]** one hundred kilowatts are now

**[17:29]** flowing into a single rack of chips,

**[17:32]** which create an immediate problem.

**[17:34]** Physics 101.

**[17:34]** Every one of those watts becomes heat.

**[17:37]** The factory is running a fever.

**[17:38]** A single flagship AI chip today dissipates

**[17:41]** over one thousand watts of heat.

**[17:43]** A chip the size of a postcard

**[17:44]** puts out the heat of

**[17:45]** a full-size space heater.

**[17:47]** Now stack seventy-two of them into one

**[17:49]** rack, plus their memory and networking.

**[17:51]** You have got one hundred

**[17:52]** and twenty kilowatts of heat.

**[17:53]** The output of about eighty space

**[17:55]** heaters in a cabinet you could hang.

**[17:57]** Why?

**[17:57]** Because computation is heat.

**[17:59]** Every one of those trillions of

**[18:00]** calculations pushes electrons through

**[18:02]** microscopic wires and electrical

**[18:04]** resistance turns into warmth.

**[18:06]** The factory's raw material, electricity,

**[18:08]** doesn't get consumed making tokens.

**[18:10]** It gets converted almost

**[18:12]** entirely into heat.

**[18:13]** Cooling isn't a support

**[18:14]** function of AI data center.

**[18:16]** Cooling is half the job.

**[18:17]** For thirty years, the answer was air

**[18:19]** conditioning, genuinely just fancy AC.

**[18:22]** cold air pushed up through the floor,

**[18:24]** hot air sucked out the back, giant

**[18:26]** chillers and cooling towers on the roof.

**[18:28]** And air worked fine up to about

**[18:30]** thirty to fifty kilowatts per rack

**[18:32]** But we just passed that line permanently.

**[18:34]** Air physically cannot carry

**[18:36]** heat away fast enough from a one

**[18:37]** hundred and twenty kilowatt rack.

**[18:39]** You need hurricane force

**[18:40]** winds through the servers.

**[18:41]** So the industry is undergoing its

**[18:43]** biggest plumbing change in its history.

**[18:45]** The switch from air to liquid.

**[18:47]** Water carries heat about 3,000 times

**[18:49]** more effective than air per unit volume.

**[18:51]** The technology ladder in one breath.

**[18:53]** Rear door heat exchangers, a water-cooled

**[18:56]** radiator bolted to the back of the rack.

**[18:58]** A transitional patch.

**[18:59]** Direct to chip cooling, the 2026

**[19:01]** mainstream, a metal plate with

**[19:03]** liquid channels sits directly on top

**[19:05]** of each chip, connected by hoses to

**[19:07]** a CDU, a coolant distribution unit.

**[19:09]** Think of it as the rack's heart, pumping

**[19:11]** coolant to every chip and carrying

**[19:13]** the heat to the building's water loop.

**[19:15]** Nvidia's flagship racks don't

**[19:17]** offer this as an option.

**[19:18]** They require it.

**[19:19]** And at extreme immersion cooling,

**[19:21]** literally dunking entire servers into

**[19:24]** tanks of non-conductive fluid, like

**[19:26]** deep-frying a computer that never burns

**[19:28]** two quick vocabulary items

**[19:30]** investors will encounter.

**[19:31]** PUE, power usage effectiveness.

**[19:34]** It's a factory efficiency score.

**[19:36]** Total power in divided by power

**[19:38]** that actually reaches the computer.

**[19:39]** A perfect score is one point zero. Old

**[19:41]** data centers run about 2.0, a watt of

**[19:44]** cooling for every watt of computing.

**[19:45]** Modern liquid cooled facility hits 1.1.

**[19:49]** That efficiency gap times a gigawatt

**[19:51]** times electricity prices is real money.

**[19:54]** And water.

**[19:54]** Many data centers cool by evaporating

**[19:56]** millions of gallons, which is

**[19:58]** becoming a genuine permitting and

**[20:00]** political fight in dry regions.

**[20:02]** closed loop liquid systems help.

**[20:04]** But watch this issue.

**[20:05]** It decides where facilities

**[20:07]** get built. Who gets paid?

**[20:08]** Largely the same names as the power

**[20:10]** room, Vertiv is the market leader.

**[20:12]** The rare company selling both the

**[20:14]** power gear and the liquid cooling.

**[20:16]** A one-stop shop growing revenue

**[20:17]** twenty-eight percent a year

**[20:19]** at twenty percent margin.

**[20:20]** And then something remarkable happened.

**[20:22]** The two electrical giants each spend

**[20:24]** billions to buy their way into liquid

**[20:26]** cooling within months of each other.

**[20:28]** Eaton paid about nine point

**[20:29]** five billion for Boyd Thermal.

**[20:31]** Schneider Electric bought Motivair.

**[20:33]** When the electrics more

**[20:34]** disciplined industry acquires,

**[20:36]** both pay up for the same niche.

**[20:38]** They are telling you what they think

**[20:40]** every future data center looks like.

**[20:41]** smaller pure players nVent

**[20:43]** including loops and enclosures

**[20:45]** and private cool IT systems.

**[20:47]** The specialists whose cold plates

**[20:49]** ship inside many brand name servers.

**[20:51]** The liquid cooling market was

**[20:52]** about five billion in 2025.

**[20:54]** Forecasts put it at fifteen to

**[20:55]** twenty-seven billion by the early 2030s.

**[20:58]** It's the single clearest picks and

**[21:00]** shovels growth lane in this entire

**[21:01]** ecosystem because it does not care

**[21:03]** whether NVIDIA or AMD or Google wins.

**[21:06]** Heat is heat.

**[21:07]** All right, the factory has power.

**[21:09]** The fever is under control.

**[21:10]** It's time to walk onto the factory

**[21:12]** floor and meet the machine your dinner

**[21:14]** question actually runs on and the three

**[21:16]** trillion dollar company that built it

**[21:17]** This is the machine your

**[21:18]** dinner question runs through.

**[21:20]** Nvidia's GB200, NVL72 seventy-two

**[21:23]** GPUs wired together so tightly they

**[21:25]** behave as a single giant computer.

**[21:27]** It weighs about a ton and a half,

**[21:29]** draws those one hundred and twenty

**[21:30]** kilowatts we discussed, and costs

**[21:32]** roughly three million dollars.

**[21:34]** So let's open it up and to keep the

**[21:35]** parts straight, come back to the

**[21:37]** factory, specifically its kitchen.

**[21:39]** The CPU is the head chef.

**[21:41]** The central processing unit runs

**[21:42]** the operating system, takes orders,

**[21:45]** coordinates everything, brilliant

**[21:47]** at complex sequential tasks.

**[21:48]** But there's only a handful of them.

**[21:50]** For decades, the CPU was the star

**[21:52]** of computing, Intel's kingdom.

**[21:54]** In the AI server, it has

**[21:55]** been demoted to management

**[21:57]** The GPUs are 10,000 line cooks.

**[21:59]** A graphics processing unit,

**[22:01]** originally invented to draw video

**[22:03]** game graphics, contains thousands

**[22:04]** of small, simple cores that all do

**[22:07]** the same operation simultaneously.

**[22:08]** It turns out the math inside a neural

**[22:10]** network is exactly that kind of work.

**[22:13]** Billions of identical

**[22:14]** multiply and add operations.

**[22:16]** One head chef cannot do that.

**[22:18]** 10,000 line cooks, each chopping

**[22:20]** one onion at the same instant, can

**[22:22]** That accident of history, gaming

**[22:24]** graphics and AI needing the same math,

**[22:26]** is the foundation of Nvidia's empire.

**[22:29]** HBM is the countertop,

**[22:30]** high bandwidth memory.

**[22:32]** Hold that thought.

**[22:33]** It gets its own act.

**[22:34]** The SSD is the pantry.

**[22:36]** The NIC, the network interface

**[22:37]** card, is the waiter carrying

**[22:39]** dishes between kitchens.

**[22:41]** And the power supplies and the motherboard

**[22:43]** are the plumbing and wiring holding

**[22:45]** it all together. Now the companies.

**[22:47]** Nvidia finished its last fiscal year with

**[22:49]** $215.9 billion in revenue, up 65%, of

**[22:53]** which about $194 billion was data center.

**[22:55]** It controls roughly 80 to 86%

**[22:58]** of the AI accelerator market.

**[23:00]** Its gross margin in the recent quarter

**[23:01]** is about 75%, 75% on hardware.

**[23:05]** Apple, the most admired hardware

**[23:07]** company in history, runs around 46.

**[23:09]** Nvidia became the first $5

**[23:10]** trillion company last October.

**[23:12]** And depending on the week,

**[23:13]** roughly seven cents of every

**[23:15]** dollar in the S&P 500 is Nvidia.

**[23:17]** How is that margin possible?

**[23:19]** Everyone says best chips, and sure,

**[23:21]** but the real answer is a word we'll

**[23:23]** unpack fully in act eight, CUDA.

**[23:25]** Twenty years of software that every

**[23:27]** AI developer on Earth was trained on.

**[23:29]** For now, the one line version,

**[23:31]** Nvidia doesn't only sell chips.

**[23:33]** It sells the only complete factory

**[23:34]** floor system the world's engineers

**[23:36]** already know how to operate.

**[23:37]** Buying a competitor's chip means

**[23:39]** retraining your whole workforce

**[23:41]** The bear case, about forty percent

**[23:43]** of Nvidia's revenue comes from just

**[23:44]** four customers, and all four are

**[23:46]** building their own chips to replace it.

**[23:48]** Next, the challenger, AMD.

**[23:50]** AMD's Instinct GPUs are genuinely

**[23:52]** competitive on inference, more memory per

**[23:54]** chip, and by some estimates, twenty-five

**[23:57]** to forty percent better tokens per dollar.

**[23:59]** Their problem was never silicon.

**[24:01]** It's software.

**[24:02]** Their CUDA alternative, called

**[24:03]** ROCm, now hits ninety to

**[24:05]** ninety-five percent of Nvidia's

**[24:07]** performance on standard workloads.

**[24:09]** But ninety percent as good with more

**[24:10]** friction is a hard pitch when your

**[24:13]** training can cost a hundred million.

**[24:14]** AMD holds maybe five to

**[24:16]** seven percent of this market.

**[24:17]** Watchable, but it's improving.

**[24:19]** still a distant second.

**[24:21]** Intel, painful to say,

**[24:22]** is barely in this race.

**[24:24]** Still selling plenty of head chefs, but

**[24:26]** the kitchen stopped being about head chefs

**[24:28]** Now the quiet assassin, Broadcom.

**[24:30]** Here's the plot twist most

**[24:31]** retail investors miss.

**[24:33]** Those four hyperscalers building their own

**[24:34]** chips, they cannot actually do it alone.

**[24:37]** Designing a frontier AI chip

**[24:38]** takes a decade of specialized IP.

**[24:40]** So they hire Broadcom, which co-designs

**[24:43]** Google's TPU, Meta's training

**[24:44]** chip, and reportedly OpenAI's.

**[24:47]** Broadcom's books over sixty

**[24:49]** percent of custom chip market.

**[24:50]** Posted AI revenue up one hundred

**[24:52]** and six percent last quarter, with

**[24:54]** a seventy-three billion backlog.

**[24:55]** And management says it has lines

**[24:57]** of sight to a hundred billion of

**[24:58]** AI revenue in twenty twenty-seven.

**[25:00]** In our factory analogy, Nvidia

**[25:02]** sells finished kitchens.

**[25:03]** Broadcom helps the biggest

**[25:05]** restaurant chains build their

**[25:06]** own and takes a cut either way.

**[25:08]** It also dominates the

**[25:09]** switch silicon in ASIC.

**[25:11]** One company, both sides of the war

**[25:13]** Now, who actually built those racks?

**[25:15]** Not Nvidia.

**[25:16]** Nvidia designs.

**[25:17]** Super Micro integrate full liquid

**[25:19]** cooled racks faster than anyone.

**[25:21]** Revenue up 123% last quarter.

**[25:23]** Over 90% of it AI.

**[25:25]** Gross margin between six and

**[25:26]** 10%, depending on the quarter.

**[25:28]** Dell has taken over 64 billion in

**[25:30]** AI server orders with a 43 billion

**[25:33]** backlog, a staggering number as

**[25:35]** server segment margins under 9%.

**[25:37]** HPE writes supercomputing

**[25:39]** and sovereign AI deals.

**[25:40]** and beneath the brands sit the true

**[25:42]** invisible giants, the Taiwanese ODMs,

**[25:44]** original design manufacturers, Foxconn,

**[25:47]** which assembles roughly 40% of the

**[25:49]** world's AI racks, Quanta, Wiwynn

**[25:51]** Celestica.

**[25:52]** The same rack passed through many hands.

**[25:54]** Nvidia captures seventy-five

**[25:55]** percent margin on the silicon.

**[25:57]** The company that physically screws

**[25:58]** it all together keeps six to 10.

**[26:00]** In hardware, the profit

**[26:02]** lives in whatever is scarce.

**[26:03]** Chips and software are scarce.

**[26:05]** Assembly is not.

**[26:06]** But I've been hiding something from you.

**[26:08]** I said seventy-two GPUs

**[26:09]** behave as a single computer.

**[26:11]** Nobody hit that with a magic wand.

**[26:13]** Making ten thousand line cooks work

**[26:15]** as one brain is arguably the hardest

**[26:17]** engineering problem in the entire

**[26:19]** building, and it's where some of the

**[26:20]** best businesses in the ecosystem hide

**[26:22]** So why can't one GPU do the job?

**[26:25]** Simple.

**[26:26]** The model doesn't fit.

**[26:27]** A Frontier model has over

**[26:28]** one trillion parameters.

**[26:30]** Those adjacent knobs require

**[26:32]** terabytes of ultra-fast memory.

**[26:34]** The biggest GPU carries

**[26:35]** a few hundred gigabytes.

**[26:36]** so the model gets sliced

**[26:37]** across thousands of chips.

**[26:39]** And here's the consequence.

**[26:40]** To produce every single token, those

**[26:42]** chips must exchange intermediate results

**[26:45]** constantly at unimaginable speed.

**[26:48]** Back to the kitchen.

**[26:49]** Ten thousand line cooks

**[26:50]** preparing one dish together.

**[26:52]** every chef needs ingredients

**[26:53]** from other chefs every second.

**[26:55]** If passing ingredient is slow,

**[26:57]** your ten thousand chefs stand

**[26:59]** around waiting, and these are the

**[27:00]** most expensive chefs in history.

**[27:02]** at cluster scale, a network

**[27:04]** that's ten percent slower can idle

**[27:06]** billions of dollars of silicon.

**[27:08]** That's why networking is roughly

**[27:09]** forty to sixty percent of

**[27:11]** spending for every dollar of GPU.

**[27:13]** Two words to define.

**[27:14]** Bandwidth is how much

**[27:16]** data moves per second.

**[27:17]** the width of the conveyor belt.

**[27:19]** Latency is the delay for one handoff.

**[27:22]** How long a single pass takes.

**[27:23]** AI needs both everywhere at once.

**[27:26]** The wiring comes in two flavors.

**[27:28]** Scale up inside the rack is Nvidia's

**[27:31]** proprietary NVLink, an extreme speed web

**[27:33]** that makes seventy-two GPU one machine.

**[27:35]** Scale out rack to rack across the

**[27:38]** building is where a war is being fought.

**[27:40]** For years, serious AI clusters ran on

**[27:43]** InfiniBand, a specialized ultra-low

**[27:45]** latency networking that NVIDIA

**[27:47]** acquired in 2019 with Mellanox.

**[27:50]** Premium performance,

**[27:51]** premium price, one vendor.

**[27:53]** Against it, Ethernet, the open

**[27:55]** universal standard, the same family

**[27:57]** of technology as your home network.

**[27:59]** Historically slower, but backed by

**[28:01]** literally everyone who isn't NVIDIA.

**[28:03]** By early 2026, about two-thirds of

**[28:05]** new AI cluster networking is Ethernet.

**[28:08]** Open standards, given time, usually win.

**[28:10]** They just did.

**[28:11]** This is a proprietary versus

**[28:13]** open story as old as tech

**[28:15]** who profit from the nervous system?

**[28:17]** Broadcom again.

**[28:18]** Its Tomahawk chips are the

**[28:20]** merchant silicon inside most

**[28:22]** high-end Ethernet switches.

**[28:23]** Arista Networks build the switches

**[28:25]** themselves, the best in class boxes and

**[28:27]** software hyperscalers standardize on.

**[28:29]** Roughly six percent gross margin, guiding

**[28:31]** to eleven point five billion this year,

**[28:33]** with the risk that its two biggest

**[28:35]** customers are forty percent of revenue.

**[28:37]** Cisco, the incumbent, still huge,

**[28:40]** fighting to stay relevant in AI backend.

**[28:42]** Marvell plays the Broadcom

**[28:43]** playbook one tier down.

**[28:45]** custom chips for Amazon and Microsoft,

**[28:47]** plus leadership in the digital signal

**[28:49]** processor inside optical modules.

**[28:52]** And Astera Labs, one of the most

**[28:54]** remarkable margin stories in ecosystem,

**[28:56]** Makes tiny retimer chips that clean up

**[28:59]** electrical signals degrading over mere

**[29:01]** inches of circuit board at this speed.

**[29:04]** Boring, invisible, seventy-six

**[29:06]** percent gross margins.

**[29:07]** Revenue up ninety-three

**[29:09]** percent last quarter.

**[29:10]** When data moves this fast, even the

**[29:12]** space between two chips become a market

**[29:14]** And then the plot

**[29:15]** literally turned to light.

**[29:16]** copper wires can carry this speed only

**[29:18]** at a few meters before signals degrade.

**[29:20]** Fine inside a rack.

**[29:22]** Use this across a football field building.

**[29:24]** So between racks, everything converts

**[29:26]** to light through glass fiber.

**[29:28]** The device doing the conversion is the

**[29:29]** optical transceiver, a thumb-sized gadget,

**[29:32]** an electrical to light translator, and you

**[29:34]** need one at each end of every fiber link.

**[29:37]** A single large AI cluster consumes

**[29:38]** hundreds of thousands of them at

**[29:40]** hundreds of thousands of dollars

**[29:41]** each, replaced every upgrade cycle.

**[29:44]** It's a razor blade for the data centers.

**[29:46]** The names?

**[29:47]** Coherent, the market leader, Lumentum

**[29:49]** Chinese volume champion Innolight.

**[29:51]** Fabrinet.

**[29:52]** The contract manufacturer that

**[29:54]** assemble for nearly all of them.

**[29:55]** The arms dealer's arms dealer.

**[29:57]** Corning, which draws

**[29:58]** the glass fiber itself,

**[29:59]** And Amphenol, whose connectors are

**[30:02]** the knuckles of the entire system.

**[30:04]** The frontier to watch is

**[30:05]** co-packaged optics, moving the

**[30:07]** light conversion directly onto

**[30:08]** the switch chip to slash power

**[30:10]** The nervous system is built.

**[30:11]** 10,000 chips thinking as one.

**[30:14]** But there's a dirty secret

**[30:15]** on the factory floor.

**[30:16]** Most of the time, the most expensive

**[30:18]** chips in the world are waiting, not for

**[30:20]** data from across the room, for data from

**[30:22]** two centimeters away. Here's a secret.

**[30:25]** During the decode phase, the one word at a

**[30:27]** time assembly line from act two, the GPU's

**[30:30]** math cores are often not the bottleneck.

**[30:32]** For every token, the chip must

**[30:34]** pull the model's parameters and the

**[30:36]** conversation's working memory, that

**[30:38]** KV cache, from memory into its cores.

**[30:40]** The math is fast.

**[30:42]** The fetching is slow.

**[30:43]** Modern inference is what engineers

**[30:45]** call memory bandwidth bound.

**[30:47]** The line cooks are lightning,

**[30:48]** but a countertop cannot feed them

**[30:50]** ingredients fast enough, which makes

**[30:51]** the countertop one of the most valuable

**[30:53]** pieces of real estate in technology

**[30:55]** The industry's answer is

**[30:57]** HBM, high bandwidth memory.

**[30:59]** Instead of laying memory chips flat on

**[31:01]** a board a few inches from the processor,

**[31:03]** HBM stacks them vertically, eight,

**[31:06]** 12 stories high, drills thousands of

**[31:08]** microscopic elevator shafts through the

**[31:10]** silicon, and glues the whole tower directly

**[31:12]** next to the GPU on the same package.

**[31:14]** A skyscraper of memory downtown instead

**[31:17]** of suburbs of memory across a highway.

**[31:19]** The result is five to six times the

**[31:21]** bandwidth of conventional memory

**[31:22]** at five to six times the price.

**[31:24]** Nvidia happily pays.

**[31:25]** Memory is now one of the biggest

**[31:27]** cost components inside every AI chip

**[31:29]** you have heard of, and only three

**[31:30]** companies on earth can make it

**[31:32]** SK Hynix, the Korean company, owns

**[31:34]** roughly 60% of HBM and got there by

**[31:37]** out-executing its giant neighbor.

**[31:39]** It bet on HBM years before it

**[31:41]** mattered and shipped each generation

**[31:43]** first, unlocking the lion's share of

**[31:45]** Nvidia's next generation allocation.

**[31:47]** Samsung, the largest memory company

**[31:49]** overall, was embarrassingly late.

**[31:51]** Micron, the American champion, went

**[31:53]** from afterthought to selling out

**[31:55]** its entire 2026 HBM capacity in

**[31:58]** advance. Here's the stat for the act.

**[32:00]** In May 2026, all three memory makers

**[32:02]** crossed a trillion dollars in market value,

**[32:04]** combined over four trillion, roughly

**[32:06]** sixteen times what they were a decade ago.

**[32:08]** Memory used to be the most brutal

**[32:10]** commodity business in tech.

**[32:11]** Boom, bust, bankrupt, repeat.

**[32:14]** HBM changed the psychology.

**[32:16]** It sold out more than a year in

**[32:17]** advance, allocated like a scarce

**[32:19]** resource, priced like a luxury good.

**[32:21]** The open question is whether that price

**[32:23]** survives the moment all three giants

**[32:25]** finish their capacity expansions at once.

**[32:27]** Memory cycles have broken hearts before

**[32:30]** Now walk out the back of the

**[32:31]** factory to the warehouse storage.

**[32:33]** The hierarchy in one line, the closer

**[32:35]** to the chip, the faster and pricier.

**[32:37]** Cache on the chip itself, HBM beside

**[32:40]** it, regular DRAM on the motherboard,

**[32:42]** SSDs, flash drives for hot data, and

**[32:45]** at the bottom, the technology everyone

**[32:47]** declared dead ten years ago, the

**[32:49]** spinning hard drive, still unbeatable

**[32:51]** per terabyte for cold bulk data.

**[32:53]** And AI turned out to be an AI hoarder.

**[32:55]** Training datasets, model

**[32:56]** checkpoints saved every few hours.

**[32:58]** And the part nobody predicted, the output.

**[33:01]** Every conversation, every log, every

**[33:03]** generated image retained forever.

**[33:05]** Seagate CEO calls it the

**[33:07]** inference inflection.

**[33:08]** AI doesn't just consume data,

**[33:10]** It produces it endlessly.

**[33:12]** Hard drives are a literal duopoly,

**[33:14]** Seagate and Western Digital, plus flash

**[33:16]** drive players Kioxia and Solidigm.

**[33:19]** After a decade of decline, both

**[33:20]** drive makers sold out their

**[33:22]** entire production into 2027.

**[33:24]** Western Digital now ships 89% of its

**[33:26]** revenue to cloud customers and was

**[33:28]** one of the S&P 500 top performers.

**[33:30]** Consumer hard drive prices jumped

**[33:32]** 50% because AI ate the supply.

**[33:34]** A dying industry resurrected by

**[33:36]** the factory next door needing

**[33:37]** somewhere to put infinity.

**[33:39]** The machine is complete,

**[33:40]** powered, cooled, wired, fed.

**[33:42]** But a pile of perfect hardware

**[33:44]** answers exactly zero questions.

**[33:46]** Something invisible has to run the place.

**[33:48]** Everything we have toured so

**[33:49]** far, you could theoretically buy.

**[33:51]** Hardware is purchasable, where

**[33:52]** does the durable competitive

**[33:53]** advantage actually live?

**[33:55]** The stack briefly bottom to top.

**[33:57]** Linux, the free operating system running

**[33:59]** effectively every server on Earth,

**[34:01]** commercialized by Red Hat and Canonical.

**[34:04]** Kubernetes, the invisible foreman.

**[34:06]** Open source software that schedules

**[34:07]** work across thousands of machines.

**[34:09]** restarts what crashes and keeps

**[34:11]** the factory floor humming.

**[34:12]** Then the serving layer, then the models.

**[34:15]** Two stories in this act matter to

**[34:16]** investors more than all the rest combined.

**[34:18]** First story, CUDA, the twenty-year

**[34:20]** trap. In 2006, Nvidia made a

**[34:22]** decision Wall Street hated.

**[34:24]** It spent billions building a

**[34:25]** programming platform so scientists

**[34:27]** can use gaming chips for general math.

**[34:29]** For a decade, this looked

**[34:31]** like an expensive hobby.

**[34:32]** Then deep learning arrived, and

**[34:34]** every AI researcher on Earth

**[34:35]** learned to build on CUDA because

**[34:37]** it was the only mature option.

**[34:39]** Twenty years later, CUDA has

**[34:40]** millions of developers, thousands

**[34:42]** of specialized libraries, and every

**[34:44]** framework optimized for it first.

**[34:46]** Understand what this means.

**[34:47]** When AMD ships a chip with better

**[34:49]** specs, and it sometimes does, the

**[34:51]** customer isn't comparing chips.

**[34:53]** They're comparing chips plus the

**[34:55]** cost of retaining their entire

**[34:56]** engineering organization and rewriting

**[34:58]** their code with their one hundred

**[35:00]** million training run on the line.

**[35:01]** That's why seventy percent gross

**[35:03]** margins survive competition.

**[35:04]** The moat was never the silicon.

**[35:06]** The moat is the muscle

**[35:07]** memory of a million engineers.

**[35:09]** Second one, why your question

**[35:10]** costs cents, not dollars?

**[35:12]** Raw, naive inference on trillion

**[35:13]** parameter model would be super expensive.

**[35:16]** The economics only work because of an

**[35:18]** unglamorous layer called a serving engine.

**[35:20]** Software like vLLM and NVIDIA's

**[35:23]** TensorRT doing three tricks.

**[35:25]** Batching, grouping hundreds of users'

**[35:27]** questions through the chip at once.

**[35:28]** The delivery route trick from Act Two.

**[35:30]** Caching, reusing the KV working

**[35:32]** memory instead of recomputing the

**[35:34]** conversation from scratch for every word.

**[35:36]** And quantization.

**[35:38]** Rounding the model's numbers to lower

**[35:39]** precision, like shipping a slightly

**[35:41]** compressed photo, nearly identical

**[35:43]** quality, fraction of the cost.

**[35:45]** Together a three to ten times cost

**[35:47]** reduction from software alone.

**[35:49]** When OpenAI or Anthropic cuts

**[35:51]** API price 80% in a year, it's

**[35:53]** mostly this layer, not new chips.

**[35:55]** Token manufacturing costs

**[35:56]** are collapsing our curve.

**[35:57]** Remember that for the finale,

**[35:59]** because it cuts both ways.

**[36:00]** One last floor of the stack.

**[36:02]** Your dinner question doesn't

**[36:03]** need it, but enterprise AI does.

**[36:05]** RAG.

**[36:06]** Retrieval augmented generation.

**[36:08]** The model is a brilliant

**[36:09]** writer with a fixed education.

**[36:10]** RAG hands it your company's documents

**[36:13]** at question time via vector database.

**[36:15]** search engine that finds text by

**[36:17]** meaning rather than keywords from

**[36:19]** players like Pinecone, plus the data.

**[36:21]** platform Databricks and Snowflake.

**[36:23]** The librarian and the

**[36:24]** writer working together.

**[36:26]** That's the enterprise AI

**[36:27]** pitch in one sentence.

**[36:28]** And with that, the tour is over.

**[36:30]** You have seen every layer: power,

**[36:32]** cooling, silicon, light, memory, software.

**[36:35]** Now let's follow the money, all of it

**[36:36]** in one map, and then ask one question:

**[36:39]** Does any of this actually pay for itself?

**[36:41]** So here's the whole board.

**[36:42]** On the demand side, the miners, three

**[36:44]** tiers, the hyperscalers, Microsoft,

**[36:46]** Amazon, Google, Meta, Oracle,

**[36:48]** spending their combined seven hundred

**[36:50]** and twenty-five billion this year.

**[36:51]** The neoclouds, specialized GPU landlords,

**[36:54]** CoreWeave, Nebius, Lambda, Crusoe

**[36:56]** And AI labs, OpenAI now

**[36:59]** valued around $850 billion.

**[37:01]** Anthropic, which passed it at

**[37:03]** roughly $965 billion on about $47

**[37:05]** billion of annualized revenue.

**[37:06]** And XAI folded into SpaceX.

**[37:09]** Both major labs filed for IPO this June.

**[37:12]** The private market has already

**[37:13]** priced them as two of the most

**[37:14]** valuable companies on Earth.

**[37:16]** The public market is about

**[37:17]** to vote on the costs.

**[37:19]** Bernstein estimates one gigawatt

**[37:20]** of AI data center, one campus,

**[37:23]** costs about $35 billion to build.

**[37:24]** Where it goes, roughly 39% is the

**[37:27]** chips, the single biggest line, which

**[37:29]** is why Nvidia's gross profit alone is

**[37:31]** estimated near 30% of total industry cost.

**[37:34]** And that thirty-five billion

**[37:35]** machine depreciates fast.

**[37:37]** most operators write chips off over

**[37:39]** four to six years, And skeptics argue

**[37:41]** even that flatters the accounting since

**[37:43]** a five-year-old GPU competes against

**[37:45]** chips 10 times better Every year of

**[37:47]** depreciation life added or removed

**[37:50]** swings billions in reported profit.

**[37:52]** Watch that debate

**[37:53]** so follow this.

**[37:54]** NVIDIA invests billions directly into

**[37:55]** CoreWeave, into Nebius, into OpenAI.

**[37:58]** those companies use the

**[37:59]** money to buy NVIDIA chips.

**[38:01]** revenue for NVIDIA.

**[38:02]** The hyperscalers sign enormous

**[38:03]** contracts with the neoclouds.

**[38:05]** Meta alone committed roughly

**[38:06]** twenty-one billion to CoreWeave and

**[38:08]** reportedly twenty-seven billion to

**[38:09]** Nebius, which conveniently moves

**[38:11]** data center spending off Meta's

**[38:13]** balance sheet into someone else's debt.

**[38:15]** The AI lab signs compute

**[38:16]** deals with hyperscalers.

**[38:17]** OpenAI with Microsoft and Oracle.

**[38:19]** Anthropic with Amazon and Google.

**[38:21]** paid partly with money those same

**[38:23]** hyperscalers invested into them.

**[38:25]** Let's be fair, because the analytical

**[38:27]** skeptic label cuts both ways.

**[38:29]** This is not fraud, and it's not new.

**[38:31]** Vendor financing built the

**[38:32]** railroads and the telephone network.

**[38:34]** The loop has exactly one opening

**[38:36]** where fresh money is supposed to

**[38:38]** enter, end users and enterprises.

**[38:40]** You're paying twenty bucks a month,

**[38:42]** companies paying for APIs and copilots.

**[38:44]** That is the only exit that

**[38:45]** isn't recycled capital.

**[38:47]** And today, that corner is the

**[38:48]** smallest number on the board.

**[38:50]** The two biggest labs combined annualized

**[38:52]** run rates now top seventy billion.

**[38:54]** Generally spectacular.

**[38:55]** The fastest revenue ramp

**[38:56]** in software history.

**[38:58]** But the revenue they'll actually book

**[38:59]** this calendar year is a fraction of

**[39:01]** that, set against seven hundred and

**[39:03]** twenty-five billion of annual spend,

**[39:04]** with OpenAI still expected to lose on

**[39:07]** the order of fourteen billion this year

**[39:09]** The entire $7 trillion machine is

**[39:10]** the bet that a small corner grows

**[39:12]** faster than the big circle spins.

**[39:15]** So let's watch it one more time.

**[39:16]** Same two seconds, but now you can see.

**[39:18]** Your thumb hits send, and the

**[39:20]** question becomes light in glass

**[39:21]** fiber crossing state lines.

**[39:23]** arrive at a building

**[39:24]** drawing the power of a city.

**[39:26]** power from a restarted nuclear

**[39:27]** plant, a Solar turbine, a fuel

**[39:29]** cell parked behind the meter.

**[39:30]** It's chopped into tokens and

**[39:32]** fed into a three million dollar

**[39:34]** rack assembled in Taiwan, sold

**[39:36]** at a seventy-five point margin,

**[39:38]** where ten thousand line cooks

**[39:39]** fetch a trillion parameters

**[39:41]** from memory skyscrapers.

**[39:42]** While liquid coolant carries away

**[39:44]** the heat of eighty space heaters,

**[39:46]** and light speed interconnects let

**[39:48]** ten thousand chips think one thought.

**[39:50]** The software batches you with one

**[39:52]** thousand strangers, and the answer

**[39:54]** streams back token by token, each

**[39:56]** word manufactured the instant you read it

**[39:58]** Two seconds, $700 billion a year,

**[40:00]** the largest infrastructure project

**[40:02]** our species has ever attempted.

**[40:04]** So a machine can suggest you make

**[40:05]** one-pan chicken and rice. Whether

**[40:07]** that's the most important investment

**[40:09]** in history or the most expensive,

**[40:11]** Honestly, nobody on Earth knows yet.

**[40:13]** That's the whole map.

**[40:14]** If you find this video helpful,

**[40:15]** please subscribe, and I'll

**[40:17]** see you in the next one.
