What Is an AI Factory? Energy In, Tokens Out, and Who Actually Gets Paid
What the AI factory frame actually tells an operator about their own megawatts.
At GTC 2026, Jensen Huang compressed the AI infrastructure thesis into five words: "Compute is your revenue now." He was not describing a new building type. He was describing a new income statement.
The term for it, AI factory, shows up on so many vendor product pages that it is easy to dismiss as marketing. That would be a mistake. Underneath the branding is a simple question about how an operator makes money, and it applies to anyone who controls megawatts.
TLDR
- An AI factory is an accounting frame, not an architecture. Power and hardware in, tokens out, priced by a market the operator does not control.
- The gating question is whether your revenue unit is the megawatt or the token. That answer determines your capex and your risk.
- Digital twins matter most for retrofits, where the question is whether a site can reach competitive token output at all.
What an AI factory is, and what it changes
An AI factory is a facility whose measured output is token throughput rather than availability or stored capacity. NVIDIA's definition frames it as infrastructure managing the full AI lifecycle—from data ingestion through training to high-volume inference, with intelligence as the product. NVIDIA has used the term since 2024, but it was at GTC 2026 in San Jose where it became the organizing thesis of Jensen Huang's keynote.
("AI factory" is also a product name, popularized largely by Dell and NVIDIA. When a vendor uses it, they mean a validated stack you can buy. Here, we mean the category of facility.)
The distinction from a conventional data center is not merely a rebrand to avoid pushback from data center opposition. It reduces to one mechanical fact. In a colocation model, an idle rack still bills. In a factory model, idle compute earns nothing while it depreciates, which drives every utilization decision downstream.
Your supply is energy
You do not buy compute. You buy power and convert it, which makes everything downstream a conversion efficiency question. The metric is tokens per watt, or at facility scale, tokens per megawatt, a number NVIDIA now markets its platforms on.
Power is constrained, though not quite how most coverage suggests. ERCOT CTO Venkat Tirupati told GTC that roughly 230 GW of large loads sit in the interconnection queue, around 70% of them data centers, against a system peak near 85 GW. The real problem is timing. A new load can connect in 6 to 18 months. New generation takes one to two years. New transmission takes three to six. Demand arrives years before the supply built to serve it.
That gap is why flexibility became an interconnection strategy rather than a concession. A 500 MW request structured as 100 MW firm and 400 MW flexible connects faster and earns demand-response compensation. For AI load this is not theoretical: Luxor Energy and Bentaus cut a live GPU's draw to 25% in under 500 milliseconds on an ERCOT 4CP signal, without disrupting the inference workload on it.
Your product is tokens
A token is a unit of model input or output, roughly a word fragment. What matters is the accounting: token throughput is the first data center output metric that maps to revenue rather than capacity.
Hardware generation moves that number hard. NVIDIA has cited GB300 NVL72 systems delivering roughly 50x more tokens per megawatt than Hopper-era systems, a manufacturer benchmark run under manufacturer-chosen conditions and better treated as a ceiling than a planning number.
When the work happens changed too. Training ran in bursts: you booked the cluster, trained the model, stopped. The agentic workloads OpenAI's Sachin Katti described at GTC never stop. A site that earns around the clock is a better business than one that earns in bursts, and a harder one to run.
Your demand is the compute rental market
This is where the frame gets uncomfortable. Model providers set price per million tokens, which sets what a neocloud charges per GPU-hour, which sets what an operator earns per megawatt. Two steps down that chain you are a price taker, selling GPU-hours at a rate you can observe but not influence. That is why we track asking prices in the AI Hardware Price Index.
Two assumptions get built into operator models, and both are wrong:
- Rental rates always fall. H100 rates did fall after the 2023 peak, then reversed in 2026 as Hopper capacity sold out. If your model draws a straight line down from the peak, it is already wrong.
- Utilization is one number. Offtake and financed builds run near 100% because a contract guarantees it. The 70% figure that gets quoted applies to self-funded operators selling into a marketplace. Pick the wrong one and your payback is off by years.
Useful life sets the denominator. CoreWeave's Michael Intrator argued at GTC that inference disaggregation could push GPU useful life from the assumed four to five years toward eight to ten, a thesis resting on software reliably supporting heterogeneous hardware. If it holds, lifetime token output roughly doubles and capex recovery becomes the binding concern.
AI factory digital twins and the retrofit question
A digital twin here means a physics-accurate simulation of a facility, not a person's AI avatar. It pulls power, cooling, networking, and structural data into one model so teams can test design changes and failure scenarios before construction. NVIDIA's Omniverse blueprint for AI factory design is the reference implementation, and its greenfield use is well documented.
The retrofit case isn't discussed often but it's the one most operators face. You have a building with fixed constraints, so the question is whether the site can reach a competitive tokens-per-megawatt figure at all, answered before you spend capital finding out. Build the simulation and start with the floor. Can the slab carry the weight of high-density racks? Then cooling: can a liquid loop be installed, and where does the heat go? Then power distribution at the rack. Floor loading stops more conversions than cooling does, and it is the cheapest item on that list to check.
Do you sell the megawatt or the token?
The AI factory frame only changes your decisions if your revenue unit is the token. Where you sit on the value chain determines whether tokens per watt is a line in your P&L or your tenant's.
A landlord supplies power, shell, and cooling, earns in contracted kW, and carries tenant credit risk rather than utilization risk: lowest capex, lowest ceiling. An operator owns the compute and sells GPU-hours: highest capex, highest ceiling, and both utilization and residual value risk. An operator with a signed offtake contract gets paid whether the compute is busy or not, because the buyer took that risk. What limits them is financing and the quality of that buyer.
Even a pure landlord should care about tenant tokens per megawatt, because a tenant whose site converts power inefficiently cannot sustain the rent. For mining operators, power procurement, curtailment economics, and site operations map directly, and the demand-response structures being built for AI load resemble how miners already work with the grid, as in our GTC 2026 takeaways.
What the AI factory frame is actually for
The term earns its keep because it names a unit of output that maps to revenue, which uptime and PUE never did. It does not tell you to become a token producer. It tells you to know which unit you sell and price accordingly.
The AI factory frame will not make an operator money. It tells an operator where in the chain the money is actually counted, and whether they own that step or someone else does.
Hashrate Index Newsletter
Join the newsletter to receive the latest updates in your inbox.