
This article first appeared in The Edge Malaysia Weekly on September 14, 2026 - September 20, 2026
OVER the past few weeks, this column has covered the economics of data centres as well as circular financing and the off-balance sheets debts of artificial intelligence (AI) firms. With the mega initial public offering (IPO) of frontier AI lab Anthropic, which expects to raise more than US$100 billion (RM406.3 billion) at a valuation of over US$2 trillion just weeks away, and arch-rival OpenAI’s listing not too far behind, this week I will tackle the economics of frontier AI labs and hyperscalers.
Clearly, frontier labs and hyperscalers are two sides of the same economic machine. Frontier labs are trying to turn compute, talent, data and algorithms into intelligence, and then sell that intelligence. Hyperscalers such as Amazon, Google, Microsoft, Meta Platforms and Oracle, on the other hand, are trying to turn capital, electricity, GPU or graphics processing unit chips, and data centres into compute, and then sell that compute, or use it themselves, to defend higher-margin services. With over US$880 billion being spent by hyperscalers on AI infrastructure this year and nearly US$1.35 trillion in capital expenditure (capex) forecast for next year, how do the economics of hyperscalers and frontier labs work? What does it cost to build and run a frontier lab and scale it? Will bleeding-edge frontier AI labs ever be profitable? And will hyperscalers that once used their own free cash flow to pay for their capex needs and now rely exclusively on borrowings to grow, make money as compute costs drop dramatically, profits accrue mostly to AI labs and competition escalates?
Let me first break down the economics of frontier AI labs. What Anthropic and OpenAI spend on is training compute, inference and reasoning compute, very expensive talent and, of course, data they license from sources such as print media like The Wall Street Journal and The New York Times and social media platforms such as Reddit, or from scanning tens of millions of old, out-of-print physical books. Frontier labs basically build a very expensive machine, then rent its intelligence by the tasks or tokens in the hope that the value they create is a lot higher than the total compute cost per token.
It takes enormous sums of capital to become a leading frontier lab. Anthropic, the world’s most valuable lab founded in 2021, has so far raised US$132 billion in capital and accumulated US$76 billion in debt. It is currently trying to raise up to US$35 billion in debt ahead of its record-breaking IPO next month, which will raise another US$100 billion or so of equity. OpenAI has raised over US$184 billion in equity so far and has hundreds of billions in debt. I should add a caveat here. The vast majority of capital raised by frontier AI labs, more than 70% by some estimates, is through cloud-compute credits from Microsoft’s Azure, Amazon’s AWS and Google Cloud, rather than pure cash. But whether someone gives you land and machinery as their part of the stake in a factory you are building, or just hands you cash, it still counts as equity.
A frontier lab spends between 60% and 70% of its total budget on compute. Top labs also spend billions more on infrastructure and experimentation, including building out the GPU clusters, managing complex data pipelines and running smaller test models to optimise hyperparameters. Once a model is deployed, inference shifts from a capital expense to a massive operational expense. The final, full-scale training run currently costs over US$1 billion but is likely to grow to tens of billions within three years. Research firm EpochAI notes that the final run accounts for only a small portion of a lab’s total research and development compute, or around 10% for labs such as OpenAI and Anthropic. Yet a lab spending US$1 billion on a product this year is probably spending US$9 billion more on research and development (R&D) as well as experimentation. So the tens of billions in annual product spend forecast for 2029 that I mentioned earlier means top AI labs could be spending nearly US$100 billion on R&D.
A bigger cost these days is inferencing and reasoning. While training is a big one-time cost, inference is a variable cost that grows with every new user. Unlike software where serving more users costs nothing, each AI query burns real compute. So more users means higher costs. Inferencing helps the model read your questions or prompt and produces a direct answer. That requires more than a few tokens because with reasoning, the model first generates a long internal reasoning trace, or thousands of hidden “thinking” tokens, then answers your question. As more people adopt AI, more people ask questions, so we need more tokens, the main unit of text or data, such as a word, that an AI model reads, processes and generates. With more inferencing and better reasoning, we are talking about anything from 10 to nearly 100 times more tokens to answer more queries. A reasoning query is far more expensive than a simple one because the model burns a large multiple of tokens thinking.
Now let’s tally up what an AI lab would spend — at least 100,000 H200-equivalent GPUs worth between US$3.5 and US$4 billion just for training data, or up to US$4.5 billion for fully integrated service clusters. Of course you could just rent 100,000 GPU chips for up to US$4.50 per GPU hour, or an annual rental cost of US$2.1 billion per year, if you are running them continuously at full capacity. As I mentioned in my previous piece, a GPU chip’s lifespan is just four years, so you need to weigh that US$2.1 billion annual rental cost against the US$4.5 billion total ownership cost, which allows the lab to sell GPU chips in the second-hand market.
Unlike software-as-a-service (SaaS) firms that make a software once, then just collect an 80% margin forever, frontier AI labs like Anthropic or OpenAI need to build very expensive software like Claude Fable 5.1 or GPT-6 Astra, monetise it immediately and spend a huge chunk of the proceeds to build the next generation of AI model. So, a frontier AI lab is hugely capital-intensive or akin to a high-end chip foundry like Taiwan Semiconductor Manufacturing, or TSMC, which makes a lot of money and then ploughs it all back into next-generation nodes.
How do AI labs make money? There are three main revenue streams. First is consumer subscriptions. OpenAI’s ChatGPT Plus charges US$20 a month while SpaceXAI’s Grok, through an X Premium subscription, costs US$8 every month. OpenAI has nine million ChatGPT Plus subscribers while SpaceXAI has 1.9 million paying premium subscribers for Grok. The model is simple: high volume, low price, easy to cancel. AI labs also make money from Application Processing Interface or API access, where developers pay per token. It’s often called the “compute wholesale” business. Lastly, there are enterprise contracts, which Anthropic dominates, where large companies pay a lot of money for Claude access with security, support and compliance. That’s a much more high-value, sticky and high-margin business. Anthropic also has a consumer subscription business, but that’s a small part of its total revenue stream.
Anthropic had a current annualised run rate of US$65 billion at the end of July, compared with OpenAI’s annualised run rate of US$40 billion at end-August, while xAI last reported an annualised run rate of just US$0.5 billion. The Claude maker is projected to generate over US$15 billion in revenue and an operating profit of around US$1 billion during the current quarter.
Let’s now look at the economics of hyperscalers, which are basically sellers of shovels and not gold miners themselves. They rent out compute and sell software and ads. Each has a huge, profitable legacy cash cow funding the AI beta, such as search ads for Google, software for Microsoft, e-commerce for Amazon, and social media ads on Instagram and Facebook for Meta Platforms. The three big hyperscalers Amazon, Microsoft and Google started out as cloud infrastructure players. They would buy servers, build data centres, buy networking gear, provide software and rent compute. With the advent of AI, they need far more capital. Amazon will spend US$200 billion on capex this year. Seven years ago, in 2019, it spent just US$12.7 billion. So, capex is up more than 15-fold over seven years.
A cloud service provider like AWS or Azure made money by sharing. They bought enough hardware to cover average demand across thousands of customers. The machines stayed busy because none of the customers spiked at the same time. Before November 2022 when OpenAI unveiled ChatGPT, marking the start of the AI era, clouds ran on general-purpose CPU servers pooled across thousands of customers. A CPU has a handful of powerful cores that handle complicated instructions one after another, very quickly. That is what running an operating system, serving a web page or answering a database query needs. A GPU, on the other hand, has thousands of simple cores that all perform the same arithmetic at the same moment on different numbers. It was designed to share millions of pixels at once.
AI changed both the workload and hardware. An AI cluster includes GPU accelerators, high-bandwidth memory, high-speed networking, huge power systems, back-up generators, cooling, fibre optics, substations and long-term power contracts. One training run includes thousands of GPUs in lockstep for weeks, leaving no idle slice for anyone else. A GPU runs the same arithmetic across enormous grids of numbers at once, which demands short high-bandwidth links between chips and 40kW to 250kW per rack, where an ordinary rack draws 10kW to 15kW per rack. That has made specialised cooling and networking essential for all data centres.
Cloud service providers traditionally made themselves indispensable because they took enormous fixed infrastructure costs and turned it into a variable expense for customers. Microsoft’s Azure would buy a server. Whether you ran a small business, or were the CEO of a larger one, you could easily rent it on an hourly basis. AI changed it completely because the machines are much more expensive and much more specialised. Moreover, customers increasingly have enough utilisation so that owning the hardware can sometimes make economic sense.
There is one other reason why hyperscalers are hot: scarcity. If AI compute were abundant, cloud economics wouldn’t look as good as it does right now. But if demand constantly exceeds supply, capacity becomes scarce and pricing stays high, which means utilisation stays high and returns on infrastructure improve. While people are increasingly questioning the sustainability of AI infrastructure spending, it is clear that demand still exceeds supply.
For their part, hyperscalers have moved to slash costs. Since the biggest cost in the AI buildout is the GPUs and Nvidia has gross margins of 75%, their key focus has been to make their own chips and avoid paying a 75% tax to the chip behemoth. Amazon has its own Trainium chips, whose newest iteration almost matches the performance of Nvidia’s top-of-the-line Blackwell GPUs. Google has its own TPUs, or Tensor Processing Units; Microsoft has its Maia chips; and then there are MTIA, or Meta Training and Inference Accelerator chips. Most of these chips are being used for inferencing and reasoning as well as stable, internal workloads. Hyperscalers remain the biggest customers of Nvidia and still use its chips for cutting edge AI training. OpenAI and Anthropic are also designing their own chips, though they too rely on Nvidia for AI training chips. The chip giant’s annualised revenues doubled in the last quarter.
Here is what you should take away from the economics of frontier labs and hyperscalers. The frontier labs are betting that intelligence itself will eventually become a scarce and valuable resource. The hyperscalers, on the other hand, are betting that the infrastructure required to make intelligence remains scarce. And the company at the centre of it all, chipmaker Nvidia, is betting that the machines that make intelligence remain the bottleneck. That way, there will be something for everyone and the AI boom will continue for a few more years.
Assif Shameen is a technology and business writer based in North America
Save by subscribing to us for your print and/or digital copy.
P/S: The Edge is also available on Apple's App Store and Android's Google Play.