Artificial intelligence is usually discussed as if access to it were limited mainly by price. A company chooses a model, pays for an API or rents cloud infrastructure, and receives as much machine intelligence as it can afford.
That description works reasonably well when computing capacity is abundant.
It becomes less accurate when the scarce resource is not software but the physical capacity underneath it: advanced chips, power, cooling, data-center space and access to large GPU clusters. In that world, the relevant question is no longer simply what AI costs. It is who can secure enough computing capacity, for long enough, and under what rules.
This is not a hypothetical problem waiting for some distant energy crisis. Amazon Web Services now tells customers directly that demand for GPU-based machine-learning capacity has outpaced industry-wide supply and describes GPUs as a scarce resource. Its response is not a rationing committee but a commercial mechanism: customers can reserve GPU capacity in advance through products such as EC2 Capacity Blocks. (aws.amazon.com)
Governments are building their own allocation systems as well. The European Union’s AI Factories offer different access tracks for startups, small and medium-sized companies, scientific researchers and larger AI projects, with some access granted on a first-come basis and large allocations subject to technical and peer review. The United Kingdom’s AI Research Resource has offered projects between 50,000 and 1.4 million GPU hours from a finite public pool, with specific calls directed toward areas such as medical research, fusion, materials science, engineering biology and scientific discovery. (eurohpc-ju.europa.eu) (gov.uk)
These arrangements are very different from one another, but they reveal the same underlying fact. Once high-end computing capacity becomes constrained, access is allocated somehow. The mechanism may be a market price, a reservation contract, ownership of infrastructure, a government program, a peer-review panel or simply geography.
The allocation already exists. What remains unsettled is how consequential it will become.
Compute scarcity is not the same everywhere
There is no single global pool of AI computing power.
The OECD’s 2025 work on public-cloud AI compute found that seven major cloud providers operated 531 availability zones, but only 351—about 66 percent—offered at least some GPU capacity suitable for AI workloads. The United States and China together accounted for 148 AI-capable zones. Public cloud is also only one part of the market; private clusters and government supercomputers can be much more restricted and are harder to measure. (oecd.org)
At other layers of the AI infrastructure stack, concentration is even stronger. An OECD competition study published in 2025 described multiple segments of the supply chain as highly concentrated, including advanced lithography, advanced chip fabrication and GPUs. The OECD warned that shortages in chips, energy, networking or water could slow downstream innovation and create risks of resource hoarding or preferential allocation, particularly where firms control both infrastructure and downstream AI services. (oecd.org)
That does not prove that cloud providers are systematically denying capacity to smaller competitors. The OECD explicitly treats the issue as a competition risk rather than a universal description of current conduct.
The distinction matters because scarcity can take several forms at once. A chip may exist but not be available in the required cloud region. GPU servers may be installed but waiting for a grid connection. A data center may have power but insufficient networking or cooling capacity. A company may technically be able to rent GPUs but only at a price or contract duration it cannot finance.
“Who gets compute?” is therefore not answered at one gate. It is answered repeatedly throughout the infrastructure stack.
In commercial markets, money buys more than compute
The obvious allocation mechanism is price.
But AI infrastructure markets increasingly sell something more valuable than raw GPU-hours: certainty.
AWS Capacity Blocks, for example, let customers reserve clusters of accelerated computing capacity in advance for specified periods. In its U.S. government cloud, AWS says customers can reserve capacity up to eight weeks ahead and for durations of up to six months. The product exists precisely because on-demand availability cannot always be assumed. (aws.amazon.com)
For a large company, a long reservation or multiyear infrastructure agreement can transform an uncertain future supply problem into a contractual asset. For a smaller company, the same strategy may require capital it does not have.
This is one reason compute scarcity is different from an ordinary shortage of consumer goods. A wealthy buyer does not merely buy more units. It can purchase priority, duration and predictability.
Recent infrastructure deals illustrate the scale at which some AI companies are trying to secure that predictability. Reuters reported in July 2026 that AI startup Reflection had signed a computing agreement worth more than $1 billion with Nebius after entering another major capacity agreement the previous month. The point is not that every AI company needs billion-dollar contracts; most do not. It is that the frontier end of the industry increasingly treats future computing capacity as something worth locking down years in advance.
That dynamic can reinforce differences between firms. A well-capitalized model developer can reserve large amounts of future compute, negotiate favorable infrastructure deals and spread the cost across a large customer base. A smaller company may buy whatever remains available on shorter contracts or redesign its product around cheaper models.
Market allocation is efficient in one obvious sense: capacity tends to flow toward buyers willing to pay the most for it.
Whether willingness to pay is always the same thing as social value is a different question.
Ownership changes the allocation problem
The most powerful AI companies increasingly own or control significant parts of the computing stack themselves.
This changes the meaning of scarcity because a vertically integrated company does not have to enter the same queue as an outside customer for every unit of compute. It can decide internally whether scarce capacity should support model training, inference for paying customers, research, advertising systems or another corporate priority.
The OECD’s 2026 report on AI markets identifies this concentration as a structural feature of the emerging industry. It notes that cloud computing and specialized-chip production are highly concentrated and that complementary infrastructure—including data centers, electricity generation and grid connections—can themselves become sources of market power. (oecd.org)
Again, concentration does not automatically mean abuse. Vertical integration can lower transaction costs, accelerate investment and allow infrastructure to be used efficiently. The same companies with the greatest internal demand are often the companies investing most aggressively in new capacity.
But ownership gives a firm a form of optionality that an outside customer does not have. When resources tighten, an owner can decide which internal workload matters most before deciding what capacity remains available to others.
The allocation process may therefore occur inside corporate planning meetings long before any government begins debating whether AI should be formally rationed.
Public compute makes the value judgments visible
Government-funded computing systems expose the allocation problem more clearly because public institutions cannot simply say that the highest bidder wins.
Europe’s AI Factories are a good example. Their industrial access programs include a small “Playground” route for entry-level users, a Fast Lane route for medium-sized projects requiring up to 50,000 GPU hours and a Large Scale route for applications needing more than 50,000 GPU hours. Smaller requests can be processed rapidly, while large allocations require technical and peer-review evaluation and are intended for high-impact, high-gain projects. (eurohpc-ju.europa.eu)
The underlying EU regulation goes further. It states that certain Union access time should be used primarily to provide access to startups and SMEs for research and innovation, and that some capacity may be used to give free access to projects developing open frontier models selected through Union-wide competition. (eur-lex.europa.eu)
The United Kingdom has adopted another version of the same principle. Its AI Research Resource exists in part because the government describes advanced public AI compute as acutely underprovided for research and development. In one 2026 open-access call, up to eight million GPU hours were available on the Isambard-AI supercomputer, with individual projects able to request between 50,000 and 1.4 million GPU hours. Successful projects were subject to usage quotas intended to maintain predictable and fair access. (gov.uk)
These programs do not establish a universal hierarchy for AI.
They show something more important: once compute is publicly financed and finite, allocation becomes an explicit policy decision. Eligibility, project size, scientific merit, industrial policy and public purpose all become part of the answer to the question of who gets access.
Commercial markets hide much of this value judgment inside price. Public programs have to write it down.
Energy constraints add another allocation layer
The first article in this series examined how AI dependence intersects with electricity. The relevant point here is narrower: the grid can constrain computation even when chips and data centers are available.
The International Energy Agency’s 2026 work on electricity-system flexibility describes demand response as a mechanism through which large consumers can shift or reduce electricity use in response to grid or market conditions, often in exchange for financial incentives. Data centers are among the new concentrated loads that electricity systems increasingly have to integrate. (iea.org)
AI workloads may be unusually interesting because some forms of computation can move in time or location. A field demonstration published in Nature Energy used a 256-GPU cluster in a hyperscale cloud facility and reduced power consumption by 25 percent for three hours while maintaining the experiment’s specified quality-of-service guarantees. The result came from one system and should not be generalized to all data centers or workloads, but it demonstrated that at least some AI computing can function as flexible demand rather than an inflexible industrial load. (nature.com)
This creates an important distinction between scarcity and failure.
If a data center can delay training, move a batch job or temporarily reduce a flexible workload, an energy shortage does not necessarily make AI unavailable. It changes the amount of compute that can be delivered at a particular time.
Someone—or some scheduling system—then decides which jobs wait.
Today that decision may be made automatically according to service tiers, contracts, technical requirements and internal business priorities. Under more serious electricity constraints, the same mechanisms could become much more consequential.
There is not currently a universal rule saying that medical AI must run before advertising, or that scientific research must receive power before consumer entertainment. Writing such a hierarchy as if it already existed would be fiction.
What exists today is the infrastructure from which such prioritization could emerge.
The global allocation is already unequal
The most important divide in compute access may not be between two companies in the same cloud region. It may be between countries.
The World Bank’s 2025 Digital Progress and Trends Report found that high-income countries hosted 77 percent of global colocation data-center capacity as of June 2025. Upper-middle-income economies accounted for 18 percent, lower-middle-income economies for 5 percent and low-income countries for less than 0.1 percent. High-income countries also hosted 97 percent of the computing capacity represented in the world’s top 500 high-performance computing systems. (worldbank.org)
The World Bank does not argue that every country needs to build frontier-scale data centers. Its 2026 World Development Report makes almost the opposite point: many developing economies can gain substantially from adopting and adapting smaller, cheaper AI systems rather than attempting to reproduce the infrastructure of the largest AI powers. But the report also identifies concentration in models, chips and data centers as a dependency risk that affects which AI capabilities countries can access and adapt to local needs. (worldbank.org)
This distinction matters. Unequal ownership of frontier compute does not imply unequal access to every useful form of AI. Cloud services, open models and smaller systems can spread capability much more widely than physical infrastructure ownership alone would suggest.
Yet physical concentration still matters when the most capable systems require resources that only a small number of countries and companies can provide.
Under abundant supply, that concentration may remain largely invisible to end users. Under scarcity, geography becomes part of the allocation mechanism.
Scarcity can also create public alternatives
The market is not the only response to concentration.
The European AI Factory system, the UK’s AI Research Resource and similar public-compute initiatives exist partly because policymakers do not want access to advanced computing to depend entirely on the balance sheets of large technology companies.
The policy objective is especially clear in Europe. EuroHPC describes its AI Factories and planned AI Gigafactories as infrastructure intended to expand access for researchers, public authorities, startups, SMEs and industry while strengthening European technological sovereignty. Nineteen AI Factories and 13 associated antennas are currently being implemented or supported across Europe. (eurohpc-ju.europa.eu)
This is not proof that public compute will be allocated better than private compute. Public systems can be bureaucratic, politically influenced or poorly utilized. Peer review can favor established institutions. Strategic industrial policy can protect weak domestic projects. Public resources also remain tiny compared with the infrastructure budgets of the largest technology companies.
But public compute introduces a different principle into the allocation system: access can be granted because an institution believes a project has scientific, industrial or social value, not only because the applicant can outbid competitors.
That makes allocation criteria visible enough to debate.
The strongest objection is that compute scarcity may not last
Any argument about AI compute scarcity has to confront the possibility that today’s shortage is transitional.
The industry is spending extraordinary amounts to expand capacity. Chips are becoming more efficient. Competition is emerging among cloud providers and specialized AI infrastructure companies. Smaller models can perform tasks that once required much larger systems, while quantization, distillation and improved inference software reduce the amount of hardware needed for some workloads.
One 2026 study in Economics Letters offers an unusually concrete example of how quickly beliefs about scarcity can change. The authors examined U.S. stock-market reactions to the release of DeepSeek-R1 in January 2025 and found that firms with greater prior exposure to supplying AI compute experienced lower abnormal returns around the event. Their interpretation is that investors viewed the arrival of a lower-cost AI approach as reducing the economic value of scarce upstream compute. The study measures market expectations rather than physical GPU availability, but it shows that scarcity rents can be repriced when technical assumptions change. (sciencedirect.com)
The World Bank makes a related argument from a development perspective. Many socially useful AI applications do not require frontier-scale computing at all. Smaller systems running on cheaper infrastructure or local devices can provide medical, educational and agricultural services in places that lack large data centers. (worldbank.org)
This is the strongest reason not to treat compute like oil.
Compute is manufactured capacity. Efficiency can improve. Workloads can be redesigned. New suppliers can enter. A breakthrough can reduce the hardware required for a given task.
Scarcity can therefore shrink even while demand rises.
What may persist is not permanent shortage but unequal ability to absorb periods of shortage.
The important question is who can buy certainty
A startup that cannot get its preferred GPU can often use a smaller model, wait for capacity or change providers. A research group can modify an experiment. A consumer can accept a slower response.
A company that has committed critical operations to AI has less flexibility.
This makes guaranteed access different from raw access. The valuable asset may not be owning the fastest chip but knowing that sufficient capacity will be available when it is needed.
Organizations can obtain that certainty in different ways: owning infrastructure, reserving cloud capacity, signing long-term contracts, maintaining multiple providers, running smaller local systems or gaining access to public compute.
Each option requires resources.
Money matters. Geography matters. technical expertise matters. Political eligibility can matter in public programs. Relationships with infrastructure providers can matter at the frontier.
The phrase “democratizing AI” can therefore become misleading if it refers only to whether a model is accessible through a web browser. The deeper question is whether different users can obtain dependable amounts of the computing capacity needed to build, adapt and operate AI on their own terms.
That is a much higher standard than being able to open a chatbot.
There is no neutral allocation system
Markets allocate capacity toward customers with the greatest willingness and ability to pay. Public systems allocate it through eligibility rules, strategic priorities and review processes. Vertically integrated companies allocate it internally. Electricity systems may reward or require flexibility when grids are stressed.
None of these mechanisms is inherently neutral.
That does not mean a central authority should decide how every GPU-hour is used. Such a system would create obvious problems of bureaucracy, political favoritism and technical incompetence.
It means that pretending no allocation decision exists simply leaves the decision embedded in infrastructure ownership, contracts and prices.
The OECD has already warned that shortages and concentrated infrastructure can create risks of preferential allocation and resource hoarding. The World Bank describes global compute capacity as deeply uneven. European and British governments are explicitly building public systems because commercial markets alone do not produce the access they want for researchers and smaller firms. (oecd.org) (worldbank.org) (gov.uk)
The allocation question is therefore not arriving someday after AI becomes infrastructure.
Parts of it are already here.
If AI becomes critical infrastructure, allocation rules will matter more
The problem becomes more difficult if organizations continue moving essential work into AI-dependent systems.
A temporary shortage of video-generation capacity is mostly an inconvenience. A shortage affecting a service embedded in medical administration, industrial control, cybersecurity or financial operations would raise different questions.
There is not yet evidence for a universal social hierarchy of AI workloads, and this article does not propose one. Different countries would make different choices, and in many situations market mechanisms could continue working normally.
The point is narrower.
As AI becomes more deeply embedded in valuable economic and public functions, society will care increasingly about not only how much compute exists but who controls it, who can reserve it, who can substitute for it and who loses access first when supply is constrained.
That makes compute access a resilience issue as well as an economic one.
The first article in this series argued that an AI-dependent society could become vulnerable when energy constraints reduce machine capacity. The second examined what happens when companies allow human fallback capabilities to disappear. This third question completes the chain: if machine capacity becomes scarce while dependence continues to grow, the allocation of that capacity will determine where the loss of capability is felt.
The answer will probably not come from a single government decree.
It will emerge from thousands of contracts, cloud reservations, grid rules, capital budgets, public-compute programs and infrastructure decisions made long before a crisis occurs.
By the time scarcity becomes visible to ordinary users, much of the allocation may already have been decided.
AI Dependency & Resilience
This essay is part of a three-part series on what happens when machine intelligence becomes ordinary infrastructure: how energy constrains it, what organizations lose when human fallback disappears, and how scarce compute is allocated.
01 — Energy & Availability
When Society Depends on AI, Can It Withstand an Energy Crisis?
02 — Human Fallback
What Happens When Companies Forget How to Work Without AI?
03 — Compute Allocation
When AI Compute Becomes Scarce, Who Gets It First?