For most of the AI boom, the hardware story had one protagonist: a single class of graphics processor, made by a small number of vendors, that everyone - cloud giants, AI labs, startups, universities - queued to buy. That queue defined the industry's economics, and the vendors at its head captured historic margins. The next chapter is now visibly under way: the largest AI developers, the very customers who spent most heavily in that queue, have been designing their own inference chips, and the first of these custom parts are being reported by their makers as beating the leading merchant hardware on the measure that increasingly matters - useful AI work per watt of electricity - for the workloads they were designed around. The claims deserve the usual caution owed to benchmarks chosen by the party they flatter. The direction does not: the biggest buyers are becoming makers.
The logic of the move is textbook vertical integration, sharpened by two AI-specific facts. First, inference - serving a trained model to millions of users - is now the dominant, permanent cost of running an AI business, and unlike research it is a stable, well-understood workload, exactly the kind a custom chip serves best. A general-purpose processor pays a flexibility tax to be good at everything; a chip built for one company's models pays none. Second, the constraint on AI expansion has shifted from money to electricity - data-centre power is the scarce input - so a chip that does the same work on fewer watts is not merely cheaper but expands what is buildable at all. When efficiency per watt decides capacity, designing your own silicon stops being a cost project and becomes a growth project.
What this does to the broader market is the part worth watching from outside. The merchant chip vendors do not lose their biggest customers overnight - custom chips take years, cover only some workloads, and training remains largely on general-purpose hardware - but they do lose pricing power at the top of their customer list, and history suggests that erosion travels. Meanwhile the queue behind the giants changes shape: capacity at the contract fabs that actually manufacture all these chips becomes the contested resource, and every custom design competes for the same advanced production slots. For everyone who buys AI as a service rather than as hardware, the plausible medium-term effect is benign: more supply, more competition, and inference priced closer to its falling cost - the same pattern that played out when the largest cloud companies built their own servers and networking a decade ago.
The practical takeaways for an ordinary business are modest but real. Expect continued deflation in the price of AI inference, and be sceptical of any plan - your own or a vendor's - that assumes today's AI unit costs are stable; they are on a downward slope that vertical integration is steepening. Notice that the chips underneath your AI services are diversifying, which is quiet good news for resilience: a market that depended on one processor family had a single point of failure that the industry is now engineering away. And file the episode under a lesson older than computing: when one supplier's margins define an entire industry's cost structure, the largest customers eventually build their own. It happened to mainframes, to telecoms equipment, to servers. It is happening to AI chips on schedule.