Beyond NVIDIA: The Custom Silicon Quietly Filling AI Data Centers
Google, Amazon, Meta, Microsoft, and OpenAI are pouring billions into their own AI accelerators. Custom chips now claim more than a quarter of AI server shipments, even as NVIDIA posts record revenue.
Walk the halls of a new hyperscale campus in 2026 and the odds are still good that the racks hum with NVIDIA hardware. But the odds are shrinking. In machine rooms from Indiana to Iowa, a growing share of the accelerators drawing all that power were never sold by anyone: Google’s TPUs, Amazon’s Trainium, Meta’s MTIA, and Microsoft’s Maia are designed in-house, built by TSMC, and deployed by the million. According to TrendForce data reported by Tech Times, custom ASIC shipments are growing about 44.6 percent in 2026, nearly triple the 16.1 percent pace of merchant GPUs.
The story of AI infrastructure is no longer just NVIDIA’s story. It is a story about vertical integration, and about who keeps the margin.
Google’s decade-long head start
Google has been building Tensor Processing Units since 2015, and its seventh generation, Ironwood, is the first the company markets primarily for inference. Each Ironwood chip delivers 4,614 teraFLOPS of FP8 compute with 192 GB of HBM3E memory, according to Google, and pods scale to 9,216 chips wired together with a 9.6 Tb/s inter-chip interconnect. Google says Ironwood roughly doubles the performance per watt of its predecessor, Trillium, and is nearly 30 times as efficient as the first Cloud TPU from 2018.
The scale is what separates Google from every other custom-silicon effort. SemiAnalysis has called TPU v7 the first credible merchant-scale challenge to NVIDIA, and the customer list now backs that up. In October 2025, Anthropic announced a deal for access to up to one million TPUs, with more than a gigawatt of capacity arriving in 2026, a commitment the company valued in the tens of billions of dollars. In April 2026, Anthropic, Google, and Broadcom expanded that arrangement again, adding multiple gigawatts of next-generation TPU capacity expected to come online starting in 2027.
Amazon’s bet on Trainium, and Anthropic again
Amazon’s answer is Project Rainier, the training cluster it built with and for Anthropic. Activated in late 2025, Rainier spans data centers across the US, and according to Amazon it contains nearly 500,000 Trainium2 chips, with a goal of surpassing one million by the end of that year for training and serving Claude. AWS claims Trainium2 instances deliver 30 to 40 percent better price performance than its comparable NVIDIA-based EC2 offerings, a figure worth noting is Amazon’s own benchmark rather than an independent one.
Trainium3 arrived at re:Invent 2025 in UltraServers of 144 chips, promising roughly four times the performance of Trainium2 and about 40 percent better energy efficiency, per Amazon. Anthropic, notably, now sits on both sides of the custom-silicon race, committed to millions of chips split between Amazon’s Trainium fleet and Google’s TPUs.
Nearly 500,000 Trainium2 chips power Project Rainier, according to Amazon, making one company’s custom cluster larger than most national supercomputing programs combined.
Meta and Microsoft play catch-up
Meta’s MTIA program started narrow, accelerating the ranking and recommendation models behind its ad business. It is now sprinting. In March 2026, Meta disclosed a roadmap of four new MTIA chips co-developed with Broadcom on a six-month cadence, according to Meta and Tom’s Hardware: MTIA 300 is already in production for recommendations training, MTIA 400 is in lab testing, and the inference-focused 450 and 500 are slated for mass deployment across 2027.
Microsoft has had the rockiest road. Its next-generation Maia part, code-named Braga, slipped by at least six months after design changes, including features requested by OpenAI, destabilized the project, according to reporting cited by Data Center Dynamics. The chip finally landed in January 2026 as Maia 200: a TSMC 3 nm design with more than 140 billion transistors delivering over 10 petaFLOPS of FP4 compute in a 750 W envelope, per Microsoft. OpenAI, meanwhile, went further than requesting features. Its October 2025 pact with Broadcom covers 10 gigawatts of OpenAI-designed accelerators, with rack deployments starting in the second half of 2026 and running through 2029, according to the companies.
Why build your own chip
Three motives recur. The first is cost: NVIDIA’s data center gross margins have hovered above 70 percent, and at hyperscale, that markup is a tax measured in billions. A chip that only needs to run transformer inference well can skip the general-purpose overhead a GPU carries. The second is supply: after two years of allocation battles for H100s and Blackwells, owning a second source is insurance. The third is power. With campuses now planned in gigawatts rather than megawatts, performance per watt is the binding constraint, which is why Google, Amazon, and Microsoft all lead their chip announcements with efficiency claims rather than raw speed.
| Company | Chip (2026) | Status |
|---|---|---|
| TPU v7 Ironwood | GA, pods of 9,216 chips | |
| Amazon | Trainium2 / Trainium3 | ~500K chips in Project Rainier; Trn3 ramping |
| Meta | MTIA 300 / 400 | In production / lab testing |
| Microsoft | Maia 200 | Announced January 2026 |
| OpenAI | Broadcom co-design | Deployments from late 2026 |
The moat holds, for now
None of this has dented NVIDIA’s income statement. The company reported $193.7 billion in data center revenue for fiscal 2026, up 68 percent year over year, including a record $62.3 billion in its final quarter. Its moat is less the silicon than the software and the fabric: nearly two decades of CUDA libraries mean most inference optimizations ship on NVIDIA first, and NVLink’s 1.8 TB/s per-GPU bandwidth underpins rack-scale systems rivals are still learning to build. Porting a production workload to Trainium or TPU remains weeks to months of engineering effort.
So what share of AI compute is custom in 2026? Estimates diverge, and honest analysts say so. TrendForce puts ASIC-based systems at about 27.8 percent of AI server shipments this year, while revenue-based estimates cited by industry trackers such as Introl and Silicon Analysts still credit NVIDIA with anywhere from 70 to 86 percent of the accelerator market, depending on whether hyperscalers’ internal chips are counted at cost or at market value. The trend line, though, points one direction. NVIDIA’s slice of a rapidly growing pie keeps shrinking at the edges, and every hyperscaler has decided the same thing: the most expensive chip in the data center is the one you have to buy from someone else.