Special Report · AI Infrastructure · October 10, 2026
Custom silicon is becoming a strategic tool for controlling AI infrastructure economics — not simply a way to design a faster chip. Hyperscalers want greater control over performance, power consumption, supply availability, and the cost of serving AI workloads at scale.
This report examines why the largest cloud and technology companies are moving beyond a GPU-centric architecture, what that shift does to the economics of AI computing, and which semiconductor suppliers capture the most durable value as a result.
Evidence standard. Figures in this report are labeled as disclosed (company announcements, earnings releases, technical documentation) or estimate (attributed analyst research, industry reporting, or CODEW-derived analysis). Forward-looking statements are identified as such and should not be read as reported results.
For most of the last decade, the question a hyperscaler asked about AI hardware was simple: how many GPUs can we get, and how fast? That question assumed a market structure in which one vendor supplied the compute layer and everyone else built on top of it.
That assumption no longer holds. Google, Amazon, Meta, and Microsoft have each committed to designing their own AI accelerators, and each has moved those programs from experimentation into production. The shift is not driven by a belief that these companies can out-design the incumbent GPU vendor on raw performance. It is driven by something more consequential: the realization that at sufficient scale, the architecture of the compute layer becomes a control point over the economics of the entire business.
Custom silicon is best understood as vertical integration applied to compute. When a company designs its own accelerator, it gains control over four things it otherwise negotiates for: performance characteristics, power envelope, supply availability, and cost per unit of work. None of those are chip-design problems in the conventional sense. They are business-model problems, and the chip is the instrument for solving them.
The consequence is that this is not a story about whether custom chips beat GPUs. It is a story about how value redistributes across the AI infrastructure stack when the largest buyers stop being purely buyers.
The CODEW Lens: Hyperscalers are not trying to win the chip market. They are trying to stop paying someone else's margin on the largest cost line in their business. Those are different objectives, and only the second one explains the scale of the investment.
1. Why the GPU-Only Strategy Stopped Being Enough
A general-purpose GPU is a bundle. The buyer gets compute, a software ecosystem, networking, and a vendor roadmap — priced together and delivered on the vendor's schedule. For a company running experimental workloads at modest scale, that bundle is excellent value. Flexibility is worth paying for when you do not yet know what you need.
Three things change as AI becomes a core operating expense rather than a research budget.
First, the workloads stabilize. Recommendation and ranking models, search ranking, ad targeting, and increasingly inference serving have well-understood computational shapes. They do not need the full flexibility of a general-purpose processor. A workload that is stable and high-volume can be served more efficiently by hardware optimized for it.
Second, the volume becomes large enough to justify the fixed cost. Designing an advanced-node accelerator carries a substantial non-recurring engineering cost, plus the cost of a software stack, plus the organizational cost of building a silicon team. At low volume, that investment never pays back. At hyperscale volume, the arithmetic inverts.
Third, supply becomes a strategic variable rather than a procurement detail. A company that depends entirely on one accelerator supplier for its AI roadmap has ceded control over its own deployment schedule. Diversification is not only about price. It is about being able to build when you need to build.
Together, these three conditions describe the position every major hyperscaler now occupies. The move to custom silicon is not a bet on a technology trend. It is the predictable response of very large buyers to a very large, concentrated cost line.
The CODEW Lens: The GPU bundle is priced for flexibility. Hyperscalers stopped needing flexibility on their largest workloads years ago. The custom silicon race is what happens when a buyer outgrows the bundle it is being sold.
2. The Economics: When Does Designing Beat Buying?
This is the analytical question that determines whether custom silicon is a permanent feature of the market or a phase. The answer is not a fixed threshold. It is a relationship between four variables.
Non-recurring engineering cost — design, verification, IP licensing, masks, and the software stack that makes the chip usable.
Volume — total units deployed over the program's life, not units in the first quarter.
Efficiency gain — the performance-per-dollar or performance-per-watt advantage the custom design delivers on the target workload.
Opportunity cost — the engineering capacity consumed by the program, and the risk of being wrong about the workload.
The crossover point is where cumulative savings from the efficiency gain exceed the non-recurring cost plus the opportunity cost. This is why the arithmetic works differently for different companies and different workloads. A hyperscaler deploying hundreds of thousands of accelerators across a stable workload will reach that crossover. A mid-sized enterprise with variable workloads will not — and should keep buying the bundle.
There is a second-order effect that is frequently overlooked. Custom silicon does not eliminate cost; it relocates it. A hyperscaler that designs its own accelerator still pays for fabrication, memory, advanced packaging, and networking — and it now also pays for design. What it removes is the merchant vendor's margin on the compute die and, more importantly, the vendor's control over the roadmap. That control is worth more than the margin at hyperscale, because it determines when and how fast the company can deploy.
The framing that matters: the question is not "is custom silicon cheaper than a GPU." It is "at what volume, on what workload, and over what time horizon does the fixed investment pay back — and who bears the risk if the workload assumption is wrong." (Framework — CODEW-derived)
Note that none of these companies discloses the non-recurring cost of its accelerator programs, and published analyst estimates vary widely by node, program scope, and whether software development is included. Any specific figure presented as fact should be treated with caution. (Estimate — analyst research, not disclosed)
The CODEW Lens: Custom silicon is not a cost-reduction story. It is a control story with a cost-reduction side effect. The companies pursuing it are buying optionality on their own roadmap, and the price is a fixed investment they cannot recover if the workload changes.
3. Why One Chip Cannot Serve Every Workload
The shift toward custom silicon is inseparable from the divergence of AI workloads. Training, inference, recommendation, and agentic systems place different demands on hardware, and those demands pull chip design in different directions.
| Workload | Primary constraint | Design implication |
|---|---|---|
| Training | Raw throughput and interconnect bandwidth | Large die, high memory bandwidth, scale-out fabric |
| Inference | Latency, cost per token, power efficiency | Smaller die, aggressive quantization, memory hierarchy optimization |
| Recommendation/ranking | Throughput per watt at massive sustained scale | Specialized memory access patterns, embedding-heavy design |
| Agentic AI | Mixed workload, frequent control flow, memory capacity | Less deterministic; favors flexibility over specialization |
The table explains an important detail in the disclosed roadmaps. Google's eighth-generation TPU family includes separate designs for training and inference — a structure that only makes sense if the two workloads have genuinely diverged at the silicon level. (Disclosed)
It also explains why agentic workloads are the least likely candidate for near-term specialization. Control-flow-heavy, less predictable workloads benefit from general-purpose flexibility. A chip optimized for one inference pattern may be poorly suited to a workload that changes shape between releases.
The CODEW Lens: Workload divergence is what makes the hybrid outcome inevitable. There is no single architecture that serves a stable ranking model and an exploratory agent equally well — which is why the market is segmenting rather than consolidating.
4. Four Programs, Four Different Bets
Google — The Vertically Integrated Model
Google is the longest-running experiment in hyperscaler custom silicon and the only one that has operated at scale across multiple hardware generations. Its TPU program predates the current AI cycle, which gives it something the others lack: a track record of iterating a custom architecture in production.
The distinctive feature of Google's approach is vertical integration across the full stack. The TPU roadmap connects to Google Cloud, which connects to the Gemini model family, which connects to Google's own product surface. That means Google can optimize hardware for workloads it owns end to end — an advantage no merchant vendor can replicate, because a merchant vendor must serve many customers with divergent needs.
Google's disclosed eighth-generation TPU family includes separate designs for training and inference. (Disclosed) That split is the clearest public signal that workload divergence has reached the silicon level even within a single vendor's roadmap.
Notably, Google does not pursue a TPU-only strategy. It continues to offer third-party GPUs alongside its own silicon — a pragmatic position that acknowledges its internal chips serve specific workloads well without covering every use case its customers bring.
The CODEW Lens: Google's advantage is not better chips. It is that Google owns the model, the workload, and the deployment — which means it can afford to specialize further than any competitor whose chips must serve someone else's software.
Amazon Web Services — Building an Alternative Compute Platform
AWS's custom silicon strategy is organized around Trainium for training and Inferentia for inference — a two-track approach that mirrors the workload split visible in Google's roadmap. (Disclosed) Trainium2 instances reached general availability, moving the program from limited access to a broadly available product.
AWS's stated objective is straightforward: lower the cost of running AI workloads on its cloud. The strategic logic differs from Google's in an important respect. Google designs chips for workloads it controls. AWS designs chips for workloads its customers control — which means it must win developer adoption, not just internal deployment.
That is the central challenge of the AWS program. A chip can be cheaper per unit of work and still lose if the cost of porting software and retraining engineers exceeds the savings. AWS's approach has been to build the alternative stack alongside the incumbents rather than in place of them, offering custom silicon as one option among several.
The CODEW Lens: AWS has the hardest version of this problem. It must persuade customers to move workloads onto unfamiliar hardware. Google only has to persuade itself.
Meta — Optimizing AI at Enormous Scale
Meta's MTIA program is the clearest example of custom silicon justified purely by internal scale. Meta does not sell cloud compute. It runs recommendation and ranking models across billions of daily interactions — an enormous, stable, well-understood workload with a direct connection to revenue.
Meta states that MTIA has been deployed at scale for recommendation and ranking workloads. (Disclosed) That is a meaningful claim, because it means the program has moved past pilot status into production serving. It also describes the ideal case for specialization: high volume, stable workload, and measurable efficiency gains.
Meta's expanded partnership with Broadcom spans chip design, advanced packaging, and networking (disclosed) — a scope that illustrates how the enabler relationship works in practice. Meta defines the workload and the architecture. Broadcom implements it and supplies the surrounding silicon. This is the co-development model examined in detail in the CODEW Company Deep Dive on Broadcom.
The CODEW Lens: Meta's program works because the workload and the revenue are directly connected. Not every company can point to a specific business metric that a chip improvement moves. Meta can.
Microsoft — Custom Silicon for Azure AI
Microsoft's Maia program follows an inference-focused strategy, oriented toward improving the economics of serving AI on Azure and across Microsoft's AI products. The emphasis on inference rather than training reflects a reasonable bet: inference is the workload that scales with usage, and usage is where the recurring cost sits.
Microsoft occupies a distinctive position among the four. It is simultaneously a major customer of third-party accelerators, an operator of a large cloud platform, and a company whose AI products are deeply tied to an external model provider. That combination makes its internal silicon strategy less about displacing suppliers and more about establishing a credible alternative for the workloads where it makes economic sense.
Microsoft has positioned Maia as offering better performance per dollar than competing alternatives (company claim — not independently verified). Claims of this kind are common in silicon announcements and should be treated as directional until benchmarked by third parties on production workloads.
The CODEW Lens: Performance-per-dollar claims in press releases describe the best case on a favorable workload. The number that matters is performance per dollar on the workload the customer actually runs.
5. Software Is the Real Constraint
Every custom silicon program eventually encounters the same obstacle, and it is not transistor density. It is the cost of moving software.
A chip is only useful if developers can write for it. That requires compilers, libraries, kernels, profiling tools, debugging support, and documentation. It requires that existing frameworks work without modification, or that the modifications are worth making. It requires a community of engineers who have used the hardware before and can help others use it.
This is the incumbent's strongest defense, and it is not a technical moat so much as an accumulated one. A mature software ecosystem represents years of engineering investment by thousands of people, and it compounds. Every library that supports the incumbent makes it marginally easier to keep using it. Every hour a developer spends learning an alternative is an hour not spent on their actual problem.
The hyperscalers' structural advantage here is that they do not need to win the developer market to succeed internally. Google can port its own models. Meta can optimize its own recommendation stack. The software cost is real, but it is paid once by a company that owns both the hardware and the workload, rather than repeatedly by external developers evaluating an unfamiliar platform.
This is why internal deployment leads external adoption, and why AWS's challenge is structurally harder than Google's or Meta's. Selling custom silicon as a cloud product means asking customers to absorb a software migration cost that the hyperscaler absorbed internally.
The CODEW Lens: The performance gap between custom and general-purpose silicon is a design problem, and it is being solved. The software gap is an ecosystem problem, and ecosystems move on a different clock.
6. Supply Chain: Control Moves, It Does Not Disappear
One of the more persuasive arguments for custom silicon is diversification: reduce dependence on a single accelerator supplier and gain leverage over pricing and allocation. That argument is valid but incomplete, because it addresses only one link in a longer chain.
A company that designs its own accelerator still depends on:
A leading-edge foundry — very few options at the most advanced nodes, and capacity is allocated years in advance.
Advanced packaging capacity — often the tightest constraint in the chain, and shared across the entire industry.
High-bandwidth memory — a concentrated supplier base and a persistent bottleneck.
Networking silicon — increasingly supplied by the same enablers who implement the accelerator.
Custom design shifts the bottleneck rather than removing it. It converts a negotiation with one merchant vendor into a set of dependencies across a supply chain where several links have fewer alternatives than the original arrangement did.
There is a partial exception. Designing your own accelerator does give you a second source for the compute die, and in a shortage that matters enormously. But if the constraint sits in packaging or memory rather than compute, the diversification benefit is smaller than the investment implies.
The CODEW Lens: Supply chain control is a real benefit and a partial one. The honest framing is that custom silicon diversifies the compute layer while concentrating the enabler layer — which is precisely why the enablers are the most interesting part of this market.
7. The Enablers: Who Actually Captures the Value
The most consequential analytical insight in this report is also the least intuitive. If hyperscalers succeed in designing their own accelerators, the companies that benefit most may not be the hyperscalers. They may be the suppliers that make custom design possible.
Four companies occupy that position.
Broadcom supplies custom accelerator design, high-speed Ethernet switching, connectivity, and advanced packaging expertise. Its disclosed Q3 fiscal 2026 results — $29.6 billion in total revenue and $16.7 billion in AI semiconductor revenue — establish that this business operates at scale rather than as an emerging initiative. (Disclosed) Broadcom's position is strengthened by occupying two layers of the same deployment: the accelerator and the fabric connecting accelerators.
Marvell Technology competes in custom compute silicon, optical connectivity, and data-center infrastructure. Recent reporting on its Google partnership illustrates how hyperscaler relationships and supplier diversification are reshaping the market (media report — not company-confirmed in full detail). The relevant structural point is that hyperscalers have historically cultivated multiple suppliers for critical components, which supports the case for at least two credible custom-silicon partners while capping what either can charge.
TSMC provides the advanced manufacturing and packaging capabilities that underpin the entire custom-chip ecosystem essentially. Every hyperscaler program described in this report ultimately depends on the same supplier. That is a position of considerable leverage, and it is not disrupted by anything in the current competitive dynamic.
Nvidia remains the incumbent benchmark for integrated AI computing — GPUs, networking, systems, and software delivered as a coherent platform. The custom silicon movement is frequently framed as a threat to Nvidia. A more accurate framing is that it constrains Nvidia's pricing power at the high end of the market while leaving the flexibility premium intact for everyone else.
| Layer | Substitutability | Pricing power |
|---|---|---|
| Foundry & packaging | Very low | Very high |
| Custom implementation | Low | High |
| Networking silicon | Low to moderate | High |
| Merchant accelerators | Moderate, rising | Moderate, under pressure |
| Accelerator design | High — in-housing is the strategy | Low, and structurally declining |
The pattern is clear. Value concentrates in the layers that cannot be in-housed. A hyperscaler can build a design team. It cannot build a foundry. It can write a compiler. It cannot conjure packaging capacity.
The CODEW Lens: The custom silicon race is not a contest between hyperscalers and Nvidia. It is a transfer of value from the layer that can be in-housed to the layers that cannot. Foundries, packaging, and implementation partners win. Merchant accelerator pricing at the high end is what gets compressed.
8. Displacement or Hybrid?
The question of whether custom silicon displaces GPUs is usually posed as a forecast. It is better approached as a structural observation, because the evidence already points in one direction.
Google offers third-party GPUs alongside its own TPUs. That is a disclosed strategic choice, not a transitional compromise. (Disclosed) A company with the longest-running and most mature custom silicon program in the industry has concluded that it still needs both.
The reason is not technological. It is economic. Custom silicon requires a stable, high-volume workload to justify its fixed cost. The AI workload landscape is not stable — it is changing faster than at any point in the industry's history. A company that commits entirely to specialized hardware has made a bet that its workloads will not change shape. That is a bet no rational operator makes at this moment.
The likely equilibrium is therefore a hybrid architecture in which:
Stable, high-volume workloads — recommendation, ranking, established inference — migrate to custom silicon.
Frontier training and experimental workloads remain on general-purpose accelerators.
Emerging workloads — agentic systems, novel architectures — stay flexible until their shape is understood.
This is not a compromise outcome. It is the rational allocation of specialized and general-purpose hardware across a workload portfolio with different characteristics. The companies that manage that allocation well will have lower cost structures than those that commit to either extreme.
The CODEW Lens: The hybrid outcome is not a draw. It represents a permanent ceiling on merchant accelerator pricing at the high-volume end and a permanent floor under demand for flexibility everywhere else. Both matter, and they affect different parts of the market.
9. Where the Durable Value Sits
The investment implications follow from the structural analysis rather than from any forecast about which architecture wins.
Durable value sits where substitution is hardest. Advanced manufacturing, packaging capacity, and high-bandwidth memory have concentrated supplier bases and constrained capacity. They benefit from every architecture, every vendor, and every hyperscaler program. They do not need the custom silicon thesis to be correct in order to grow.
Durable value sits in implementation. The companies that turn a customer's architecture into working silicon — Broadcom, Marvell — occupy a position that is genuinely difficult to in-house because it requires accumulated IP across many programs. Hyperscalers can build design teams, but building the SerDes library, the packaging expertise, and the manufacturing relationships takes years.
Durable value sits in networking. Networking content per deployment scales with cluster size rather than with design wins. It grows without requiring new customers, and it benefits whether the endpoints are custom accelerators or merchant GPUs.
Pricing pressure sits in the layer that is being in-housed. Merchant accelerator pricing at the high-volume end faces structural pressure from a growing set of credible internal alternatives. That does not mean the merchant business declines — the flexibility segment remains large and defensible. It means the premium that the merchant vendor could charge for being the only option is gone, and it is not coming back.
A caution on valuation. The custom silicon narrative is well understood and widely discussed. Where a company's growth depends on hyperscaler capital expenditure remaining elevated, that assumption is already embedded in expectations. The asymmetry is unfavorable: confirmation produces no re-rating, while a capex digestion period or a program cancellation produces an outsized reaction. This applies across the enabler layer, not only to any single company. (CODEW-derived analysis)
The CODEW Lens: The safest position in a race is not the fastest runner. It is the company supplying the track. Foundry, packaging, memory, and implementation partners are the track. They get paid regardless of who wins.
Conclusion: A Control Strategy, Not a Chip Strategy
The custom silicon race is best understood as the largest buyers in technology asserting control over the largest cost line in their business. Google, AWS, Meta, and Microsoft are not trying to become semiconductor companies. They are trying to stop being purely price-takers in the layer that determines their margins, their deployment schedules, and their product roadmaps.
That objective is achievable, and the disclosed evidence shows meaningful progress. Meta has deployed MTIA at scale in production workloads. Google has iterated TPUs across multiple generations and split its roadmap by workload. AWS has moved Trainium2 to general availability. Microsoft has a production inference accelerator. None of these are announcements of intent. They are operating programs.
What remains unproven is the durability of the outcome. A custom silicon program that works during a period of expanding capital expenditure has not yet demonstrated that it works during a contraction. The fixed costs are committed; the volumes that justify them are not guaranteed. That is the risk that sits beneath every program described in this report, and it is the risk that will be tested before the decade is out.
The defensible conclusion: custom silicon is becoming a permanent feature of AI infrastructure rather than a temporary experiment — but it will settle into a hybrid market, not a replacement one. Value will concentrate in the layers that cannot be in-housed: foundry, packaging, memory, networking, and custom implementation. The layer being in-housed will lose its pricing premium, not its market.
The CODEW Lens: The companies designing their own chips are not trying to win the semiconductor market. They are trying to stop paying for someone else's roadmap. Understanding that distinction is the difference between reading this as a competitive battle and reading it as a procurement strategy at unprecedented scale.
What to Watch
| Signal | Reinforces the thesis | Weakens the thesis |
|---|---|---|
| External availability | Custom accelerators offered broadly to cloud customers | Programs remaining internal-only after multiple generations |
| Workload scope | Custom silicon expanding into training from inference | Programs confined to a single narrow workload |
| Software ecosystem | Third-party frameworks supporting custom silicon natively | Porting costs remaining prohibitive for external developers |
| Supply chain | Packaging and memory capacity expanding to meet demand | Programs delayed by constraints outside the hyperscaler's control |
| Capex behavior | Custom silicon volumes holding through a spending pause | Programs scaled back faster than merchant purchases |
| Enabler concentration | Multiple credible implementation partners winning programs | Implementation consolidating to a single supplier per layer |
The CODEW Stat
4 hyperscaler programs · 3 enabling layers · 1 hybrid outcome Special Report is a long-form strategic analysis series from The CODEW. It examines an industry-level question rather than a single company, connecting economics, technology choices, supplier relationships, and the distribution of value across a market. It sits alongside the Company Intelligence family, which covers individual companies through Company Deep Dive and Company Analysis.
Coverage in this series is based on company announcements, earnings disclosures, technical documentation, SEC filings, attributed market research, and original reporting. All figures are labeled as disclosed or estimated. Broadcom financial figures cited are drawn from its September 2, 2026 earnings release and represent reported results. Performance-per-dollar claims referenced from vendor announcements are company claims and have not been independently verified on production workloads. Framework and valuation analysis are CODEW-derived and should not be read as investment advice.
Reviewed by Erwin Castro
on
Saturday, October 10, 2026
Rating:

No comments: