Technology Fundamentals · Core Technology Explainer · October 9, 2026
Compare GPUs, TPUs, and NPUs to understand how different AI chips handle training, inference, memory, performance, and energy efficiency.
![]() |
| Image credit: Pexels |
GPU, TPU, and NPU are three answers to the same question: what should a processor look like if its main job is running neural networks? They arrive at different conclusions because they were designed for different constraints — one for programmability, one for efficiency at data center scale, and one for a power budget measured in watts.
The three acronyms are frequently compared as if they were competing products on the same shelf. They are not. They occupy different deployment tiers, carry different software burdens, and are good at different things. This explainer covers what each architecture is, how their designs actually diverge, where memory and energy shape the outcome, and how to decide which one fits a given workload.
1. What GPUs, TPUs and NPUs Are
A GPU — graphics processing unit — is a processor built around thousands of relatively simple cores that execute the same instruction across many data points at once. It was designed to render pixels, and the discovery that this structure maps almost perfectly onto neural network math is what created the modern AI hardware market. Its defining trait is general programmability: a GPU can be made to run nearly any workload, not just the one it was designed for.
A TPU — tensor processing unit — is an application-specific integrated circuit built by Google for its own AI workloads. It is not a product category so much as a specific program: Google began deploying TPUs internally in 2015 and later made them available through its cloud. TPUs strip out graphics functionality entirely and organize computation around large matrix operations, which makes them efficient on AI workloads and less adaptable to anything else.
An NPU — neural processing unit — is a small, power-efficient accelerator integrated into a system-on-chip alongside a CPU and GPU. It appears in smartphones, laptops, tablets, cameras, and cars. NPUs are inference-only in practice: they run trained models rather than training them, and they are designed around the assumption that the available power and thermal budget is tiny.
The taxonomy is looser than it appears. "GPU" and "TPU" name real hardware classes with clear lineage, while "NPU" is closer to a marketing and functional label — vendors use it for quite different internal designs. Any comparison between the three should be read with that in mind. For the broader category they all belong to, see What Is an AI Accelerator?
The CODEW Lens: These are not three competitors on one shelf. They are three designs built for three different power budgets.
2. Architectural Differences
A GPU is built on the principle of latency hiding through parallelism. Thousands of threads run concurrently, and when one stalls waiting on memory, the scheduler switches to another. This is why GPUs tolerate irregular memory access and branching better than dedicated AI silicon — they have the machinery to keep many divergent paths in flight. The cost is overhead: a substantial share of the chip is devoted to scheduling, caches, and control logic rather than arithmetic.
A TPU takes the opposite approach. Its core structure is a systolic array — a grid of multiply-accumulate units through which data flows in a fixed, rhythmic pattern, with results passed from one unit to the next. There is very little control logic because the dataflow is largely predetermined by the compiler. That determinism is the source of the TPU's efficiency: almost all the silicon is doing arithmetic, and almost none is deciding what to do next. It is also the source of its rigidity. Workloads that do not fit the dataflow pattern perform poorly, and there is no scheduler to fall back on.
An NPU is narrower still. It typically consists of MAC arrays and fixed-function blocks specialized for convolution, attention, and activation operations, tightly coupled to the SoC's memory and power management. Many NPUs work primarily on quantized integer arithmetic rather than floating-point, which reduces both energy per operation and data volume. The design is optimized for a single metric: useful inference per watt, within a thermal envelope of a few watts.
The three architectures therefore differ along a consistent axis. GPU maximizes flexibility. TPU maximizes efficiency on a known workload. NPU maximizes efficiency within a hard power constraint. Each trade-off is deliberate, and none is universally correct.
3. Training vs. Inference: Which Architecture for Which Workload
The training/inference split — covered in detail in AI Inference vs. AI Training — maps closely onto the architecture question.
Training is GPU territory, with TPUs as a meaningful secondary option for organizations already inside Google's ecosystem. Training demands flexibility because model architectures keep changing, and it demands scale because frontier runs span thousands of devices. GPUs provide both, along with an interconnect ecosystem that has been tuned for exactly this pattern. TPUs are competitive here — Google trains its own models on them — but they require commitment to Google's toolchain.
Inference is where the three split most cleanly. In a data center, large GPUs remain the default because they handle any model and are already qualified; TPUs and hyperscaler silicon serve the subset of workloads stable enough to justify specialization. At the edge, the NPU is effectively the only option — a data center accelerator cannot operate inside a phone's thermal budget, and a phone cannot afford to send every inference request to a server.
The practical summary: GPUs do everything, TPUs do data center AI well within one ecosystem, and NPUs do on-device inference and nothing else.
4. Performance Considerations
Comparing accelerators by headline throughput is close to meaningless. Published TFLOPS and TOPS figures are quoted at specific numerical precisions, sometimes assume sparse computation, and describe peak capability rather than achievable performance on real models. A chip advertised at twice the throughput of another may deliver less on a workload that does not map well to its architecture.
The metrics that actually matter depend on the job. For training, the relevant questions are time-to-train on a representative model, scaling efficiency as devices are added, and interconnect bandwidth — because a cluster that cannot keep its accelerators synchronized wastes most of its capacity. For inference, the questions are tokens per second, time to first token, latency under load, and cost per million tokens at the concurrency the application actually experiences.
Underneath all of this sits software. A chip's theoretical performance is only accessible if compilers, kernel libraries, and framework support can generate efficient code for it. This is why GPUs have remained dominant even where competitors ship comparable silicon: the realized performance gap is frequently larger than the specification gap, and it comes from tooling rather than hardware.
The CODEW Lens: Peak throughput is a marketing number. Realized performance is a software number.
5. Memory and Bandwidth
Memory architecture is where the three designs diverge most sharply, and it explains much of their behavioral difference. For the underlying technology, see What Is HBM and Why Does It Matter for AI?
Data center GPUs pair large compute arrays with High Bandwidth Memory, delivering bandwidth measured in terabytes per second. This is essential because these chips are frequently memory-bound rather than compute-bound, particularly during inference. Capacity is also generous — enough to hold substantial models entirely on-device, which is often the deciding factor in what can be served from a single accelerator.
TPUs also use HBM and carry unusually large on-chip memory relative to their compute, which reduces how often the array must reach out to external memory. The combination of a deterministic dataflow and ample on-chip storage lets TPUs sustain high utilization on the operations they are built for.
NPUs have no such luxury. They share LPDDR memory with the rest of the SoC, which means bandwidth is limited and contended with the CPU, GPU, and display pipeline. This single constraint shapes nearly everything about on-device AI: models must be small, weights are aggressively quantized to 8-bit or 4-bit, and context windows are kept short. Where a data center accelerator is limited by how fast it can read a 70-billion-parameter model, a phone NPU is limited by the fact that it cannot hold one at all.
6. Energy Efficiency
Efficiency comparisons between these architectures are frequently unfair, because they are measured against different constraints. A data center GPU operating at several hundred watts may be more efficient per operation than a competing design while still consuming vastly more power overall. An NPU at two watts is not more efficient than a GPU in any absolute sense — it is operating in a regime where nothing else could run at all.
What can be compared is efficiency per operation, and here the ordering is reasonably clear. General-purpose GPUs carry overhead from scheduling, caching, and control that specialized silicon does not. TPUs devote more of their area to arithmetic and less to coordination, which generally gives them an efficiency edge on the workloads they target. NPUs push this furthest, using fixed-function blocks and low-precision integer arithmetic to minimize energy per inference.
The largest single lever across all three is numerical precision. Reducing a computation from 16-bit floating point to 8-bit integer cuts both the energy per operation and the volume of data that must be moved, often by a factor of several. Most of the practical efficiency gains in inference over the past several years have come from precision reduction rather than from new silicon — and it works precisely because inference tolerates lower precision far better than training does.
7. Cloud vs. Edge Deployment
Deployment location constrains architecture more than any other factor. Cloud deployments have power, cooling, and space in abundance relative to the chip, so the design can prioritize throughput and capacity. Edge deployments have none of those, so the design must prioritize efficiency above all else.
Cloud is where GPUs and TPUs live, alongside a growing set of hyperscaler-designed accelerators. These chips sit in racks, connect through high-speed fabrics, and are typically rented rather than owned. The relevant constraints are power delivery per rack, cooling capacity, and interconnect topology. Edge is where NPUs live: phones, laptops, tablets, cameras, vehicles, industrial equipment. The relevant constraints are battery life, thermal dissipation, cost per unit, and — increasingly — privacy, since on-device inference means data never leaves the device.
Between these two sits a middle tier that is growing: regional data centers and on-premises inference appliances serving latency-sensitive or data-residency-constrained workloads. This tier uses a mix of architectures depending on volume and requirements, and it is where the case for specialized inference silicon is strongest. For the broader infrastructure picture, see The CODEW's AI Infrastructure Special Report.
8. Nvidia GPUs: The General-Purpose Default
Nvidia's position in AI rests on a stack rather than a chip. Its accelerators are strong — successive generations have pushed memory capacity, bandwidth, and low-precision throughput — but the durable advantage is the software ecosystem built over nearly two decades. Libraries, compilers, and framework integrations mean that most AI code in the world was written to run on Nvidia hardware first, and often only.
That ecosystem effect is why switching costs are high even when an alternative is technically competitive. Porting a training pipeline or an inference stack to different silicon involves revalidation, performance tuning, and often rewriting kernels that were hand-optimized for Nvidia's architecture. The switching cost is a software cost, and it does not appear in any specification sheet.
Nvidia's generational cadence has also shaped the market's expectations. Each new architecture arrives with substantial gains in memory capacity and low-precision throughput, and buyers plan capacity around that rhythm. Competitors must not only match the current generation but keep pace with the next one. For company-level analysis, see Nvidia Company Deep Dive; for the merchant alternative, see The CODEW's AMD coverage.
9. Google's TPUs: Custom Silicon at Hyperscale
Google has run TPUs in production longer than any other organization has run custom AI silicon, and the program's logic is straightforward: if you are consuming enough AI compute, designing your own chip can be cheaper than buying one — and it lets you co-design hardware and software in ways a merchant vendor cannot. Google controls the model architectures, the training frameworks, the compilers, and the data centers, so it can optimize across all of them simultaneously.
The TPU program's limitations are structural rather than technical. TPUs are not sold as merchant silicon; they are consumed internally or rented through Google Cloud. Using them requires committing to Google's toolchain and, in practice, to Google Cloud as a platform. For organizations already there, the economics and performance can be compelling. For organizations with multi-cloud or on-premises requirements, the TPU is not a general option regardless of its merits.
The TPU's larger significance is as a template. Every major cloud provider now runs a comparable program — AWS with Trainium and Inferentia, Microsoft with Maia, Meta with its internal accelerators. The shared thesis is that vertical integration beats merchant purchasing for workloads a provider understands deeply and runs at sufficient volume. Whether that thesis extends beyond the hyperscalers is the open question.
The CODEW Lens: Custom silicon works when you own the whole stack. Most companies do not.
10. NPUs in PCs, Smartphones and Edge Devices
NPUs have been in smartphones for years — handling photography, speech recognition, and keyboard prediction — but they entered mainstream discussion when PC vendors began marketing them as a defining feature. The current generation of laptops ships with NPUs rated in the tens of TOPS, and those figures are now printed on spec sheets alongside processor speed and memory.
The TOPS framing deserves skepticism. It measures peak integer operations at a specific precision, ignores memory bandwidth, and says nothing about whether a real model can run at acceptable speed. A laptop NPU with a high TOPS rating but shared, limited memory bandwidth will struggle with anything beyond small quantized models, regardless of the number on the box.
What NPUs genuinely enable is a class of workloads that could not otherwise exist: real-time transcription, image processing, local assistants, and privacy-preserving inference where data never leaves the device. These are meaningful capabilities, and they are growing. What they are not is a substitute for data center accelerators — the model sizes and context lengths involved are a different order of magnitude, and that gap is set by memory, not by compute.
11. When Each Architecture Makes Sense
A practical decision framework follows from everything above. Choose GPUs when you need flexibility, when model architectures are changing, when you are training at any meaningful scale, or when you need access to a mature software ecosystem. This covers the large majority of production AI today, and it is why GPU demand has proven so durable.
Choose TPUs or comparable custom silicon when you are already committed to a hyperscaler platform, when your workloads are stable and well-understood, and when your volume is high enough that the efficiency gain justifies the loss of portability. This is a narrow set of conditions, which is why custom silicon has not displaced merchant accelerators despite years of investment.
Choose NPUs when inference must happen on a device — because of latency, connectivity, privacy, or cost per query at very high volume — and when the model is small enough to fit within the memory and power budget. This is not a compromise choice; for on-device workloads there is no alternative.
The framework is deliberately unexciting. Most organizations will use GPUs for the bulk of their AI work, may use custom silicon if their cloud provider offers a compelling case, and will use NPUs incidentally through the devices their employees and customers already own.
12. Why the AI Chip Market Is Diversifying
Several forces are pushing the market apart. Cost pressure is the strongest: as inference grows to dominate AI compute spending, the incentive to find a cheaper way to serve models increases, and specialization is the most direct route to lower cost per token. Supply concentration is the second: dependence on a single vendor for the most critical component in the AI stack is a strategic risk that hyperscalers and governments alike have decided to reduce.
Workload standardization is the third. Diversification only makes sense if there is something stable enough to specialize against, and inference — particularly transformer-based inference — has become stable enough to justify dedicated designs in a way it was not five years ago. Sovereignty and export controls add a fourth, as governments fund domestic chip capacity for reasons that have little to do with performance per watt.
The countervailing force is software. Every additional architecture is another target that compilers, frameworks, and libraries must support, and that burden falls on a relatively small number of engineers. This is why the market has diversified in investment without yet diversifying much in production deployment. The likely outcome is not fragmentation into many equal players but layering: a small number of general-purpose architectures handling most work, with specialized silicon occupying specific tiers where the economics clearly justify it. For ongoing coverage of these shifts, see The CODEW's Semiconductor Watch.
The CODEW Lens: Silicon diversifies faster than software does. That gap is what keeps the market concentrated.
FAQ
Q: Is a TPU faster than a GPU?
Not universally. TPUs can be more efficient on the specific workloads they were designed for, particularly at scale within Google's stack. On general workloads, or in environments that require flexibility and portability, GPUs are usually the stronger choice. The comparison depends entirely on what is being run and where.
Q: Can an NPU run a large language model?
Small quantized ones, yes. NPUs are constrained by shared memory bandwidth and limited capacity, so they handle models in the low billions of parameters or below. Larger models must be served from a data center or split between device and cloud, which reintroduces the latency and privacy trade-offs that on-device inference was meant to avoid.
Q: Why do so many companies still buy GPUs if custom silicon is more efficient?
Because efficiency is only one variable. GPUs run any model, work with any framework, and are available for rent at any scale on any cloud. Custom silicon requires committing to a platform, accepting limits on what can be run, and often revalidating an entire software stack. For most organizations, that trade is not worth it.
The CODEW Takeaway
GPUs, TPUs and NPUs are not three products competing for one slot. They are three designs shaped by three different constraints — programmability, data center efficiency, and a few watts of thermal budget — and each is well matched to its tier. GPUs dominate because they are flexible and well-supported, not because they are the most efficient. TPUs work because Google owns the entire stack around them. NPUs exist because on-device inference has no alternative. As inference grows to dominate AI spending, the pressure to specialize will increase, but software portability will keep the market more concentrated than the hardware investment suggests. The architecture that wins a given deployment is the one whose constraints match the problem — and that answer changes with the tier, not with the benchmark.
The CODEW Lens: Match the constraint, not the benchmark. GPU for flexibility, TPU for scale inside one stack, NPU for the edge.
The CODEW Stat
~2 W to ~1,000 W The power envelope spanning the AI chip market — from an NPU inside a smartphone operating within a couple of watts of thermal budget, to a flagship data center accelerator drawing several hundred watts and up. Same class of mathematics, roughly three orders of magnitude apart in power. No single architecture can serve both ends.
Reviewed by Erwin Castro
on
Thursday, October 08, 2026
Rating:

No comments: