Skip to content
Wednesday, October 7, 2026
iInnovate MagSTARTUPS · INNOVATION · GADGETS · AI
AI

What Actually Makes AI Hardware Fast: GPUs, TPUs, and Matrix Math

AI hardware is built around one dominant operation: multiplying large matrices of numbers, in parallel, over and over. Google's introduction to Cloud TPU defines the chips as custom-developed application-specific integrated circuits for machine learning, and NVIDIA's March 18, 2025 Blackwell…

Mei-Ling Chen · May 11, 2026 · 6 min read
ShareXFacebookLinkedInTelegramEmail
Hands steady a glowing accelerator board over an open server chassis in a warm graphite lab, off-white cable bundles and one teal status light catching the frame.
Hands steady a glowing accelerator board over an open server chassis in a warm graphite lab, off-white cable bundles and one teal status light catching the frame.

AI hardware is built around one dominant operation: multiplying large matrices of numbers, in parallel, over and over. Google's introduction to Cloud TPU defines the chips as custom-developed application-specific integrated circuits for machine learning, and NVIDIA's March 18, 2025 Blackwell Ultra announcement describes a rack of 72 GPUs and 36 CPUs built for the same workload.

Why did matrix multiplication end up owning the chip?

Neural networks compute with tensors — grids of numbers representing weights, activations, and gradients — and nearly every step of training and inference reduces to matrix multiplies and accumulations. Google's TPU documentation states that TPUs are optimized for workloads dominated by matrix computations, and it is equally blunt about what they are bad at: linear algebra with frequent branching, lots of element-wise operations, or requirements for high-precision arithmetic. The chip is a specialist, and the software stack exists to keep the specialist fed.

That feeding happens through compilation. The TPU documentation explains that code goes through the XLA compiler, which turns the linear algebra, loss, and gradient portions of a model graph into TPU machine code while everything else executes on the host machine. On the NVIDIA side, the same specialization appears as tensor cores inside GPUs — matrix engines placed next to general-purpose graphics hardware — which is why a GPU can still render a game while a TPU cannot, yet both excel at the same multiply-accumulate flood. The architectural difference is proportion, not kind.

How do the three chip types divide the work?

Google's documentation lays out the split explicitly, and it is the clearest published guide to choosing AI hardware that exists. CPUs handle quick prototyping that requires maximum flexibility, simple models, small batch sizes, and workloads limited by input-output or networking bandwidth. GPUs fit models with a significant number of custom PyTorch or JAX operations and medium-to-large batches. TPUs take models dominated by matrix math with no custom operations in the main training loop, training that runs for weeks or months, and ultra-large embeddings common in ranking and recommendation workloads.

Workload signalBest-fit hardware (per Google's docs)
Quick prototypes, tiny models, I/O-bound trainingCPU
Many custom PyTorch/JAX operations, medium-large batchesGPU
Weeks-long training, huge batches, giant embedding tablesTPU
Frequent branching, element-wise algebra, high-precision mathNone of the above — TPUs explicitly unsuited

The table doubles as a warning label. The documentation's list of what TPUs are not suited for is as specific as the list of what they excel at, which is the honest way to evaluate any AI accelerator: the benchmark that matters is the one shaped like your model, not the one shaped like the vendor's marketing deck.

What does a modern AI system look like at rack scale?

NVIDIA's announcement of Blackwell Ultra, made at GTC on March 18, 2025, shows the current ceiling. The GB300 NVL72 connects 72 Blackwell Ultra GPUs and 36 Arm Neoverse-based Grace CPUs in a rack-scale design, and the company claims 1.5 times the AI performance of the previous GB200 NVL72. NVIDIA also claims the platform increases the revenue opportunity for AI factories by 50 times compared with Hopper-based systems — a company-claimed figure tied to throughput assumptions, not audited results.

The announcement's framing matters for another reason: it targets test-time scaling inference — the art of applying more compute during inference to improve accuracy, in NVIDIA's own words — for reasoning, agentic, and physical AI. The newest hardware is optimized not only for training runs but for the inference explosion that follows them, when a deployed model answers queries millions of times a day, each answer potentially extended by deliberate multi-step reasoning.

Why can't software just catch up instead?

Because the bottleneck is arithmetic volume, not code quality. Reasoning models that deliberate over longer chains multiply the inference load; NVIDIA's announcement quotes its chief executive saying reasoning and agentic AI demand orders of magnitude more computing performance. Training a large model is a months-long stream of matrix multiplies over trillions of parameters, and no compiler trick removes the underlying operations — compilation, as the TPU docs describe it, only maps them onto the silicon as efficiently as possible.

That is also why the accelerators keep getting larger and more tightly connected. A single chip's memory and interconnect limit how big a model slice can be, so vendors build racks and pods — groups of chips acting as one logical accelerator — which Google's documentation describes as slices that workloads scale across with minimal code changes. At that point networking becomes part of the computer: the speed of the links between chips determines how often any one of them sits idle waiting for data. The system, not the chip, is the unit of AI compute.

What should a buyer actually take away?

Match the hardware to the shape of the workload, using the maker's own documentation as the first filter. If a model spends its time in standard matrix operations at large batch sizes, TPU-class or tensor-core hardware earns its cost. If it is riddled with custom operations, branching logic, or unusual precision requirements, general-purpose GPUs or even CPUs may finish first — the TPU documentation says so directly, and that candor is rare enough to be worth trusting.

Second, treat multi-x performance claims as company-claimed until an independent, named evaluation says otherwise. NVIDIA's 1.5x figure compares specific systems under specific conditions, and the 50x revenue-opportunity figure is a business projection layered on top of hardware claims. Third, budget for the whole system — interconnect, memory, and the compiler stack — because the documented fit between workload and silicon is exactly where the differences between a fast deployment and an expensive idle cluster hide.

What about memory — why does it dominate the bill?

Because matrix multiplies move more data than they compute, relatively speaking. Every multiply needs its operands fetched from memory and its result written back, and on models with hundreds of billions of parameters, the parameters themselves are the data. Chip designers respond by stacking high-bandwidth memory next to the compute engines and by linking chips into the slices and racks described above, so that memory capacity grows with the accelerator count rather than bottlenecking behind it.

The consequence for buyers is that accelerator counts are only half the sizing exercise. A cluster with too little memory per chip runs models in more, smaller shards — adding communication that can dominate the runtime — while a cluster with too much memory per chip idles capacity. The same documentation-first rule applies: model size, batch size, and precision determine memory demand, and the vendor's system pages state the memory configuration per slice explicitly. Reasoning workloads sharpen the tradeoff further, since test-time scaling keeps activations in memory for longer deliberation chains.

iInnovate Mag is an independent publication and is not affiliated with any company mentioned in this article.

Sources

  1. Introduction to Cloud TPU — Google Cloud
  2. NVIDIA Blackwell Ultra AI Factory Platform Paves Way for Age of AI Reasoning — NVIDIA Newsroom

More from our brands

Part of the VUGA Network