Signal & Supply
โ† Archive
August 21, 2026 COMPUTE

Why AI Runs on GPUs, Not CPUs.

Both are chips that do math. But a modern AI model does one kind of math โ€” the same small operation, repeated billions of times, all at once. That single fact decides the layout of the entire data center.

Key takeaway A CPU is a few strong workers optimized to finish one hard task fast (latency). A GPU is a swarm of simple workers optimized to finish a mountain of identical tasks together (throughput). AI is a mountain of identical tasks โ€” so the building fills with GPUs.

Watch the race

1,024 identical tasks ยท CPU: 4 fast lanes vs GPU: 64 slower lanes ยท same total work โ€” press Run

CPU4 cores
Done: 0% Ticks: 0
GPU64 cores
Done: 0% Ticks: 0

The plain version

Think of a CPU as four master chefs. Each one can cook almost any dish โ€” sear, braise, plate a dessert โ€” fast and well, switching between totally different tasks without missing a beat. That flexibility is exactly what you want for running an operating system, a web browser, a spreadsheet: a grab-bag of different jobs arriving in no particular order.

A GPU is a different bet: a thousand line cooks, each trained to do exactly one thing โ€” flip a burger. Ask any single line cook to plate a tasting menu and they're useless. But ask all thousand to flip a thousand identical burgers at once, and nothing beats them. AI is the stadium crowd, not the tasting menu: training and running a model is mostly one boring operation โ€” multiply two grids of numbers together and add up the results โ€” repeated billions of times per second. Crucially, each of those multiplications doesn't depend on the others, so all thousand line cooks can work at once instead of waiting in line.

That's why a modern AI data center is racks and racks of GPUs, with a much smaller number of CPUs acting as traffic cops โ€” loading data, scheduling jobs, and handing work off to the GPU swarm. Once you've committed to that layout, the real bottleneck stops being "how many chips can we buy" and becomes power and cooling: a thousand line cooks all working flat-out generate a lot of heat and draw a lot of electricity.

The expert version

A CPU spends its transistor budget on control: deep pipelines, branch prediction, out-of-order execution, and large caches, all in service of minimizing the latency of a single instruction stream. A GPU spends roughly the same die area on the opposite bet โ€” thousands of simple ALUs executing in SIMT (single-instruction, multiple-thread) lockstep, trading per-thread latency for aggregate throughput across many threads at once.

The AI workload matches the GPU's bet almost exactly. The core primitive of both the forward and backward pass through a transformer is dense matrix multiplication (GEMM), dominated by independent multiply-accumulate operations โ€” ideal for massive parallelism. Modern accelerators add Tensor Cores, units that compute a small matrix tile (e.g., 16ร—16) per instruction in reduced precision (FP16, BF16, or FP8), which is the source of most of the headline TFLOPS figures vendors quote. GPUs hide memory latency not through caching but through thread oversubscription โ€” when one warp stalls on a memory fetch, the scheduler simply switches to another warp that's ready to run โ€” which is why performance is frequently memory-bound rather than compute-bound, and why GPUs pair their ALUs with wide, fast HBM instead of large caches.

At frontier scale, a single model no longer fits on one GPU's memory, so training spans thousands of GPUs and the bottleneck shifts from the chip itself to the interconnect โ€” NVLink within a node, InfiniBand or equivalent across nodes โ€” that keeps thousands of parallel workers synchronized.

Memory bandwidth, to scale โ€” how fast each chip can feed its own math units
Server CPU
~0.4 TB/s
Data-center GPU
~6 TB/s

Why it matters for tech + supply chain: how much AI you can run is really how many GPUs you can power and cool, which is why constraints moved from chips to megawatts, land, and memory.

Why it matters for tech + supply chain: the throughput-over-latency bet at the chip propagates into HBM demand, power delivery, network topology, and cost per token; downstream shortages trace back to it.