Why AI Models Need a GPU: A Beginner Explainer (October 2026)

A plain-language walk through parallelism, matrix multiplication, VRAM, memory bandwidth, and why a CPU can run an LLM but rarely fast enough.

Updated October 10, 2026

Every guide on this site starts from a number: how much VRAM your graphics card has. That framing is correct for picking hardware, but it skips the question behind it. A reader who lands on best-gpus-for-local-ai or how-much-vram-does-a-local-llm-actually-need with no background has to take the word of the rest of the site that GPUs matter. This page is the answer to “why”, in plain language, with the four mechanics that actually explain the gap between a CPU and a GPU for running an LLM.

What a GPU is, in one paragraph

A CPU is built for serial work: a few very fast cores that branch, predict, and serve one program well at a time. A GPU is built for parallel work: thousands of small arithmetic units that do the same operation on many numbers at once. The Wikipedia article on CUDA describes GPUs as having “evolved into highly parallel multi-core systems allowing efficient manipulation of large blocks of data,” and frames the trade-off directly: this design “is more effective than general-purpose central processing units (CPUs) for algorithms in situations where processing large blocks of data is done in parallel.” Machine learning, the page lists, is one of those situations.

That is the whole reason a GPU helps with AI. An LLM, when it is generating one new token, is multiplying two big matrices. Doing that on a CPU is possible. Doing it on a GPU is parallel, so it finishes sooner.

Why matrix multiplication matters

A transformer layer is, mechanically, a sequence of matrix multiplications with a few non-linear steps between them. The Hugging Face Transformers optimization guide states the rule of thumb plainly: “Loading the weights of a model having X billion parameters requires roughly 2 * X GB of VRAM in bfloat16/float16 precision.” For a 70-billion-parameter model that is about 140 GB just to put the weights in memory.

The guide also explains why memory bandwidth, not arithmetic throughput, sets the ceiling on how fast tokens come out. During autoregressive generation, each new token requires reading the entire model plus the growing key-value cache, and the speed of generation is dominated by how fast the chip can stream those weights out of memory rather than by how many FLOPs it can do. That is the reason a high-bandwidth card that has fewer raw FLOPs can sometimes outpace a compute-heavy card with narrower memory.

Tim Dettmers’ “Which GPU for Deep Learning?” puts the same point more bluntly: “Tensor Cores are very fast. So fast, in fact, that they are idle most of the time as they are waiting for memory to arrive from global memory,” and “When comparing two GPUs with Tensor Cores, one of the single best indicators for each GPU’s performance is their memory bandwidth.” The arithmetic units are not the bottleneck. The pipe feeding them is.

What Tensor Cores add on top

Modern NVIDIA GPUs have a second, narrower set of units on each compute block called Tensor Cores. The NVIDIA Ampere architecture white paper describes them as performing “fused multiply-add (FMA) operations on matrix fragments in a single instruction,” with the A100 generation doing “256 FP16/FP32 FMA operations per clock” per Tensor Core and 1024 per SM. That is what makes the FP16 Tensor Core peak of 312 TFLOPS on A100 possible, compared to the FP32 FMA path on V100.

For a beginner the practical takeaway is small. A GPU without Tensor Cores can still run an LLM. A GPU with Tensor Cores runs the same LLM in roughly the same wall-clock time but at much higher throughput, which matters when a model is being served to more than one user at once. For a single home user the difference between “with Tensor Cores” and “without” is usually a small fraction of total throughput, while the difference between “any GPU” and “no GPU” is often a factor of ten.

Why VRAM, not total RAM, is the headline number

The model weights have to live somewhere the GPU can read them at full bandwidth. That somewhere is the GDDR6X or HBM memory physically attached to the GPU, called VRAM. System RAM is reachable from the GPU through the PCIe bus, but at a fraction of the speed.

This is the reason every guide on this site talks about VRAM in gigabytes rather than system RAM in gigabytes. Two cards can have the same nominal memory size and behave very differently because of the type: GDDR6X on consumer GeForce cards versus HBM2e or HBM3 on data-center cards. HBM is wider and faster per byte, which is why a 24 GB RTX 3090 trades blows with a 24 GB card it has no business competing with on paper.

CardMemoryBandwidthCUDA coresTGP
RTX 309024 GB GDDR6X936 GB/s10,496350 W
RTX 409024 GB GDDR6X1,008 GB/s16,384450 W

Bandwidth figures are from the Wikipedia specifications tables for each card (RTX 3090: “936 GB/s”; RTX 4090: “1008 GB/s”). Memory size, CUDA core counts and TGP figures are from each card’s NVIDIA specifications page.

The bandwidth column is the one that matters for token generation. A higher number there is, all else equal, faster tokens.

What a CPU can do, and what it cannot

A modern CPU can still run an LLM. llama.cpp, the engine under Ollama and most local AI software, is MIT-licensed and supports “AVX, AVX2, AVX512 and AMX support for x86 architectures” alongside its GPU backends. Generation on a CPU is slower per token but produces identical output, which is why local-ai-without-a-gpu works at all.

The catch is throughput. A CPU reads model weights out of system RAM, which is typically tens of gigabytes per second for a desktop. A consumer GPU reads weights out of VRAM at several hundred gigabytes per second. That order-of-magnitude gap is why a CPU can produce one to two tokens per second on a 7B model where a GPU produces thirty to sixty.

Apple Silicon lives in the middle. The MacBook Pro M4 Max page lists 410 GB/s memory bandwidth for the 32-core GPU variant and 546 GB/s for the 40-core GPU variant, both feeding a unified memory pool up to 128 GB. That is well below a discrete GPU’s bandwidth but well above a CPU’s, which is why M-series Macs are a useful middle ground that run-llms-locally-on-apple-silicon covers in detail.

What quantization changes, and what it does not

Quantization shrinks the weights so they fit in less memory. The Transformers guide describes the trade-off: 8-bit and 4-bit quantization achieve “computational advantages without a considerable decline in model performance.” Inference time is often not reduced, though, because the quantized weights are dequantized on the fly to bfloat16 for the matrix multiplication. What changes is the memory footprint, which is what makes a 15-billion-parameter model fit on a 24 GB card.

That is why what-quantization-costs-you is its own page. Quantization makes the model fit; it does not by itself make the model faster.

Bottom line

A GPU helps with AI for four concrete reasons. It does the same arithmetic in parallel across thousands of cores instead of a few. It has dedicated Tensor Cores that fuse matrix multiplies into one instruction. It has high-bandwidth VRAM sitting next to those cores, so the weights stream in fast enough to keep them busy. And it can hold the entire model plus a usable context window in that memory, so nothing has to be swapped out to system RAM mid-generation. A CPU can run the same model and produce the same output, but reads from system RAM at a fraction of the bandwidth, which caps how many tokens per second it can produce. Apple Silicon sits between the two with unified memory and per-chip bandwidth in the few-hundred GB/s range. The pages this one links to translate these mechanics into actual hardware picks.