When a graphics chip (GPU) runs a program, it does not do exactly what a CPU does but faster. It does something qualitatively different: it duplicates one thread of execution hundreds of times and launches them all at once over thousands of pieces of data. Understanding that difference is understanding why your graphics card has thousands of “cores” while a CPU has only a few dozen, and why the two complement each other instead of competing.
The CPU thinks, the GPU divides
A modern CPU (for example, a chip with 8 or 16 cores) is designed to run one instruction after another with the least possible delay. It spends silicon and transistors on branch prediction, register renaming, instruction reordering and large, fast caches. All that hardware exists so that a single thread reaches the end of its path as quickly as possible.
A GPU does not chase that goal. Its hundreds of execution blocks are individually much simpler and slower than a CPU core. The play is one of volume: instead of speeding up one thread, it runs thousands of them in parallel and hides the latency (the wait time of each operation) of one by filling the gap with the work of the others. It is the same trick as an office with many counters: each task takes the same time, but the total comes out sooner.
One program, many hands: the SIMT model
The architecture that makes this possible is called SIMT (Single Instruction, Multiple Threads). It is a generalisation of the older SIMD (Single Instruction, Multiple Data). The core idea: one control thread governs a group of threads that execute the same instruction at the same time, each on its own data.
In vendor terms, that control thread is called a warp on NVIDIA GPUs (groups of 32 threads) and a wavefront on AMD GPUs (groups of 64). Each warp advances like a file of soldiers marking the step: they all execute the same instruction at once. If the data inside the warp differ, the operations run on each register, but the instruction is common to all 32.
The execution block and its “cores”
GPU silicon is organised into compute blocks: NVIDIA calls them SM (Streaming Multiprocessor) and AMD CU (Compute Unit). Each holds dozens of ALUs (arithmetic-logic units, the circuits that add and multiply), units for special functions, and a small cache and shared memory.
When people speak of “a GPU with 10,000 cores”, there are not 10,000 independent processors like in a CPU. In reality there is a handful of SMs, and inside each one dozens of ALU lanes which, synchronised by the warps, execute elementary operations in a massive way. It is legitimate but imprecise marketing: those are 10,000 computing lanes, not 10,000 “chips”.
Divergence: the price of deciding inside the warp
The SIMT model has a known weakness called warp divergence. If a program contains an if and inside the same warp some threads take one branch and others take the other, the GPU cannot run both at once. It first runs the true branch (disabling the threads that did not take it) and then the false one. The work of the branch not taken by each thread is wasted. That is why GPU kernels perform better when written without data-dependent branches, or when at least all the threads in a warp make the same decision.
Why memory decides more than the cores
A typical GPU problem is limited not by computation but by memory access. VRAM uses extremely wide interfaces: while a CPU talks to its RAM through 64- or 128-bit channels, a professional GPU can use buses of 384 bits or more with GDDR or HBM (High Bandwidth Memory, layered memory stacked right beside the chip). The result is a bandwidth 10 to 20 times that of CPU memory.
That bandwidth is precisely what feeds thousands of ALUs. A kernel is considered “memory-bound” when data arrives slower than the ALUs consume it; the worst combination is writing a GPU kernel that does very little arithmetic per byte read.
Latency hiding and occupancy
A GPU hides latency through occupancy: the more warps an SM can hold with their data ready, the sooner it can switch from one to another when one is waiting on memory. While one warp requests a piece of data, the scheduler launches another warp whose data is already cached. Adding more threads does not always speed things up: each consumes registers, and if an SM runs out of registers no more warps fit and occupancy drops.
Tensor cores: the next step
For AI and floating-point work, modern GPUs add specialised units called tensor cores on NVIDIA (and similar units on AMD, such as those in the RX line). They are ALUs designed for matrix multiplication, the central operation of neural networks, performing a small matrix multiply-and-accumulate (for example 4×4) in a single cycle. That is why training large models is measured in TFLOPS rather than in the speed of a single core.
The dancing partners
In practice, CPU and GPU do not compete: they take turns. The CPU runs the general logic, decides what to do, and hands the GPU huge, regular operations (rendering pixels, multiplying matrices, transforming geometry) that fit the SIMT model. The cost of moving data between RAM and VRAM (or, in integrated laptops, of sharing the memory) is what decides whether the division pays off. Understanding the GPU is, above all, understanding that massive parallelism is not “more speed” but a different contract: thinking in terms of synchronised threads doing the same thing, and of the bytes crossing every second.





