The Problem: Idle Time Waiting for Slow Operations
Even with pipelining and instruction-level parallelism, both discussed earlier in this series, a processor core frequently stalls while waiting for a slow operation to complete, such as a cache miss that requires fetching data from main memory. During these stalls, execution hardware that could otherwise be doing useful work sits idle.
The Core Idea: Multiple Threads Sharing One Core
Hardware Multithreading allows a single physical processor core to hold the state of multiple independent Threads (separate streams of instructions) simultaneously, switching between them to keep execution hardware busy even when one thread is stalled waiting on a slow memory operation.
Fine-Grained Multithreading
Fine-Grained Multithreading switches between threads on every single clock cycle, cycling through the available threads in a round-robin fashion. This approach can effectively hide latency from short stalls, since some other thread's instruction can be issued in the very next cycle, but it slightly reduces the execution speed of any single thread running in isolation, since that thread only gets a fraction of the cycles.
Coarse-Grained Multithreading
Coarse-Grained Multithreading switches threads only when the currently running thread encounters a costly stall, such as a cache miss requiring access to main memory. This avoids the small per-cycle overhead of switching every cycle, but reacts more slowly to shorter stalls, since a brief delay might not be worth the cost of a full thread switch.
Simultaneous Multithreading: Issuing from Multiple Threads at Once
Simultaneous Multithreading (SMT), combined with the superscalar, multiple-issue hardware discussed earlier in this series, goes further by issuing instructions from several different threads within the very same clock cycle, filling otherwise unused issue slots that a single thread alone could not fully occupy on its own.
Without SMT (single thread):
Cycle 1: 2 of 4 issue slots used, 2 idle
With SMT (two threads sharing the core):
Cycle 1: Thread A uses 2 slots,
Thread B uses the 2 otherwise-idle slotsWhy This Differs From Adding More Cores
Hardware multithreading is not the same as the multicore approach discussed earlier in this series. Multithreading shares the same physical execution resources, such as ALUs, among multiple threads within a single core, whereas multicore designs duplicate entire cores. Multithreading is generally cheaper in terms of chip area but provides a more modest performance improvement, since threads on the same core still compete for shared resources like cache space.
Why This Technique Matters for Modern Processors
Hardware multithreading is widely implemented in commercial processors specifically because real-world workloads frequently stall on memory access, discussed earlier in this series regarding the memory hierarchy, and multithreading provides a relatively low-cost way to recover a meaningful fraction of that otherwise-wasted execution capacity.