Quick answer: For I/O-bound work, threads and asyncio give near-linear speedups — 20 blocking sleeps took about 1.01 s serially and 0.06 s across 20 threads (16–17×). For CPU-bound work the GIL serialises bytecode execution, so four threads gave no gain (and 0.84× on long tasks), while four warmed processes reached 1.5–1.7× on short tasks and 2.5–3.1× once each task was long enough to amortise process start-up. Choose threads or asyncio for waiting, processes or a free-threaded build for computing, and never assume a statement like counter += 1 is atomic just because the GIL hides the race in testing.
Part 12 of our Python series — Module 3, the deep dive. Previous: bytecode and the adaptive interpreter. Series start: environment setup.
What the GIL actually is
CPython has one global interpreter lock. A thread must hold it to execute bytecode, so within one process, only one thread ever runs Python instructions at a time. The lock is released in two important situations:
- during blocking I/O (
socket,open,time.sleep), which is why threaded network code scales; - inside C extensions that explicitly release it (NumPy,
hashlib,zlib, many database drivers).
It is not released in the middle of an arithmetic loop, so Python threads buy concurrency against waiting, not parallelism against computing. Check the build you are running:
import sys, sysconfig
print(sys.version.split()[0]) # 3.14.6
print(sys._is_gil_enabled()) # True on a standard build
print(sysconfig.get_config_var("Py_GIL_DISABLED")) # 0 = GIL compiled in
sys._is_gil_enabled() exists from 3.13, which is also the first version shipping an official free-threaded build (python3.13t). On a free-threaded build the value is False and threads can execute bytecode in parallel — at the price of a slower single-threaded interpreter and third-party packages that may not be thread-safe yet.
The decision table
| Workload | Best tool | Why |
|---|---|---|
| Network calls, disk, database | ThreadPoolExecutor | GIL released during the wait |
| Thousands of concurrent sockets | asyncio | no thread stack per task |
| Pure-Python number crunching | ProcessPoolExecutor | separate interpreters, real cores |
| NumPy / C-extension heavy | threads | the extension releases the GIL |
| Shared mutable state, no locks possible | processes | isolation by design |
Measured: I/O-bound work scales
Twenty tasks that each sleep 50 ms:
| Strategy | Time | Speedup |
|---|---|---|
| serial loop | 1.018 s | 1.0× |
ThreadPoolExecutor(max_workers=20) | 0.062 s | 16.6× |
asyncio.gather over asyncio.to_thread | 0.072 s | 14.2× |
The threaded version is bounded by the slowest task, not by the sum — that is the whole point. Note the last row: asyncio.to_thread() runs a blocking function in a worker thread and awaits it, which is the correct escape hatch when you must call a synchronous library from async code. With truly async I/O (await asyncio.sleep, aiohttp) you avoid the thread entirely and generally go faster still.
The modern futures API keeps this short:
from concurrent.futures import ThreadPoolExecutor
def fetch(url):
return url, len(url) # stand-in for a network call
with ThreadPoolExecutor(max_workers=8) as pool:
for url, size in pool.map(fetch, ["a", "b", "c"]):
print(url, size)
executor.map preserves input order; as_completed yields results as they finish when order does not matter.
asyncio in the shape you should actually write it
import asyncio
async def fetch(name, delay):
await asyncio.sleep(delay)
return f"{name}:{delay}"
async def main():
try:
async with asyncio.timeout(1.0):
async with asyncio.TaskGroup() as tg:
tg.create_task(fetch("a", 0.1))
tg.create_task(fetch("b", 0.2))
except* TimeoutError:
print("timed out")
asyncio.run(main())
Three things changed here relative to old tutorials: asyncio.TaskGroup (3.11+) is the structured-concurrency API that cancels siblings when one task fails; asyncio.timeout (3.11+) replaces wait_for; and exceptions from a task group arrive as an ExceptionGroup, caught with except*. Prefer these over bare gather() in new code.
The one rule that matters: an async function that never awaits anything is just a slower function, and blocking inside a coroutine (a requests call, a big time.sleep) freezes the entire event loop for every task. Use asyncio.to_thread at the boundary.
Measured: CPU-bound work does not — until you use processes
Four identical CPU-heavy tasks run three ways. The interesting variable is how much work each task contains, because that is what the fixed overheads compete against:
| Task size | Serial | 4 threads | 4 warmed processes |
|---|---|---|---|
| 3,000,000 iterations | 0.45 s | 0.41–0.45 s (0.99–1.09×) | 0.28–0.41 s (1.1–1.7×, noisy) |
| 20,000,000 iterations | 3.5 s | 3.9 s (0.84–0.91× — slower!) | 1.2 s (2.4–3.1×) |
| 4 processes, cold pool (3M) | 0.44 s | — | 0.52 s (0.79× — slower!) |
Four lessons in one table:
- Threads never help pure Python computation. On short tasks they are a wash; on long ones they are measurably slower, because every 5 ms the GIL is handed between threads and each hand-off costs the receiver time to reacquire it. Adding threads to CPU-bound code can lose you 10–16%.
- A cold process pool is slower than doing nothing special. On Windows and macOS each worker is a brand-new interpreter that re-imports your module; four spawns cost more than the 0.44 s of work. Warm the pool — submit a trivial job first, or keep it alive across requests — and the real cores appear.
- Process speedup grows with task size. Same pool, same machine: 1.1–1.7× on short 3M-iteration tasks (a wide range, because the fixed cost is comparable to the work) and a steady 2.4–3.1× on 20M-iteration tasks. Chunk small jobs into fewer, bigger tasks.
- Even at its best this is not 4×. Memory bandwidth, the parent process, and your 20-logical-CPU sandbox all take a cut. Measure; never assume the core count.
Two Windows/macOS specifics that break real code:
if __name__ == "__main__":is mandatory. Thespawnstart method re-imports your main module in every child. Without the guard, the children recursively spawn their own pools.- The pool cannot be driven from a stdin script. Running
python - <<'EOF'with aProcessPoolExecutorfails withOSError: Invalid argument: '<stdin>', because the child cannot re-import<stdin>as a file. Save the script to disk first — the same reason notebooks and REPLs make multiprocessing awkward.
Race conditions: why the GIL is a trap
Because only one thread runs bytecode at a time, most read-modify-write code appears correct even though it is not:
| Variant | Expected | Measured | Lost |
|---|---|---|---|
counter += 1, 4 threads × 200k | 800,000 | 800,000 | 0 |
| same, thread switch every 1 µs | 800,000 | 800,000 | 0 |
| read, yield, then write | 800,000 | 200,000 | 600,000 |
The first two rows are the dangerous case: the bug is real in both, but the GIL’s switch interval (5 ms by default) means a thread rarely yields between the load and the store, so the corruption is invisible. Widen the window — read into tmp, release the GIL with time.sleep(0) or any I/O call, then write back — and 75% of the updates vanish. Under a free-threaded build the same code corrupts without any widening at all.
list.append and dict[k] = v are atomic in CPython because their C implementations never yield, but nothing about x = x + 1, x *= 2, or collection.sort() is guaranteed. The discipline is to keep shared state behind a lock or, better, to not share it at all: pass immutable data to workers and collect returned results.
import threading
lock = threading.Lock()
def safe_increment(state):
with lock: # context manager releases even on exception
state["total"] += 1
Programming guidelines worth internalising: hold locks for the shortest possible region, never acquire two locks in different orders in different threads (that is the classic deadlock), and pass timeout= to lock.acquire() so a deadlock fails loudly instead of hanging forever.
Free-threaded Python, briefly
The experimental free-threaded build removes the GIL entirely: threads then scale on CPU-bound pure Python, and the correctness reasoning above stops being optional — the runtime no longer hides your races. Take it seriously, but treat it as a test target, not a default: benchmark your own workload, check that each dependency is thread-safe, and keep a GIL build in production until your stack is officially supported.
Complete executable example
Save as concurrency_lab.py and run python concurrency_lab.py (it needs to be a file, not stdin, because of the process pool):
# concurrency_lab.py -- measure the four execution models on your own machine
import asyncio
import sys
import threading
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def section(title):
print("\n" + title)
print("=" * len(title))
def io_task(seconds=0.05):
time.sleep(seconds) # releases the GIL: perfect for threads
return seconds
def cpu_task(n=3_000_000):
total = 0
for i in range(n): # pure Python: holds the GIL
total += i * i
return total
N = 20
def io_benchmarks():
section("1. I/O-bound: 20 tasks x 50 ms")
start = time.perf_counter()
[io_task() for _ in range(N)]
serial = time.perf_counter() - start
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=N) as pool:
list(pool.map(io_task, [0.05] * N))
threads = time.perf_counter() - start
async def gather_io():
await asyncio.gather(*(asyncio.to_thread(io_task) for _ in range(N)))
start = time.perf_counter()
asyncio.run(gather_io())
async_time = time.perf_counter() - start
print(f"serial {serial:6.3f}s 1.0x")
print(f"threads {threads:6.3f}s {serial / threads:.1f}x")
print(f"asyncio+to_thread {async_time:6.3f}s {serial / async_time:.1f}x")
assert threads < serial / 4, "threads should beat serial I/O by a wide margin"
def cpu_benchmarks():
section("2. CPU-bound: 4 tasks, short and long")
for work in (3_000_000, 20_000_000):
start = time.perf_counter()
[cpu_task(work) for _ in range(4)]
serial = time.perf_counter() - start
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as pool:
list(pool.map(cpu_task, [work] * 4))
threads = time.perf_counter() - start
with ProcessPoolExecutor(max_workers=4) as pool:
pool.map(cpu_task, [1] * 4) # warm up: pay spawn cost here
start = time.perf_counter()
results = list(pool.map(cpu_task, [work] * 4))
processes = time.perf_counter() - start
print(f"work {work:>10,} serial {serial:6.3f}s "
f"threads {threads:6.3f}s ({serial / threads:5.2f}x) "
f"processes {processes:6.3f}s ({serial / processes:5.2f}x)")
assert threads > serial * 0.6, "threads must not speed up pure-Python CPU work"
print("checksum:", sum(results))
assert serial / processes > 1.2, "warmed processes should beat serial by a clear margin"
def race_demo():
section("3. Race condition: a narrow window, then a wide one")
rounds, workers = 200_000, 4
expected = rounds * workers
def run(label, body):
state = {"count": 0}
threads = [threading.Thread(target=body, args=(state,)) for _ in range(workers)]
start = time.perf_counter()
for t in threads:
t.start()
for t in threads:
t.join()
elapsed = time.perf_counter() - start
print(f"{label:26s} expected {expected:>8} got {state['count']:>8} "
f"lost {expected - state['count']:>8} ({elapsed:.2f}s)")
return state["count"]
def narrow(state):
for _ in range(rounds):
state["count"] += 1 # load, add, store -- tiny window
def wide(state):
for _ in range(rounds):
seen = state["count"] # read
time.sleep(0) # yield the GIL: another thread can run
state["count"] = seen + 1 # write stale value back
narrow_count = run("counter += 1", narrow)
wide_count = run("widened window", wide)
lock = threading.Lock()
def locked(state):
for _ in range(rounds):
with lock:
state["count"] += 1
locked_count = run("counter += 1 + lock", locked)
assert narrow_count >= expected * 0.95, "GIL makes the narrow form look correct"
assert wide_count < expected * 0.5, "widening the window must expose the race"
assert locked_count == expected, "the lock must restore correctness"
if __name__ == "__main__": # REQUIRED for ProcessPoolExecutor on Windows/macOS
section(f"Python {sys.version.split()[0]} on {__import__('os').process_cpu_count()} CPUs")
io_benchmarks()
cpu_benchmarks()
race_demo()
print("\nAll concurrency-lab assertions passed.")
Line by line: the module-level guard is not decoration — without it the spawned workers re-import the module and start their own pools; the narrow-window assertion allows a 5% margin, because a rare unlucky thread switch can drop an update even there; io_task sleeps, which releases the GIL, so the 20-thread pool finishes in roughly one task’s time; cpu_task never yields, so threads stall at 1× while processes (warmed with a fake one-iteration job) get real cores; race_demo first proves the narrow += looks correct and then breaks it with time.sleep(0) in the middle, which is the cleanest demonstration that the GIL guarantees bytecode atomicity, not statement atomicity; the final assertions turn each measured claim into a failing test rather than a comment.
Common mistakes
- Threading CPU-bound code and wondering why nothing improved. Check with a profile first; if the time is in Python computation, use processes.
- Starting a process pool while the GIL guard is missing. Symptom: runaway recursive spawning or
BrokenProcessPool. - Blocking inside a coroutine. One
requests.getin an async function freezes every other task; wrap it inasyncio.to_thread. - Sharing mutable state between threads without a lock because “the tests passed”. Increase the interleaving with
sys.setswitchinterval(1e-6)to reproduce. - Treating
list.appendresults as proof of thread safety. Some operations happen to be atomic; almost none are documented as such.
Key takeaways and challenge
- GIL: one thread executes bytecode at a time; it is released for I/O and inside C extensions.
- I/O-bound → threads or asyncio (measured 16–17× and 14× here). CPU-bound → processes.
- Threads make CPU-bound code no faster and can make it 10–16% slower.
- Cold process pools can be slower than serial; warm them, and chunk small jobs so the speedup grows (1.7× → 3.1× here).
counter += 1is not atomic; the GIL merely makes the race hard to trigger.- Use
TaskGroup+asyncio.timeoutin new async code, and never block the event loop.
Challenge: take a real slow script of yours and decide per section, I/O or CPU. Convert the I/O sections to a ThreadPoolExecutor and the CPU sections to a warmed ProcessPoolExecutor, then measure end to end. If your CPU sections use NumPy, try threads first — the extension releases the GIL and you skip all the pickling.
Want help making a real project concurrent? Ampersand Academy does one-to-one Python training and profiling sessions.
What does the GIL actually prevent?
Only one thread executes Python bytecode at a time, so threads do not speed up CPU-bound Python. The lock is released during blocking I/O and inside C extensions such as NumPy.
When should I use threads instead of asyncio?
Use threads when you must call blocking libraries or share simple state with locks. Use asyncio for thousands of concurrent sockets, and asyncio.to_thread to bridge a blocking call into async code.
Why did my multiprocessing pool run slower than the serial version?
A cold pool spawns fresh interpreters that re-import your module, and on Windows that cost exceeded the work measured. Warm the pool first and give each task enough work to amortise the overhead.
Is counter += 1 thread safe in Python?
No. It is a load, add and store, so two threads can read the same value. The GIL usually makes the window too small to trigger, which is why the bug survives testing; use a lock or avoid shared mutable state.
Do I need the if __name__ guard with ProcessPoolExecutor?
Yes on Windows and macOS, because spawn re-imports the main module in each child. Without the guard children start their own pools, and a script run from stdin fails outright.

