Site icon Ampersand Tutorials

Python GIL, Threads, asyncio and Multiprocessing (Measured)

Quick answer: For I/O-bound work, threads and asyncio give near-linear speedups — 20 blocking sleeps took about 1.01 s serially and 0.06 s across 20 threads (16–17×). For CPU-bound work the GIL serialises bytecode execution, so four threads gave no gain (and 0.84× on long tasks), while four warmed processes reached 1.5–1.7× on short tasks and 2.5–3.1× once each task was long enough to amortise process start-up. Choose threads or asyncio for waiting, processes or a free-threaded build for computing, and never assume a statement like counter += 1 is atomic just because the GIL hides the race in testing.

Part 12 of our Python series — Module 3, the deep dive. Previous: bytecode and the adaptive interpreter. Series start: environment setup.

What the GIL actually is

CPython has one global interpreter lock. A thread must hold it to execute bytecode, so within one process, only one thread ever runs Python instructions at a time. The lock is released in two important situations:

It is not released in the middle of an arithmetic loop, so Python threads buy concurrency against waiting, not parallelism against computing. Check the build you are running:

import sys, sysconfig
print(sys.version.split()[0])                    # 3.14.6
print(sys._is_gil_enabled())                     # True on a standard build
print(sysconfig.get_config_var("Py_GIL_DISABLED"))  # 0 = GIL compiled in

sys._is_gil_enabled() exists from 3.13, which is also the first version shipping an official free-threaded build (python3.13t). On a free-threaded build the value is False and threads can execute bytecode in parallel — at the price of a slower single-threaded interpreter and third-party packages that may not be thread-safe yet.

The decision table

WorkloadBest toolWhy
Network calls, disk, databaseThreadPoolExecutorGIL released during the wait
Thousands of concurrent socketsasynciono thread stack per task
Pure-Python number crunchingProcessPoolExecutorseparate interpreters, real cores
NumPy / C-extension heavythreadsthe extension releases the GIL
Shared mutable state, no locks possibleprocessesisolation by design

Measured: I/O-bound work scales

Twenty tasks that each sleep 50 ms:

StrategyTimeSpeedup
serial loop1.018 s1.0×
ThreadPoolExecutor(max_workers=20)0.062 s16.6×
asyncio.gather over asyncio.to_thread0.072 s14.2×

The threaded version is bounded by the slowest task, not by the sum — that is the whole point. Note the last row: asyncio.to_thread() runs a blocking function in a worker thread and awaits it, which is the correct escape hatch when you must call a synchronous library from async code. With truly async I/O (await asyncio.sleep, aiohttp) you avoid the thread entirely and generally go faster still.

The modern futures API keeps this short:

from concurrent.futures import ThreadPoolExecutor

def fetch(url):
    return url, len(url)      # stand-in for a network call

with ThreadPoolExecutor(max_workers=8) as pool:
    for url, size in pool.map(fetch, ["a", "b", "c"]):
        print(url, size)

executor.map preserves input order; as_completed yields results as they finish when order does not matter.

asyncio in the shape you should actually write it

import asyncio

async def fetch(name, delay):
    await asyncio.sleep(delay)
    return f"{name}:{delay}"

async def main():
    try:
        async with asyncio.timeout(1.0):
            async with asyncio.TaskGroup() as tg:
                tg.create_task(fetch("a", 0.1))
                tg.create_task(fetch("b", 0.2))
    except* TimeoutError:
        print("timed out")

asyncio.run(main())

Three things changed here relative to old tutorials: asyncio.TaskGroup (3.11+) is the structured-concurrency API that cancels siblings when one task fails; asyncio.timeout (3.11+) replaces wait_for; and exceptions from a task group arrive as an ExceptionGroup, caught with except*. Prefer these over bare gather() in new code.

The one rule that matters: an async function that never awaits anything is just a slower function, and blocking inside a coroutine (a requests call, a big time.sleep) freezes the entire event loop for every task. Use asyncio.to_thread at the boundary.

Measured: CPU-bound work does not — until you use processes

Four identical CPU-heavy tasks run three ways. The interesting variable is how much work each task contains, because that is what the fixed overheads compete against:

Task sizeSerial4 threads4 warmed processes
3,000,000 iterations0.45 s0.41–0.45 s (0.99–1.09×)0.28–0.41 s (1.1–1.7×, noisy)
20,000,000 iterations3.5 s3.9 s (0.84–0.91× — slower!)1.2 s (2.4–3.1×)
4 processes, cold pool (3M)0.44 s—0.52 s (0.79× — slower!)

Four lessons in one table:

Two Windows/macOS specifics that break real code:

Race conditions: why the GIL is a trap

Because only one thread runs bytecode at a time, most read-modify-write code appears correct even though it is not:

VariantExpectedMeasuredLost
counter += 1, 4 threads × 200k800,000800,0000
same, thread switch every 1 µs800,000800,0000
read, yield, then write800,000200,000600,000

The first two rows are the dangerous case: the bug is real in both, but the GIL’s switch interval (5 ms by default) means a thread rarely yields between the load and the store, so the corruption is invisible. Widen the window — read into tmp, release the GIL with time.sleep(0) or any I/O call, then write back — and 75% of the updates vanish. Under a free-threaded build the same code corrupts without any widening at all.

list.append and dict[k] = v are atomic in CPython because their C implementations never yield, but nothing about x = x + 1, x *= 2, or collection.sort() is guaranteed. The discipline is to keep shared state behind a lock or, better, to not share it at all: pass immutable data to workers and collect returned results.

import threading

lock = threading.Lock()

def safe_increment(state):
    with lock:                 # context manager releases even on exception
        state["total"] += 1

Programming guidelines worth internalising: hold locks for the shortest possible region, never acquire two locks in different orders in different threads (that is the classic deadlock), and pass timeout= to lock.acquire() so a deadlock fails loudly instead of hanging forever.

Free-threaded Python, briefly

The experimental free-threaded build removes the GIL entirely: threads then scale on CPU-bound pure Python, and the correctness reasoning above stops being optional — the runtime no longer hides your races. Take it seriously, but treat it as a test target, not a default: benchmark your own workload, check that each dependency is thread-safe, and keep a GIL build in production until your stack is officially supported.

Complete executable example

Save as concurrency_lab.py and run python concurrency_lab.py (it needs to be a file, not stdin, because of the process pool):

# concurrency_lab.py -- measure the four execution models on your own machine
import asyncio
import sys
import threading
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor


def section(title):
    print("\n" + title)
    print("=" * len(title))


def io_task(seconds=0.05):
    time.sleep(seconds)          # releases the GIL: perfect for threads
    return seconds


def cpu_task(n=3_000_000):
    total = 0
    for i in range(n):           # pure Python: holds the GIL
        total += i * i
    return total


N = 20


def io_benchmarks():
    section("1. I/O-bound: 20 tasks x 50 ms")

    start = time.perf_counter()
    [io_task() for _ in range(N)]
    serial = time.perf_counter() - start

    start = time.perf_counter()
    with ThreadPoolExecutor(max_workers=N) as pool:
        list(pool.map(io_task, [0.05] * N))
    threads = time.perf_counter() - start

    async def gather_io():
        await asyncio.gather(*(asyncio.to_thread(io_task) for _ in range(N)))

    start = time.perf_counter()
    asyncio.run(gather_io())
    async_time = time.perf_counter() - start

    print(f"serial            {serial:6.3f}s  1.0x")
    print(f"threads           {threads:6.3f}s  {serial / threads:.1f}x")
    print(f"asyncio+to_thread {async_time:6.3f}s  {serial / async_time:.1f}x")
    assert threads < serial / 4, "threads should beat serial I/O by a wide margin"


def cpu_benchmarks():
    section("2. CPU-bound: 4 tasks, short and long")

    for work in (3_000_000, 20_000_000):
        start = time.perf_counter()
        [cpu_task(work) for _ in range(4)]
        serial = time.perf_counter() - start

        start = time.perf_counter()
        with ThreadPoolExecutor(max_workers=4) as pool:
            list(pool.map(cpu_task, [work] * 4))
        threads = time.perf_counter() - start

        with ProcessPoolExecutor(max_workers=4) as pool:
            pool.map(cpu_task, [1] * 4)          # warm up: pay spawn cost here
            start = time.perf_counter()
            results = list(pool.map(cpu_task, [work] * 4))
            processes = time.perf_counter() - start

        print(f"work {work:>10,}  serial {serial:6.3f}s  "
              f"threads {threads:6.3f}s ({serial / threads:5.2f}x)  "
              f"processes {processes:6.3f}s ({serial / processes:5.2f}x)")
        assert threads > serial * 0.6, "threads must not speed up pure-Python CPU work"

    print("checksum:", sum(results))
    assert serial / processes > 1.2, "warmed processes should beat serial by a clear margin"


def race_demo():
    section("3. Race condition: a narrow window, then a wide one")
    rounds, workers = 200_000, 4
    expected = rounds * workers

    def run(label, body):
        state = {"count": 0}
        threads = [threading.Thread(target=body, args=(state,)) for _ in range(workers)]
        start = time.perf_counter()
        for t in threads:
            t.start()
        for t in threads:
            t.join()
        elapsed = time.perf_counter() - start
        print(f"{label:26s} expected {expected:>8}  got {state['count']:>8}  "
              f"lost {expected - state['count']:>8}  ({elapsed:.2f}s)")
        return state["count"]

    def narrow(state):
        for _ in range(rounds):
            state["count"] += 1        # load, add, store -- tiny window

    def wide(state):
        for _ in range(rounds):
            seen = state["count"]      # read
            time.sleep(0)              # yield the GIL: another thread can run
            state["count"] = seen + 1  # write stale value back

    narrow_count = run("counter += 1", narrow)
    wide_count = run("widened window", wide)

    lock = threading.Lock()

    def locked(state):
        for _ in range(rounds):
            with lock:
                state["count"] += 1

    locked_count = run("counter += 1 + lock", locked)

    assert narrow_count >= expected * 0.95, "GIL makes the narrow form look correct"
    assert wide_count < expected * 0.5, "widening the window must expose the race"
    assert locked_count == expected, "the lock must restore correctness"


if __name__ == "__main__":       # REQUIRED for ProcessPoolExecutor on Windows/macOS
    section(f"Python {sys.version.split()[0]} on {__import__('os').process_cpu_count()} CPUs")
    io_benchmarks()
    cpu_benchmarks()
    race_demo()
    print("\nAll concurrency-lab assertions passed.")

Line by line: the module-level guard is not decoration — without it the spawned workers re-import the module and start their own pools; the narrow-window assertion allows a 5% margin, because a rare unlucky thread switch can drop an update even there; io_task sleeps, which releases the GIL, so the 20-thread pool finishes in roughly one task’s time; cpu_task never yields, so threads stall at 1× while processes (warmed with a fake one-iteration job) get real cores; race_demo first proves the narrow += looks correct and then breaks it with time.sleep(0) in the middle, which is the cleanest demonstration that the GIL guarantees bytecode atomicity, not statement atomicity; the final assertions turn each measured claim into a failing test rather than a comment.

Common mistakes

Key takeaways and challenge

Challenge: take a real slow script of yours and decide per section, I/O or CPU. Convert the I/O sections to a ThreadPoolExecutor and the CPU sections to a warmed ProcessPoolExecutor, then measure end to end. If your CPU sections use NumPy, try threads first — the extension releases the GIL and you skip all the pickling.

Want help making a real project concurrent? Ampersand Academy does one-to-one Python training and profiling sessions.

What does the GIL actually prevent?

Only one thread executes Python bytecode at a time, so threads do not speed up CPU-bound Python. The lock is released during blocking I/O and inside C extensions such as NumPy.

When should I use threads instead of asyncio?

Use threads when you must call blocking libraries or share simple state with locks. Use asyncio for thousands of concurrent sockets, and asyncio.to_thread to bridge a blocking call into async code.

Why did my multiprocessing pool run slower than the serial version?

A cold pool spawns fresh interpreters that re-import your module, and on Windows that cost exceeded the work measured. Warm the pool first and give each task enough work to amortise the overhead.

Is counter += 1 thread safe in Python?

No. It is a load, add and store, so two threads can read the same value. The GIL usually makes the window too small to trigger, which is why the bug survives testing; use a lock or avoid shared mutable state.

Do I need the if __name__ guard with ProcessPoolExecutor?

Yes on Windows and macOS, because spawn re-imports the main module in each child. Without the guard children start their own pools, and a script run from stdin fails outright.

Exit mobile version