Deterministic testing for multithreaded Python

· LWN.net

17 min read Original article ↗
This article brought to you by LWN subscribers

Subscribers to LWN.net made this article — and everything that surrounds it — possible. If you appreciate our content, please buy a subscription and make the next set of articles possible.

Python's support for multithreaded programs has improved considerably over the last few years with the advent of the "free-threaded" version of the language. But testing multithreaded programs is notoriously difficult, because the underlying host system determines the thread-execution ordering, which adds an element of non-determinism. At PyCon US, Larry Hastings gave a talk (YouTube video) about his blanket project, which is meant to provide mechanisms for deterministic testing of multithreaded Python code.

He began with a scenario where attendees were working on a new Python library, probably destined for the Python Package Index (PyPI). They had heard about "this no-GIL thing"—free-threaded Python which does not have a global interpreter lock (GIL)—and wanted to ensure their new library would work in a multithreaded environment. So they look into it and realize that "all I need to do is add a couple of locks"; he said that would likely suffice in many situations.

But sometimes life becomes more complicated and a race condition is found. That is a solvable problem, he said; often it just requires adding some code to recognize that the condition is occurring and handle it in some fashion. Once that is done, another problem rears it head, however: test coverage.

In 2026, projects should have a unit-test suite with 100% coverage, Hastings said. But that is hard to do when there is a line in the code that only runs rarely, under weird circumstances that are completely outside of the control of the developer. He showed a test-coverage report with a 99% that "is sitting there mocking you"; "you can't sleep at night, you've failed, your life is over". The answer is that "you are going to get it to 100% because you are going to use something called 'blanket'".

Blanket

Blanket is a new library, so new in fact that he had done the initial release the night before the talk. Why that name? "Because it is coverage using threads", he said, to groans, laughter, and applause.

[Larry Hastings]

Using the Python threading module is "giving yourself a problem", he said. While that is "a little glib", there are seven synchronization primitives in the module (e.g. threading.Lock(), threading.RLock(), and threading.Semaphore()) but they are all non-deterministic. For example, if three threads all want to acquire the same lock, the order in which they do so is unknown—it depends on the scheduler in the operating system. "You cannot predict what order your threads are going to run in."

The solution is to use blanket's deterministic replacements of the same seven primitives that are all available from the module's Scenario object (e.g. scenario.Lock()). That will make the program deterministic and allow the developer to control the order of events. It will "make any race condition reproducible on demand".

He showed a short Python program that appears in the "Quickstart" section of the blanket documentation. It creates three threads (using threading.Thread objects), each of which is going to acquire() and release() a Lock (implicitly using a with block) and then wait() on a Barrier object. A barrier is initialized with a number of waiters, when that number of threads have called wait() for the barrier, all of the waiters are allowed to run—in some order. For his example, the threads each use the worker() function as the thread; it makes the synchronization calls and reports that it acquired the lock and that it is past the wait:

    import random
    import threading
    
    lock = threading.Lock()
    barrier = threading.Barrier(3)

    def worker(name):
        with lock:
            print(f"worker {name} got the lock")
        barrier.wait()
        print(f"worker {name} is past the barrier")

The three threads, named "A", "B", and "C", are added to a list, which is shuffled (using random.shuffle()), used to start() the threads, and, finally, used to await their completion using join():

    A = threading.Thread(target=worker, args=('A',))
    B = threading.Thread(target=worker, args=('B',))
    C = threading.Thread(target=worker, args=('C',))

    threads = [A, B, C]
    random.shuffle(threads)
    
    for t in threads:
        t.start()
    for t in threads:
        t.join()

As can be seen in successive runs of the program, the order of which threads acquire the lock and get past the barrier is different each time, which he emphasized with four slides showing different orderings. For example:

    worker C got the lock
    worker A got the lock
    worker B got the lock
    worker B is past the barrier
    worker C is past the barrier
    worker A is past the barrier

There are six separate orderings for the lock acquisition and six for getting past the barrier, which makes for 36 different possible orderings of the test. If a race condition only happens in one of those orderings, it is difficult (or painful) to reproduce it reliably, but blanket can help.

He went through the program, modifying it to use blanket. The first step is to import blanket and then to create the scenario object, lock, and barrier:

    scenario = blanket.Scenario()
    lock = scenario.Lock()
    barrier = scenario.Barrier(3)

The worker(), thread creation, and shuffling operations are all the same from the earlier version of the code. The idea is that the worker() threads still do exactly what they did before, but will be using blanket versions of the lock and barrier. It is in the latter part where things change more:

    lock_api = scenario.api(lock)
    barrier_api = scenario.api(barrier)

    with scenario:
        for t in threads:
            t.start()
        list(lock_api.relay(B, A, C))
        lock_api.unblock(lock.release, C)
        with barrier_api.cycle(C, A, B):
            pass
    for t in threads:
        t.join()

The lock_api and barrier_api objects provide a "control surface" to manage the behavior of the underlying object. The code then "enters the scenario" via the with scenario: block and starts the threads as before. The next four lines are where the major differences lie; they use the high-level blanket API.

The relay() function of the lock_api object returns an iterator that steps through its thread arguments, releasing the lock (if it is held), then acquiring the lock, in the thread-order specified. It actually yields each thread after the acquisition step, but that is not used in the code, which just calls list() to iterate. That does mean that the final lock release is not done, which is what the unblock() call handles. The cycle() call on the barrier_api object prepares the threads for waking up after the barrier; more complicated things can be done inside that with block, but by default the order of the arguments governs the order in which the threads are woken up after the barrier.

The most important thing to note from the code is that the lock should be acquired in the BAC order and the barrier should be exited in the CAB order. As might be guessed, his next slide showed exactly that (as did the next three for emphasis).

Scenario

An object of the Scenario class is at the top level in blanket (e.g. scenario = blanket.Scenario()) and all seven of the threading synchronization primitives have blanket equivalents that can be reached from the object (e.g. lock = scenario.Lock()). One interesting thing about the primitives in blanket is that they are classes that are bound to the scenario instance as a bound inner class. He likens those to methods, which are likewise bound to their instance, and believes they solve a lot of problems, though "I have been using these for years and nobody seems to care".

The wrapper objects that blanket returns look the same as those returned by threading, which he calls "masquerading". He has added a way to tell the difference for debugging purposes, however.

    >>> t_lock = threading.Lock()
    >>> b_lock = scenario.Lock()
    >>> t_lock
    <unlocked _thread.lock object at 0x7c62de036d10>
    >>> b_lock
    <unlocked _thread.lock object at 0X7C62DE1CE510>
The ID for a blanket lock is printed in upper case, so that is a "tell", he said. Otherwise, both t_lock and b_lock are instances of threading.Lock, though of course only b_lock is an instance of scenario.Lock.

The scenario provides three levels of APIs: low-level, medium-level, and high-level. For the most part, users will stick with the high-level API, Hastings said, but they need to understand the other two levels, as well. The high-level abstraction is convenient, "but it is important to understand the low-level abstractions underneath you because all abstractions are leaky sooner or later".

The "transaction" is one of the "foundational, low-level concepts" that underlie blanket. It is "an object that represents a particular call to a particular method on a particular primitive at a particular time". If a thread calls Lock.acquire(), the transaction represents the thread, the lock object, and the call type (acquire). Transactions are created frequently, "under the covers"; he showed what one "looks like" by creating a scenario.Lock(), acquiring it, and retrieving an entry from the log that the scenario object maintains:

    <Lock.acquire RETURNED 'MainThread'
    start_time=377537.982058657 result=True
    blocking=False for Lock 1 0X79C14A4D2660>

It shows that the transaction is for an acquire, it is in the RETURNED state, ran on the main thread, and started at a specific time. The result holds the return value or the exception if one was raised. The transaction is not currently blocking and it was done on Lock 1, which is an internal name that blanket automatically gives to a lock. Locks can be renamed by test suites, which makes it easier to read the test output, he said.

Transactions are state machines that only operate in a forward direction. "If you are in a particular state, you know you will never visit a previous state." Outside of a scenario, a transaction simply moves from the COMMIT state, meaning that the underlying method has been called, to the RETURNED state. But inside the scenario (i.e. in a with scenario: block), transactions switch from "unregulated" to "regulated". There are two additional states each inserted ahead of the original two, so BLOCKED comes before COMMIT and PAUSED comes before RETURNED.

The two new states, which are called "parking states", are what allow blanket to control the behavior. They are entered before a blocking call to the underlying synchronization primitive in order to await permission to continue. The way tests work with blanket is that the main thread acts as the scheduler for the worker threads that are using the synchronization primitives. The main thread is running the unit-test suite, orchestrating the order of operations between the worker threads from within the scenario. After the test completes, the main thread exits the scenario and checks the assertions for the test case.

The most surprising thing he found out while developing blanket was that he could simply wrap the existing synchronization primitives and not have to simulate or reimplement them. All that turned out to be needed was the two stopping points: block and pause. Blocking before every method call is "the bread and butter of blanket"; it allows controlling what happens. "You don't need to control the internal state, just who calls the lock, in what order." That implicitly controls the state of the lock, Hastings said.

So blanket uses the real threading primitives, but wraps them. It blocks before making each call, allowing blanket to establish the order for the threads to call into the primitive. It also pauses after the call, which is what allows blanket to adjust the order of which threads resume after using a Barrier or similar primitive.

The other fundamental concept in blanket is scenario.wait(), which is a method of the scenario object that is modeled after "a wonderful API from Win32" called WaitForMultipleObjects(). The wait() method is passed a number of "signaling objects" that represent a binary signal that is either true or false. If all of the objects passed are false, the caller sleeps until at least one of them becomes true, which causes wait() to return a set of the passed-in objects that have signaled.

There are multiple kinds of signaling objects that can be passed, starting with a thread handle, which will signal if a transaction is running on it; that allows the main thread to see what the transaction is. The thread handle can be wrapped in a Terminated() object, so it will signal if the thread terminates. The Call() object can be used to signal when a specific thread is calling a particular method on a given primitive (e.g. Call(thread, lock.acquire)). A transaction handle can be passed to wait for the transaction to complete; they can be wrapped in Paused() to signal only when the transaction is in that state. The Not() wrapper can be used to negate the signal being returned, as well.

Low to high

Hastings presented a simplified version of his starting example in order to show the levels of the blanket API. The first part of the example remained the same through five separate rewrites:

    import blanket
    scenario = blanket.Scenario()
    lock = scenario.Lock()

    def worker(name):
        with lock:
            print(f"worker {name} got the lock")

    A = scenario.thread(worker, 'A')
    B = scenario.thread(worker, 'B')
    C = scenario.thread(worker, 'C')

Instead of using a lock and a barrier, the simplified version simply acquires and releases the lock, while reporting that it did so. The threads he uses in the example are slightly different as well; they are "managed" threads using scenario.thread(), which wraps the underlying threading.Thread objects. A distinguishing feature of the scenario.thread() objects is that they are started (i.e. thread.start()) when the scenario is entered and joined (thread.join()) when it is exited. The first low-level version was:

    with scenario:
        for thread in (B, A, C):
            for i in range(2):
                scenario.wait(thread)
                tx = scenario.transaction(thread)
                tx.unblock()
                scenario.wait(tx)

That produced the expected output (as would all of the examples): "B", then "A", then "C" acquire the lock. It steps through the threads in order, then does a series of operations twice, once for acquire and the second for release. It waits for the thread to enter a transaction, retrieves the transaction, unblocks it so it can continue, and waits for the transaction to complete. That is expressed a bit more directly in the next version:

    lock_api = scenario.api(lock)

    with scenario:
        for thread in (B, A, C):
            lock_api.unblock(lock.acquire, thread)
            scenario.wait(lock.release)
            lock_api.unblock(lock.release, thread)

This version uses the lock API object to unblock an acquire operation on the thread. It then waits for a release operation (on any thread, but the test only has a single possibility) and unblocks it. The next version is at the middle level of the blanket API:

    lock_api = scenario.api(lock)

    with scenario:
        for thread in (B, A, C):
            d = scenario.Driver(thread)
            d()
            d.skip()
            d()
            d.finish()
            d()

It uses a Driver object to push the thread through the various states; each call to d() drives the thread to the next point where a decision is needed. skip() and finish() calls provide the decisions to unblock the acquire and release operations. The next middle-level-API version uses a driver underneath:

    lock_api = scenario.api(lock)
    
    with scenario:
       for thread in (B, A, C):
          scenario.skip(thread, lock.acquire, lock.release)

The scenario.skip() call is passed a thread and multiple operations and will drive the thread to complete those operations by, effectively, running the driver as in the previous example. His last version, which uses the high-level blanket API is quite terse:

    lock_api = scenario.api(lock)

    with scenario:
        for thread in lock_api.relay(B, A, C):
            pass

That uses the relay() mechanism that was employed in the example at the start of the talk. It drives the threads through the acquire and release operations on the lock. These high-level APIs are tailored to the type of the primitive, so relay() is used for Lock and RLock objects, while cycle() is used for Condition, Event, and Barrier objects (as was also seen in the first example). Semaphore and BoundedSemaphore have an allocate() function to drive them through their acquisition and release.

Moar features

After that whirlwind tour of blanket, Hastings went through a list of other features, starting with timeout management. That allows programs to cause timeouts to happen (with expire()), cause them not to happen (with disregard()), or restore the timeout to its original state (revert()). He also briefly described "raw handles", which are synchronization objects that avoid the management blanket provides; the scheduler can interact with the underlying primitives that way.

Each blanket primitive (e.g. scenario.Lock()) is actually made up of four separate objects: the object masquerading as a threading primitive, the API object (e.g. scenario.api(lock)), the raw handle (scenario.raw(lock)), and an internal-only core object. Each of the blanket methods needs to acquire a single lock before they do their job. That lock lives in the Scenario core ("score") object and is called score.lock. That lock is a bit like the GIL, but it is lightly contended; it is a deliberate tradeoff to make blanket deterministic, he said.

He completed the list with two different code-injection mechanisms that blanket supports. scenario.inject() can be used to monkey-patch blanket into a module that directly uses the threading primitives, so that the module will use blanket instead. "Some galaxy-brained individual" who is writing their own synchronization primitives can use blanket.injector() to insert calls into the Python bytecode of those primitives, which will allow blanket-style testing of the custom synchronization mechanism.

Hastings completed his prepared talk with the credo of blanket, which is:

Your test should be effectively single-threaded. If it isn't, you haven't blanketed hard enough. Slow it down.

Before he took questions, he had a few extra slides to present. The first was about whether there are other technologies similar to blanket. "Only kind of"; there is a "rich area of computer-science research" into something called "stateless model checking" (SMC). It is like blanket, but instead of writing a scheduler, SMC systematically or stochastically explores the different combinations of ordering between the various synchronization mechanisms. As he understands it, one of the main focuses of current research is on how to reduce the size of the space that needs to be explored, "because it explodes quickly". He mentioned CHESS, Shuttle, and Loom as examples. The closest to blanket is perhaps Coyote, Hastings said.

His final slide was a picture of Mickey Mouse as a sorcerer from or inspired by Fantasia. As might be guessed, he used that as the backdrop to describe his use of "AI" (presumably LLMs) on the project. It was his first time working with those tools; "it has been an enormous amount of work and I was enormously more productive". The project has been "pair-programmed with AI"; "the tests are kind of vibe testing and the docs are kind of vibe docking". It was so productive, in fact, that he rewrote blanket five times along the way, making it simpler and improving the API as he did. Along those lines, Hastings gave a presentation at EuroPython 2026 (YouTube video) about blanket 2.0, which he has not, as yet, shipped.

There was a question about performance, but that is not something he cares much about because speed is not really a goal of the project, Hastings said. Blanket is meant for unit tests, not for running in production; if there are obvious inefficiencies, he will fix them if he can without harming the "intentionally not fast" use case. As the credo notes, the idea is to slow things down.

[I would like to thank the Linux Foundation, LWN's travel sponsor, for its assistance with my trip to Long Beach, CA for PyCon US.]

Index entries for this article
ConferencePyCon/2026
PythonFree-threading
PythonTesting