Skip to content

Chapter 24

The tests

Since chapter 1 this book has cited tests, many by name; this chapter shows their suite whole. By its end you will know its seven layers, from one routine of the game to the finished page in two browsers, each with what it compares and against what; the machinery that lets a run of the suite give the same answers however its tests are spread over a machine; why the suite runs in two phases, and how one run of it is held to another; and what a green run proves, and what it leaves to the eye and the ear.

Seven layers

The suite is about 930 tests under tests/, run by pytest, Python's framework for tests. We arrange them here in seven layers, groups of tests that each hold the port at one scale against one reference, because a fault is best caught at the smallest scale that can see it, where its failure says the most: a routine and an input under the oracle, a pass in the open loop, a fault of the picture only in a browser.

Layer What it compares Against Told in
one routine results, flags and touched memory, on random inputs the original under the emulator; the ROM's floating point; the blitter model 5, 12
the original observed the instrument's claims; the notes' findings; a key's effect the headless original itself 6, 19
the front end files, music and drawing calls, VBlank by VBlank the headless original's run 19
the missions state, drawing, chance, palettes and sound, every pass and tick a recording of the original 8, 18
the whole game replayed the state's hash after every input sample hashes stored with a demo 17
the core itself its two builds, save states, the page file, the generated files itself and the repository 22
the page requests, pixels, box, clock, sound, saves, pause, fullscreen what the core gave it 23

The layers follow the specification's levels of verification where tests exist, the drawing within the routine's layer and the sound within the missions' and the front end's; the picture's level is not made as such, the replay's hashes covering the display memory (chapter 22), and the core itself is added. Chapter 21 grouped the same modules by what each is about; here they stand by what each compares.

Two dashed frames, the emulator phase round six layers and the page phase round the page, and a legend: held against the original, the port's own record, or what the core gave the browser.

The seven layers and the two phases they run in. A phase is a way of running the suite, not a layer: the first six layers run together over the processor cores, the page alone after them.

One routine

Eight modules, one for each milestone that ported pure routines and one for the floating point, hold the routines one at a time by chapter 5's differential tests: the original's code under the emulator and the port's C in the native library on the same random inputs, compared in their results, from M3 on in the memory either side touched, from M4 on in the flags. The later modules also assert that their cases ran the code no mission script reaches. The floating point is held to the ROM's library in three builds, native, as WebAssembly and under a checker of undefined behaviour; the blit through its register programme, replayed by chapter 12's model of the blitter.

The original observed

Seven modules hold the headless original to what it claims and observe the notes' findings under it, two of them holding the port's sound as well. Two runs of one run description give the same dump; an observer changes no step; the model of Paula changes no step of a run in which nothing sounds. A run description of tests/runs/ that presses a key is compared with one that does not, the pair's baseline, so that whatever differs belongs to the key.

The front end and the missions

One module replays the front end: the headless original's schedule of the title sequence and the menus goes through the port VBlank by VBlank, the fades taking no time, since the harness's contain no wait and the port's are a setting (chapter 7), and every file opened, call of the music and drawing call must come at the same VBlank. The music is held the same way through the outer loop, on and switched off, by a module counted with the original's own tests.

Seven modules fly the missions. Each mission script is first made into a recording: the script run once under the headless original with a dump after every pass and tick, observers on the drawing routines, the reads of the entropy stream and the addresses each tick wrote. The port replays it in chapter 8's two loops, the open and the closed, comparing after every pass and tick what chapter 8 listed, with the completeness lists, every map's setup and the tables' capacities beside them. A recording is the costliest thing the suite makes, one mission's 38 seconds in chapter 6, and all the tests of its script share it.

The whole game, replayed

One module plays back a demo the port recorded, chapter 17's, kept in tests/replays/ with its seed file, the schedule of its playback and a hash of the whole save state after every input sample: from the program's start, through the front end and the idle rank selection, into the attract demo's playback of a short recorded flight. Played in the native library and in Node's WebAssembly, it must give the stored hashes in both: the whole program from its start into a mission, and the page's own build held where the loops and the oracles run natively. The state holds the display memory (chapter 22), so a change in what is drawn changes the hashes; but they are the port's own record, never compared with the original.

The core itself

Eight modules test the port as a program, held to its own rules, since the original has nothing to compare them with. Two hold the core's two builds to each other: four VBlanks a tick, as many audio frames as the emulated time lasts, the same picture, a WebAssembly that imports nothing; a state saved in one core loads into a second, and a foreign one is refused. A third continues states saved inside a mission in a new core on another seed (chapter 22). The rest hold the page file, which asks the network for nothing, the generated files, the ROM check's messages, the keyboard assist's rules and the tests' own isolation.

The page

Two modules open the finished page from a file:// address, as a double click opens it, in Chrome and Firefox without a window, headless. A program in Node drives each browser through its remote control, Chrome's DevTools protocol and Firefox's WebDriver BiDi, a W3C specification still in draft, over the WebSocket Node has built in, so that nothing need be installed. Keys are pressed through the driver, never dispatched from a script in the page, since a scripted event activates nothing and a page not activated may not start its sound (chapter 23). The browsers are muted from outside, because the owner works at the same computer (chapter 10), and from outside so that the page's sound still flows to be measured.

A driver drives its browser once and prints what it measured; the tests in Python judge it, reading the geometry off the page's layout, never from the shell's word, so that the shell cannot vouch for its own mistakes. The exact picture the shell hands them is the framebuffer through each row's palette, unscaled (chapter 23), where the canvas holds it enlarged and reduced. What the drivers hold:

Area What is held
the file it loads only itself; the console shows no error
the picture the box in the machine's proportions, in both standards and after resizes; the canvas and a screenshot against the exact picture, on both renderers
the clock 50 VBlanks a second within 2 on PAL, 60 on NTSC
the sound nothing built before an activating key; music and effects arriving as audio frames that are not silent
the game the stick's keys; a mission flown from the keyboard; the enemy's fighter found in the sky by its frames' pixels, since no colour alone told it apart
the storage a game saved and loaded after a reload; the demo recorded, reloaded and played as recorded
the shell's own the help screen, the pause sign, a moment hidden, an absence, fullscreen

Under a true scale factor

Only a screenshot shows what the browser did after the shell was finished. Chrome can pretend a screen's density through its protocol. It then places the canvas on whole CSS pixels, and where the box's edge falls on half a CSS pixel the picture is shifted by a device pixel or resampled once more, an artefact that would hide a real fault. So the scale factor, two device pixels to a CSS pixel as on a Retina screen, is given on Chrome's command line, where Chrome behaves as on a real display, and the window puts the box's edge on half a CSS pixel, the case the pretence gets wrong.

Three questions are asked of a screenshot that no few chosen points could answer. Every pixel: at the centre of the block each framebuffer pixel is shown as, the screenshot must carry its colour. Where the picture lies: its position, fitted from its own colour edges to a fraction of a device pixel, must be the box the page reports. Whether the edges are hard: between two neighbouring blocks' centres, at most one device pixel may be of neither colour, where one smooth step would smear the edge over most of a block. Here is the first; look at lines 6 and 7, which find the screenshot's pixel at the centre of every block, and at the last line, which keeps the worst colour channel, allowed 2 in 255 and found at 0:

# tests/picture.py, lines 98-107
def whole_picture_differences(source, screenshot, box):
    """For every framebuffer pixel, how far the screenshot is from it at its block's centre."""
    left, top, width, height = box
    rows, columns = source.shape[:2]
    xs = block_centres(left, width, columns)
    ys = block_centres(top, height, rows)
    assert xs.min() >= 0 and xs.max() < screenshot.shape[1], 'the box is not in the screenshot'
    assert ys.min() >= 0 and ys.max() < screenshot.shape[0], 'the box is not in the screenshot'
    shown = screenshot[np.ix_(ys, xs)].astype(int)
    return np.abs(shown - source.astype(int)).max(axis=2)

The three hold only where a framebuffer pixel is at least three device pixels each way, and refuse below that rather than pass, for there a block's centre carries its neighbours. At a scale factor of 2 a window of 1280 by 900 CSS pixels gives four across and seven and a half down; at ordinary density the window must be large, as the headless Firefox runs make it.

The frame time and the visible window

The frame-time test walks the page into a mission at a screen's size and times every animation frame for five seconds. It demands a mean interval of at most 1.05 refreshes, at most five late frames, present() at most a millisecond on average, and the WebGL renderer. It measures the refresh first, on a blank page, since a page missing every other refresh would otherwise judge itself by twice the refresh. What it reads, the overlay's player line, the canvas and present(), every build of the shell offered, so it can measure the previous build too, which lacks the WebGL renderer and so must fail: a test that has never failed may be unable to (chapter 8). In a visible Firefox it fails on the frames, the present and the renderer alike (chapter 23's table).

The visible window is a mode of the Firefox module switched on by hand, nineteen tests that open a real window, run alone. It is the one check that can see a canvas fault of the graphics processor: headless Firefox composites, puts the page's layers together into the screen's picture, in software, and the fault that once gave Firefox a black picture happens only on the graphics processor. The window must stay in front and uncovered, because a covered window stops the page's clock, so it runs while the owner is away; and on a locked screen macOS moves no window into fullscreen, so its fullscreen check skips, and the small-window check behind it with it.

What makes a run hold

Each module, one Python file of tests, shares what tests/conftest.py holds: the options, the build, the core reached through ctypes, Python's way of calling C, and the fixtures that keep the tests apart. pytest first collects the tests, listing all of them before any runs, and calls functions of the suite at fixed points, its hooks; the collection hook runs once the list is made, no kin of chapter 5's hook in the emulator.

Built once

The suite builds the port itself, from the sources as they stand. pytest-xdist, a plugin, spreads the tests over test processes: processes of their own, each running a share of the tests with its own copy of the core. Each would build, and the build writes the native library in place, where a process that has it loaded while another rewrites it can crash or read a torn file. So the first test process to take a lock in the directory they share builds and leaves a file to say so, and the others find it and build nothing. The code calls a test process a worker, pytest-xdist's word; look at the lock on line 14 and at needed(), asked while it is held, which asks whether that file is still missing:

# tests/conftest.py, lines 111-125
def once_per_run(tmp_path_factory, needed, make):
    """make() if needed(), once per run of the suite however many processes run it.

    Under pytest-xdist every worker is a process of its own with a session of its own, and
    would make the same thing again, over files the others are reading.  The workers take a
    lock in the directory they all share and ask needed() while they hold it, so the first
    one makes and the others find it made.  Without xdist it is the plain call."""
    if not os.environ.get('PYTEST_XDIST_WORKER'):
        if needed():
            make()
        return
    with open(tmp_path_factory.getbasetemp().parent / 'wof-once.lock', 'w') as lock:
        fcntl.flock(lock, fcntl.LOCK_EX)
        if needed():
            make()

A control of chapter 8, built from a changed copy of the sources into a library of its own, points the suite at that library, which it loads beside the build there is, building nothing.

A fresh core for every test

A process holds one copy of the core, and chapter 9 told how a test that left it changed broke the next. So every test that takes the core starts from a fresh core: the core reset to its start and what lives beside its state put back, for about two thirds of the suite, at under five thousandths of a second each. The initialisation resets the state and two things beside it; the rest a test sets from outside:

Beside the core's state Put back by
the VBlanks of a fade step and of a pass the values read at the process's first test that takes the core
the callbacks into the test at a tick, a pass and the setup's end clearing them, and every replay as it ends
the pokes, the map list's addresses clearing them, and every replay as it ends
the stand-ins reached, the trace, the tests' copies of the state clearing them
the sound event log, the files written the core's own initialisation
the audio output rate nothing: every test that renders names its rate

This is the settings' half of chapter 9's fresh_core, which calls it; look at line 29, which reads the two settings at the process's first such test, and at lines 30 to 38, which run before every one:

# tests/conftest.py, lines 1320-1356
def fresh_settings():
    """What lives beside the core's state and survives wof_init, put back to what the process
    held before any test ran (re/notes/testing.md, "A fresh core for every test"): the
    VBlanks of a fade step and of a pass, read off the core at the first call, and the test
    instrumentation tests/shim.c sets in src/trace.c - the three hooks, the pokes, the map
    list's addresses, the stand-ins reached, the trace records and the snapshots at step S
    and at a pass's end, which only this reset forgets, never one inside a test.  The audio
    output rate survives too and is left alone: every test that renders names its rate
    before the VBlanks it takes, and a new rate empties the queue."""
    lib = _SETTINGS.get('lib')
    if lib is None:
        lib = ctypes.CDLL(str(DYLIB))
        for name, argtypes, restype in (('wof_fade_vblanks', [], ctypes.c_int),
                                        ('wof_set_fade_vblanks', [ctypes.c_int], None),
                                        ('wof_vblanks_per_pass', [], ctypes.c_int),
                                        ('wt_set_vblanks_per_pass', [ctypes.c_int], None),
                                        ('wt_set_tick_hook', [ctypes.c_void_p], None),
                                        ('wt_set_step_s_hook', [ctypes.c_void_p], None),
                                        ('wt_set_pass_hook', [ctypes.c_void_p], None),
                                        ('wt_pokes_clear', [], None),
                                        ('wt_map_addresses', [ctypes.c_void_p, ctypes.c_uint], None),
                                        ('wt_standins_reset', [], None),
                                        ('wt_trace_reset', [], None),
                                        ('wt_snapshots_reset', [], None)):
            function = getattr(lib, name)
            function.argtypes = argtypes
            function.restype = restype
        _SETTINGS.update(lib=lib, fade=lib.wof_fade_vblanks(), per_pass=lib.wof_vblanks_per_pass())
    lib.wof_set_fade_vblanks(_SETTINGS['fade'])
    lib.wt_set_vblanks_per_pass(_SETTINGS['per_pass'])
    for hook in ('wt_set_tick_hook', 'wt_set_step_s_hook', 'wt_set_pass_hook'):
        getattr(lib, hook)(None)
    lib.wt_pokes_clear()
    lib.wt_map_addresses(None, 0)
    lib.wt_standins_reset()
    lib.wt_trace_reset()
    lib.wt_snapshots_reset()

So the start is rebuilt before every test, and a replay also clears what it set when it ends, even by an exception. tests/test_isolation.py sets every item of the instrumentation, resets, and looks for each, every assertion naming what would have outlived its test:

# tests/test_isolation.py, lines 86-103, a part of test_the_fresh_core_leaves_no_test_state_behind (lines 23-103)
conftest.fresh_settings()
ported.reset_core()

ported.set_g('mission_number', 1)
ported.set_g('rank_played', 2)
lib.wof_test_poke_after_rank()
lib.wof_test_poke_after_reset()
lib.wof_test_poke_after_map()
assert (ported.g('mission_number'), ported.g('rank_played')) == (1, 2), 'a poke outlived its test'
assert lib.wof_env_map_address() == MAP_LIST_ADDRESS, 'a map address outlived its test'
assert lib.wt_standin_count() == 0, 'a stand-in count outlived its test'
assert lib.wt_trace_count() == fresh_trace, 'a trace record outlived its test'
lib.wof_test_tick_end(8)
assert ticks == [7], 'a hook outlived its test'
assert lib.wt_globals_get(1, state, len(state)) == -1, 'the step-S snapshot outlived its test'
assert lib.wt_pass_view(1) == -1, "a pass's snapshot outlived its test"
with pytest.raises(AssertionError, match='step S not reached'):
    ported.g_at_mission('mission_number')

The rest the test processes share, the files, the caches, the core and the browsers, was audited by hand, item by item (re/notes/testing.md).

Recordings made once

A recording is kept in the test process that made it, under its script, the VBlanks of a pass and its pokes, for every later test that needs it. Spread over test processes, each process that ran one of a script's tests would record it again; so the collection hook gives the loop tests one group per script and pass rate, and pytest-xdist sends a group whole to one test process. The group is a marker: a label pytest attaches to a test, by which a run selects, skips or groups it. Look at lines 4 to 7; the rest of the hook comes back in the next section:

# tests/conftest.py, lines 75-95
@pytest.hookimpl(tryfirst=True)
def pytest_collection_modifyitems(config, items):
    for item in items:
        group = recording_group(item)
        if group:
            item.add_marker(pytest.mark.xdist_group(group))
    # Without the Kickstart ROM nothing that needs it or the build can run, and the build takes
    # the system font from it: every test but those marked `without_rom` skips, with the one
    # message of tools/rom.py as its reason, as a missing browser skips its module.
    missing = romcheck.problem()
    if missing:
        skip_rom = pytest.mark.skip(reason=missing)
        for item in items:
            if 'without_rom' not in item.keywords:
                item.add_marker(skip_rom)
    if config.getoption('--slow') or os.environ.get('WOF_SLOW') == '1':
        return
    skip = pytest.mark.skip(reason='slow; run with --slow or WOF_SLOW=1')
    for item in items:
        if 'slow' in item.keywords:
            item.add_marker(skip)

Measured with eight test processes:

Tests sent Recordings made in more than one process Processor time spent on the repeats
any test to any free process 60 4,225 s
a script's tests as one group 35 1,165 s

What is still made twice belongs mostly to the completeness tests, which use every script of a milestone and so belong to no one group. The headless original's limits on the wall clock, which only ever end a recording (chapter 6), were set from times measured with the suite in parallel.

Two test processes: the build under a lock, a fresh core before each test, a recording kept for its group.

A test process's life: the build once for all, a fresh core before every test that takes one, and a recording made once and kept for its group.

Two phases

The page tests measure time and pictures, the clock, the start of the sound and what a browser really shows, and load from outside the browser disturbs all three: seven page tests once failed while another program used more than a processor core, and passed on a quiet machine. The emulator tests, so named because most run the headless original under its emulator, keep every processor core busy. So the suite runs in two phases, one after the other and never side by side, which are not chapter 7's phases of a write: the emulator phase over the processor cores, then the page phase alone.

The suite registers three markers. page marks the two browser modules, and -m 'not page' and -m page split the suite between the phases. slow marks about fifty of the longest differential runs, run only with --slow. without_rom marks the four tests that need no ROM; without it, the hook above skips every other test with the ROM check's one message as its reason, and nothing is built.

The emulator phase runs eight test processes, one for each of the machine's processor cores, with --dist loadgroup to keep the groups. The times below, all with --slow, were measured on an earlier suite of 827 tests, the serial row with its page tests:

Run Test processes Wall time, h:mm Processor time, hours Result
the whole suite, serially 1 3:00 3.0, one process 812 passed, 15 skipped
the emulator phase, machine idle 8 0:51 5.7 736 passed, 2 skipped
the emulator phase, the owner at work 12 1:11 10.7 734 passed, 2 failed
the emulator phase, the owner at work 6 1:36 7.6 736 passed, 2 skipped

Eight take about twice what the same tests take serially: a processor core runs slower when all are busy, and the recordings made twice add the rest. Twelve share processor cores, and their two failures were a recording past the limit of a whole recording as it then stood, ten minutes, since raised to half an hour; six leave processor cores idle. Both ran with the owner at work, so their times are upper bounds.

The serial run, every test in one process, was the first reference run, the run another is held to by its outcome set: every test by its id, its module, name and parameters, with its outcome, passed, failed, an error, a test that could not run, or skipped with its reason. Held to it, a parallel run shows that splitting the suite lost no test and ran none twice, and that no outcome depends on the order or the process a test ran in, which is how chapter 9's shared core came to light. pytest writes the outcome set into a junit file, an XML report; tools/junit_compare.py reads the first file as one set and the others together as the other, prints every test only one side has and every outcome that differs, and fails on any. Look at line 8, which drops what pytest-xdist appends to a grouped test's id, the group, which says where the test ran, not what it is:

# tools/junit_compare.py, lines 22-39
def outcomes(paths):
    """{test id: outcome} over the files, and the ids seen more than once."""
    seen = {}
    twice = []
    for path in paths:
        for case in ET.parse(path).getroot().iter('testcase'):
            test = '%s::%s' % (case.get('classname'), case.get('name').split('@', 1)[0])
            outcome = 'passed'
            for child in case:
                if child.tag in ('failure', 'error'):
                    outcome = child.tag
                    break
                if child.tag == 'skipped':
                    outcome = 'skipped: %s' % (child.get('message') or '').strip()
            if test in seen:
                twice.append(test)
            seen[test] = outcome
    return seen, twice

One run, counted

The release's run, and a fresh clone's, the ROM copied in and the setup script run, its emulator phase on half the machine's processor cores, since the owner was at work:

Run Test processes Emulator phase Page phase
the release, with --slow 8 804 passed, 3 skipped, in 1:01 103 passed, 20 skipped, in 0:23
a fresh clone, without --slow 4 752 passed, 55 skipped, in 0:56 103 passed, 20 skipped, in 0:22

The slow tests, the fuel script, the runs at other pass rates and the campaign's chains among them, ran only in the release's run, which so held the loops over every script. Its three skips are two long runs of the headless original and the stored replay's recorder, run only when asked for; the twenty are the visible window's nineteen tests and a lost WebGL context tried on the 2D renderer, which has none to lose. The visible window, run alone, skipped only its fullscreen check and the small-window check behind it. The clone built the page byte for byte the same. Held to the older serial reference, no test's outcome differed; the tests on one side only were those added since, the key layer's cases renumbered and a few page tests split by renderer, so the two-phase run became the reference.

The last table is counted at the book's build from pytest's collection, so that it is never stale; the oracle's and the missions' layers are the largest:

Layer Modules, with their tests Tests
one routine test_oracle_m1.py 70, test_oracle_m3.py 81, test_oracle_m4.py 40, test_oracle_m5.py 15, test_oracle_m6.py 26, test_oracle_m7.py 10, test_oracle_m8.py 23, test_oracle_ffp.py 45 310
the original observed test_headless.py 37, test_frontend.py 39, test_passes.py 4, test_objects.py 4, test_map.py 5, test_sound.py 5, test_music.py 19 113
the front end test_front_port.py 65 65
the missions test_world.py 38, test_weapons.py 42, test_enemy.py 57, test_campaign.py 18, test_loader.py 15, test_demo.py 12, test_mission.py 40 222
the whole game replayed test_replays.py 3 3
the core itself test_core_native.py 9, test_core_wasm.py 20, test_state_m4.py 5, test_dist.py 6, test_isolation.py 3, test_assist.py 47, test_generated.py 2, test_rom.py 3 95
the page test_page.py 66, test_firefox.py 59 125
the suite 34 modules 933

The test modules by what they compare, each layer's with its tests, counted from pytest's collection at the book's build.

What a green run proves

A green run here is both phases with --slow, as the release's was, the visible window run by hand beside them; the clone's run without --slow proves less. It says that after every logic tick and every pass of every mission script, in both loops, on every map the scripts reach, through the campaign's chain, the saved games and the demo, the port's state, drawing calls, draws of chance, palettes and sound events are the original's, and every address the original writes in a mission is compared or listed with its reason. The front end agrees VBlank by VBlank; every pure routine agrees with the original's on the inputs tried, the floating point with the ROM's; both builds of the core are one program. The page shows the framebuffer exactly, keeps the clock and the sound, and keeps the saved games and the demo through a reload, in two browsers. That is chapter 1's definition, logic, picture and sound, wherever an instrument reaches.

What it does not say belongs to the proof. No pixel of a mission scene is compared with the original: the headless original draws nothing, and the drawing calls, the rows' palettes and the blitter model stand for the picture. That model is the chip's documented behaviour, not derived from the original, and no emulator exact to the cycle has checked it; its line mode holds the original's line routine to the port's, so the game's one line, the cable, rests on the documentation too (chapter 20). Code no script runs is held by the oracle's cases where a test can hold it, or is a stand-in that would fail loudly. A fade's step and a pass's VBlanks are settings, one set by eye, the other filmed in a quiet scene; the music's tempo rests on a timer's assumed starting value, and a directory's order on the documentation (chapter 1). No test drives Safari, where the page also runs. Fullscreen on a real screen and the smoothness of scrolling as a person sees it were the owner's eyes, the sound as heard the owner's ears (chapter 10).

Every test is named for what it holds, and a run of the suite holds what each of the book's claims from the instruments rests on: that is why this book could cite a test by name for each.

For the developer

The switches are environment variables: WOF_SLOW_HEADLESS=1, the two long determinism runs; WOF_REPLAY_WRITE=1, the stored replay recorded anew; WOF_CORE_LIBRARY, a control's library; WOF_FRAMES_PAGE, the frame-time test on another build; WOF_FIREFOX_VISIBLE=1, the visible window. The table of modules comes from book/tools/suite.py and book/suite.toml.

What comes next

The chapter in one sentence: the suite holds the port at seven layers, from one routine to the page in two browsers, keeps every test free of what ran before it, and a green run proves chapter 1's definition wherever an instrument reaches, leaving a mission's pixels to the blitter model and the eye and the ear to the owner. Chapter 25 hands it to you: the setup, the ROM, the build and both phases on your own machine, the tools, and how to change the port without breaking the proof.

Further reading

The files named here are in the repository, github.com/sy2002/wof-wasm.

Outside the repository: pytest and pytest-xdist; the Chrome DevTools Protocol and WebDriver BiDi; Martin Fowler's "Eradicating Non-Determinism in Tests", on tests that leave state behind.