Reliable Visual Regression Testing for Humans and Coding Agents
See how JupyterLab made visual regression testing reproducible on Linux machines, cut CI from 55 to 15 minutes and let contributors update snapshots.
See how JupyterLab made visual regression testing reproducible on Linux machines, cut CI from 55 to 15 minutes and let contributors update snapshots.
JupyterLab uses visual regression testing to catch unintended changes to the interface before a release: it compares about 350 reference screenshots (snapshots, in Playwright's terms) on every pull request. The tests use Galata, JupyterLab's test framework built on Playwright. Until February 2026 the suite took 55 minutes, and now it takes 14 to 16. Flaky tests, which fail once and pass on retry, went from 17 per run in January to 2.4 in August. Any contributor can now request new reference images with a comment. A local run on Linux with the CI fonts produces the same pixels as CI, and on Fedora 44 with its default fonts, 85% of the screenshots match.
In the three months to 25 September 2026, 28 of the 30 merged pull requests that added or changed a UI test in JupyterLab were AI-assisted.1 A developer who works with an agent on interface code needs a test that runs on their machine, a result they can trust, and an answer in minutes.
This post describes what we changed and which parts you can reuse for Playwright tests on GitHub Actions.
| Before | Now | |
|---|---|---|
| Suite run time | 42 min to 1 h 25 min, 55 min on average | 14 to 16 min |
| Regenerating screenshots | about 45 min, started by a maintainer | about 1 min, requested by any contributor |
| Local run matches CI | no | yes, with the CI fonts |
| Looking at a failure | download and unzip the report, start a web server | click a badge in the pull request |
| Flaky tests per run | 17 (January 2026) | 2.4 (August 2026) |
| Runs with a hard failure | 54% (January 2026) | 14% (August 2026) |
Fixing the flaky tests uncovered real bugs. Many tests were flaky because of a bug in JupyterLab that appeared only under CI timing. Examples are a notebook that took focus while it initialized, a race when several settings change quickly and a debugger that did not show its variables after F9. While pinning the fonts for the tests, we found three parts of the interface that ignored the configured font. All 19 bugs found this way are fixed.2
We also ran the Chromium tests on two setups other than the CI runner: Fedora 44 with its default fonts, and the Ubuntu 26.04 runner that CI will have to move to.3
| Before | Now | |
|---|---|---|
| Fedora 44: screenshots that match CI | 1% | 85% |
| Fedora 44: tests that pass | 56% | 93% |
| Ubuntu 26.04 runner: screenshots that match CI | 35% | 100% |
| Ubuntu 26.04 runner: tests that pass | 71% | 99.8% |
If you would like to submit a PR to JupyterLab with a coding agent's assistance:
A screenshot of the interface is mostly text, and text rendering depends on the machine. We found these causes:
install-deps installed about 80 system packages on CI, among them fonts that developer machines do not have (e.g. xfonts-cyrillic). The theme asked for system-ui first, so the font used depended on the installed packages.JupyterLab 4.6 fixes all of them. The Galata helper extension ships its own fonts as npm dependencies, so the lockfile pins their versions:
"@fontsource/dejavu-sans": "^5.2.5",
"@fontsource/dejavu-mono": "^5.2.5",
"@fontsource-variable/noto-sans-sc": "^5.2.10"
It applies them over the system fonts, and sets the properties that the browser would otherwise choose:
:root {
--jp-code-font-family-default: 'DejaVu Mono' !important;
font-kerning: normal;
-webkit-font-smoothing: none;
-moz-osx-font-smoothing: none;
font-optical-sizing: none;
}
The tests start Chromium with --disable-lcd-text, which turns off subpixel antialiasing, and --disable-webgl, which makes the terminal use its DOM renderer.
We stopped running install-deps. The only font it installed that the tests needed was one for Chinese characters, and Noto Sans SC in the list above took its place.
The terminal emulator, xterm.js, has no API for kerning or text rendering, so the tests set them on the canvas context:
ctx.fontKerning = 'normal';
ctx.textRendering = 'geometricPrecision';
The data grid did not apply the configured font-family to its canvases. Users saw this too, so we fixed it in JupyterLab.
The console banner is off in the default Galata settings. IPython now prints a static banner when SOURCE_DATE_EPOCH is set.
The new fonts changed every existing screenshot, and one pull request updated 310 reference images. No version of the fonts reproduced the old images, and neither did fonts copied from the CI runner. The update was due anyway: moving CI from Ubuntu 22.04 to 24.04 broke about 150 tests through font versions alone. With the fonts pinned, a runner upgrade no longer changes the fonts in the screenshots.
We ran the Chromium tests in Fedora 44 and openSUSE Tumbleweed containers on GitHub Actions, against the reference images from the Ubuntu runner. With the runner's font packages installed (DejaVu, Liberation, Lato and Noto Color Emoji), both matched all 255 screenshot comparisons. openSUSE also needed the three fontconfig rules that turn off hinting for small DejaVu text, which Ubuntu and Fedora ship with the font.
With the fonts that Fedora 44 installs by default, 39 screenshots still differ, all of Mermaid diagrams. JupyterLab shows a Mermaid diagram as an SVG image, and an image cannot use the fonts of the page, so its text uses a system font. Two Vega charts had the same problem because they asked for sans-serif, and the test chart now sets DejaVu Sans in its Vega config.
We tried reusing the Linux reference images on the macOS and Windows runners of GitHub Actions, with a sample of 20% of the screenshots, and at first 4% matched. We overrode navigator.platform, because macOS menus showed shortcuts as symbols and were up to 63 pixels narrower. We also turned off glyph hinting and subpixel positioning on Linux with --font-render-hinting=none and --disable-font-subpixel-positioning. After that no element differed in size, but only 14% of the screenshots matched on macOS and 4% on Windows. FreeType, CoreText and DirectWrite draw the same glyphs differently, and neither the operating systems' font smoothing settings nor Chromium's --text-contrast and --text-gamma switches changed that. The option left is a tolerance per platform: with maxDiffPixelRatio: 0.01, 59% of the failing screenshots would pass on macOS and 53% on Windows.
The old workflow for new reference images rebuilt JupyterLab, ran the whole suite with --update-snapshots and pushed the result. A maintainer had to start it, and it took about 45 minutes to produce a few images.
The failed run already has those images. Playwright writes the actual screenshot next to the expected one for every failed comparison, and its JSON reporter records which reference file each one belongs to. We enabled the JSON reporter on CI:
reporter: process.env.CI
? [['blob'], ['json', { outputFile: 'test-results/report.json' }]]
: [['list'], ['html', { open: 'on-failure' }]],
A script, unpack_snapshots.py, copies each image to the path of its reference.
Now a contributor writes this comment on their pull request:
please open PR to update snapshots
After a maintainer approves the run, a bot waits for the test run of the head commit to finish, takes the screenshots from its artifacts, and opens a pull request against the contributor's branch. A new request on the same pull request replaces the one in progress. The contributor accepts the new images by merging it. After the test run, the bot takes about a minute. The time grows with the number of changed screenshots and does not depend on the number of tests.

The bot covers Galata screenshots, JSON snapshots, documentation screenshots and example snapshots.
Pushing to the contributor's branch needs a token with write access, in a job that anyone can start with a comment. For that reason the old workflow was limited to maintainers. Our bot commits to its own fork and opens a pull request against the contributor's branch, so it has no write access to the repository. The write happens when the contributor merges.
A GitHub App cannot do this. An App can open pull requests inside its organisation, but not against a fork owned by someone else, so we use a plain bot account with a personal access token.
The workflow also limits what it does with files from the pull request:
galata/**/*-snapshots/ and examples/**/*-snapshots/) are accepted. The job that holds the token checks each path again, and refuses a path that a symbolic link redirects.--no-verify alone skips the pre-commit hook, but not post-checkout or post-commit.A Playwright HTML report is a directory of HTML, scripts, screenshots and videos. As a GitHub Actions artifact it is a zip, so to look at a failure you downloaded it, unzipped it, started python -m http.server and opened localhost. Most reviewers did not.
In February 2026, GitHub Actions added artifact uploads without a zip (archive: false), and it serves a single uploaded HTML file directly. We asked Playwright to make its report self-contained, and the maintainers declined because the need is niche. Our action, inline-playwright-report, does it instead. The report keeps its assets in a base64 zip inside index.html, so the action rewrites the src attributes inside that zip.
The bot comment on every JupyterLab pull request now has a badge: green when all tests pass, orange with the number of flaky tests, red with the number of failures. The link opens the report filtered to the flaky or failed tests.

When a pull request adds or changes a test, the workflow runs that test again with video recording, and the video goes into the same report.
A pull_request workflow from a fork cannot write comments. The usual solution is a second workflow, triggered by workflow_run, that has write permissions and reads an artifact from the first one. Code from the fork produced that artifact, so its content is untrusted. The comment built from it appears under a trusted bot account.
The comment takes two values from the artifact: the link to the report, and the failing and flaky counts. Without checks, a fork could make the trusted bot post a phishing link, or end the markdown link early and add its own text to the comment. The action makes these checks:
new URL() instead of inserting the string. This removes newlines and encodes angle brackets, so the value cannot escape the markdown link.https://github.com/you/yourrepo/blob/<sha>/evil.html can be attacker content.Unknown otherwise.We also run zizmor on the workflow files, in CI and as a pre-commit hook. It correctly flags the workflow_run trigger, so that line has an inline exemption that explains the reason.
Flakiness measured on pull requests includes the effects of each change. Since December 2025, JupyterLab runs the suite every six hours on main. The inputs are the same every time, so any difference between runs is flakiness.

Every Monday, a script reads the last 28 scheduled runs and posts a table to a public issue. The table lists which tests failed, which passed only on retry, how often, and in which browser. It shows when the count starts to rise again.
To check a fix, you can start the workflow by hand with a test name pattern for Playwright's --grep and a repeat count of up to 50.
Most flaky tests, outdated screenshots and quirks of the snapshot updating workflows had one of a few causes, and the contributing guide now lists best practices for writing UI tests:
expect.soft. Otherwise the first failure hides the others, and the update takes several CI rounds.waitForTimeout().Lint rules enforce some of these:
no-wait-for-timeout, no-element-handle, no-networkidle, prefer-to-have-count and prefer-web-first-assertions.no-restricted-syntax rules against screenshot({ path }), because the result should go through toMatchSnapshot(), and against test.describe.configure({ mode: 'serial' }), because serial tests cannot be split across shards.The Playwright rules highlighted 118 violations in 36 files. We enabled it gradually, suppressing the existing violations with inline eslint-disable comments and slowly working through the old ones in follow-up PRs.
With six shards per browser, the suite takes 15 minutes instead of 55.4 Set fullyParallel: true, or Playwright cannot balance the shards.

The setup step runs once per shard, so each extra minute there costs six minutes of runner time. We replaced two lines:
- playwright install-deps
- playwright install chromium
+ playwright install chromium --only-shell
install-deps installs about 80 apt packages that headless tests do not use, and the apt step failed often enough to cancel jobs. If you test with Firefox or WebKit, remove install-deps but do not add --only-shell: it applies to Chromium only.
We also removed Python test dependencies that the browser tests did not use, and a second frontend build. The browser cache key pointed at a missing file, so the cache was not refreshed when Playwright was updated. We fixed the key.
When one shard fails, you can re-run only that shard. A re-run produces blobs for its own shard only, so the merge job starts from the merged blobs of the previous attempt.
Two composite actions in jupyterlab/maintainer-tools (BSD-3-Clause) work in any repository:
These parts need a few changes before you can reuse them:
galata, the test directory name in its paths, with your own.@fontsource packages and CSS from an entry point that only the tests use.waitForTimeout calls remain behind eslint-disable comments.This work was a collaboration with Quansight PBC. At OpenTeams, @krassowski led the work, and @MUFFANUJ and @Darshan808 worked on the flaky test fixes, the CI tooling and the lint rules.
The Jupyter Foundation funded this work. @jtpio opened the 2023 issue that described the problem, @bollwyvl asked for a readable CI report, and @jasongrout set up the scheduled runs that give the flakiness numbers in this post.
Thank you to the reviewers: @jtpio, @jasongrout, @brichet, @Yann-P, @HaudinFlorence and @mfisher87.
Counted from the AI usage section of the JupyterLab pull request template. The section asks whether AI generated some or all of the content. The count covers pull requests merged into main from 25 June to 25 September 2026 that changed a test file in galata/test. It leaves out backports and pull requests from bots. The other two pull requests did not answer. ↩
Found through flaky tests:
Found while pinning the fonts:
Measured on GitHub Actions on 26 September 2026 with the Chromium tests of the jupyterlab project, in a Fedora 44 container and on the Ubuntu 26.04 runner. "Before" is JupyterLab's main branch on 4 February 2026, before the first change of this work, with its own lockfile and Playwright version, and Python packages as of that date. A test or screenshot that passes on the retry counts as passing, as on CI.
The median of the 51 runs on main from the update to Playwright 1.58 on 26 January 2026 to the change to six shards on 4 February. After the update, the median Firefox job took 50 minutes instead of 42. The median Chromium job took 43 minutes instead of 42. From October 2025 to the update, the median run took 43 minutes. The Firefox shards are still about 20% slower than the Chromium shards, so we compare with the runs that used the same Playwright version. ↩
Benchmarks, deep dives, and lessons from OpenTeams engineers. Leave your email, or follow the feed.