Repository navigation
Tests: a fast tier by default, a reliable 4-minute slow tier behind --slow - #4
Merged
Merged
Conversation
The suite took 11:22 on four cores, and 25 tests (61 with their parameters) were two thirds of it: the twelve tutorials alone 350 s, the pages' figures 50 s, and a handful of numerical checks at 4-19 s each. They are marked `slow` and skipped unless `--slow` is given, so an ordinary run is 3:40 serially and about 1:10 with `-n 4`; `--slow -n 4` runs everything in 5:42, as before a PR. Under xdist, each worker capped OpenBLAS at one thread: four workers that each started four took 9:36, barely better than one. pytest-xdist joins the `test` extra. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM
…work Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM
A GUI test that drops a widget in a reference cycle left it for the cyclic collector, which runs in whichever thread allocates next -- in the failing runs, a live fit's TIFF reader thread. ~QWidget closes its window, and QWindow::close waits for the GUI thread to flush window-system events; the GUI thread was the test, waiting for that fit, so the worker hung (py-spy: QWindowSystemInterface::flushWindowSystemEvents under tifffile in the reader). That was test_live failing in about half of the parallel --slow runs, and the 300 s timeouts in test_gui_view3d and test_gui_render_axes. conftest disables automatic collection for the session and collects every 100th test on the main thread, as pyqtgraph's GarbageCollector does; the fast tier's time is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM
Work stealing by default (conftest's pytest_xdist_make_scheduler): xdist's own scheduler gave one worker the tutorials -- 372 s of tests against ~100 s for each of the others -- and --slow -n 4 took 6:50; stealing takes 3:45. With --slow the slow tests run first, so the long ones start early. The tutorial test runs the storyboards in draft (SMAPPY_TUTORIAL_DRAFT): shots at 1x and JPEG, a quarter of the pixels and 120 ms to encode where WebP took 550. The published tutorials are unchanged. CLAUDE.md: the slow tier is parallel again now that it does not hang, and says that automatic collection is off in the tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The suite had grown to 1322 tests and took 11:22 on four cores. Running it in parallel hung or failed in about half of the runs.
pytest tests -n autopytest tests -n auto --slowNo test was removed or weakened.
Changes
Two tiers. About 60 tests are marked
slow: the tutorials, the docs-figure tests and the 11 heaviest numerical checks. They took two thirds of the time and are skipped unless--slowis given. They are skipped rather than deselected, so the summary still counts them.The parallel hang, fixed. A GUI test that left a Qt widget in a reference cycle left it for Python's cyclic garbage collector. The collector runs in whichever thread allocates next; here that was a later test's TIFF reader thread. The widget's destructor calls
QWindow::close, which waits for the GUI thread to flush window events. The GUI thread was the test, itself waiting for that fit, so the worker deadlocked. py-spy showedQWindowSystemInterface::flushWindowSystemEventsundertifffilein the reader thread. This caused:test_livefailing in about half of the parallel--slowruns;test_gui_view3dandtest_gui_render_axes.The bug predates this PR:
main's setup failed the same way. The fix is pyqtgraph'sGarbageCollectorremedy:conftest.pyturns off automatic collection and collects on the main thread every 100 tests. Before the fix, about half of the parallel runs failed; after it, 8 of 8 passed. The fast tier's time is unchanged.Balanced parallel runs. Three changes, all in
conftest.py:pytest_xdist_make_schedulerhook. xdist's own scheduler gave one worker 372 s of tutorials while the other three idled after about 100 s.--slowis given.Draft tutorials in the test.
SMAPPY_TUTORIAL_DRAFT=1takes screenshots at 1× instead of 2× and saves them as JPEG (120 ms per shot) instead of WebP (550 ms). The published tutorials are unchanged. Most of a tutorial's remaining time is the GUI really working (fits, renders, simulating its data), which this test is there to exercise.Smaller changes:
pytest-xdistadded to thetestextra.--slowresult; and notes that automatic garbage collection is off in the tests.Test results
--slow -n auto: 1321 passed, 4 skipped (3:39, 3:44, 3:44 in the last three runs)-n auto: 1260 passed, 65 skipped (1:24)🤖 Generated with Claude Code
https://claude.ai/code/session_01NjvC6By9NAqXj9KRfwWdLM