Fix benchmark workflow - #1349
Merged
Merged
Conversation
asv's default build command runs `pip wheel` without `--no-deps`, which drops a wheel for every dependency into the build cache directory. asv then aborts with "Found multiple wheels in ... Cannot decide correct one", so build the project wheel on its own.
xarray will change `Dataset.dims` to return a set of dimension names, so indexing it by name only works today via a FutureWarning.
asv changes its default build command between releases, which is how the benchmark job broke without anything changing in this repo. Its version is also absent from the recorded results, so an upstream change to the timing machinery would silently shift the baseline rather than fail.
The counting functions return lazy dask arrays, so compute the result to measure the counting itself rather than the building of the task graph. Graph construction accounted for the 0.12s these have reported since 2021, while the counting it stood in for takes roughly three times as long. Warm up the numba kernels on a tiny dataset in setup, so that jit compilation is not measured by the first timed run.
Contributor
|
Tick the box to add this pull request to the merge queue (same as
|
Contributor
Author
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1349 +/- ##
=========================================
Coverage 100.00% 100.00%
=========================================
Files 46 46
Lines 2994 2994
=========================================
Hits 2994 2994 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
jeromekelleher
approved these changes
Sep 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

The benchmarks workflow has been red. Two independent problems:
benchmarks/asv.conf.jsonleavesbuild_commandcommented out, so asv usesits default, which runs
pip wheelwithout--no-depsand drops a wheel forevery dependency into the build cache. asv 0.6.6 turned more than one wheel
there into a hard error:
First commit sets
build_commandexplicitly, which also stops the jobdepending on asv's defaults. A later commit pins asv, since its version is
not recorded in the results.
count_call_allelesandcount_cohort_allelesreturn lazy dask arrays andthe benchmarks never computed them, so they were timing the construction of
the task graph rather than the counting — which is why they have reported an
almost unchanging 0.12s for years. Both now compute the result, with a numba
warmup in
setupso jit compilation is not measured.The second change moves them from ~0.12s to ~1.1s and resets their history, as
the old values were measuring something else.
Verified by a full successful run of the workflow on a fork.