feat: Updates for 10x Genomics Atera bundles - #426
stephenwilliams22 wants to merge 9 commits into
Conversation
Adds a native reader for 10x Genomics Atera datasets: cell-by-gene expression table, cell/nucleus segmentation labels and boundary polygons, per-transcript locations, and morphology images. - Table (`cell_feature_matrix.zarr.zip`/`csc_cell_feature_matrix.zarr.zip`) is read directly with `anndata`'s own zarr IO (`anndata.io.read_elem`/ `anndata.io.sparse_dataset`), with `X` lazily backed by dask. - Cell/nucleus labels are lazily backed by dask; boundary polygons are read via `shapely.from_ragged_array`, with an optional tiled/pyramided partial-read path (`read_cell_boundaries`) for large bundles. - Transcripts are read lazily per spatial tile via dask-delayed. - Morphology images are read via `tifffile`'s native OME-TIFF tile grid, avoiding materializing full-resolution planes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #426 +/- ##
==========================================
- Coverage 65.79% 60.17% -5.63%
==========================================
Files 26 28 +2
Lines 3263 3761 +498
==========================================
+ Hits 2147 2263 +116
- Misses 1116 1498 +382
🚀 New features to boost your workflow:
|
|
@stephenwilliams22 could you please always ensure that the pre-commit checks are green? Ideally a PR should always have green CI before anyone has a look. Thanks! |
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@Zethson sorry about that. we should be good to go now. |
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds a native reader for 10x Genomics Atera datasets: cell-by-gene expression table, cell/nucleus segmentation labels and boundary polygons, per-transcript locations, and morphology images.
cell_feature_matrix.zarr.zip/csc_cell_feature_matrix.zarr.zip) is read directly withanndata's own zarr IO (anndata.io.read_elem/anndata.io.sparse_dataset), withXlazily backed by dask.shapely.from_ragged_array, with an optional tiled/pyramided partial-read path (read_cell_boundaries) for large bundles.tifffile's native OME-TIFF tile grid, avoiding materializing full-resolution planes.Enables massive datasets on a basic laptop (2M cells, 11 billion transcripts). Basic analysis workflow (loading, plotting, HVG, PCA, normalization, clustering, etc) max memory usage was 7GB on a new mac in ~15min.