Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
67 changes: 67 additions & 0 deletions .github/workflows/big-endian.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
# Run the byte-order-sensitive tests on a big-endian host: s390x, emulated with QEMU.
# Emulation makes the full suite take about 40 minutes, so this job runs only the tests
# that exercise data types, codecs and metadata, and only when that code changes (plus
# a weekly run to catch anything the path filter misses).

name: Big-endian

on:
pull_request:
branches: [ main ]
paths:
- "src/zarr/codecs/**"
- "src/zarr/core/buffer/**"
- "src/zarr/core/dtype/**"
- "src/zarr/core/metadata/**"
- "src/zarr/metadata/**"
- "tests/test_codecs/**"
- "tests/test_dtype/**"
- "tests/test_metadata/**"
- "ci/big-endian/**"
- ".github/workflows/big-endian.yml"
schedule:
- cron: "0 3 * * 1"
workflow_dispatch:

permissions:
contents: read

concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true

jobs:
s390x:
name: s390x (QEMU)
runs-on: ubuntu-latest
timeout-minutes: 90
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0 # hatch-vcs needs the tags to resolve a version
persist-credentials: false
- uses: docker/setup-qemu-action@99012661954931238ded8c8b007157a8430204e1 # v4.4.0
with:
platforms: s390x
- uses: docker/setup-buildx-action@f87e5991a6d7451dcb8d9637bfbc97413f497069 # v4.4.1
- name: Build test image
uses: docker/build-push-action@c3c9e263c25d99ce0380d002d59b67737d91b0dc # v7.4.0
with:
context: ci/big-endian
platforms: linux/s390x
tags: zarr-big-endian
load: true
cache-from: type=gha,scope=big-endian
cache-to: type=gha,scope=big-endian,mode=max
- name: Run tests
run: |
docker run --rm --platform linux/s390x -v "$PWD":/src:ro zarr-big-endian bash -euc '
cp -r /src /work && cd /work
pip install -q --no-deps --no-build-isolation .
python -c "import sys; assert sys.byteorder == \"big\", sys.byteorder"
python -m pytest -p no:cacheprovider --hypothesis-profile ci -n auto \
tests/test_dtype tests/test_dtype_registry.py \
tests/test_codecs tests/test_metadata \
tests/test_array.py tests/test_v2.py tests/test_info.py \
tests/test_pipeline_parity.py tests/test_cli/test_migrate_v3.py
'
7 changes: 7 additions & 0 deletions changes/4435.bugfix.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
Fix byte-order bugs found by running the test suite on a big-endian (s390x) host.

- Zarr V3 data type metadata has no byte order, so V3 data types without an explicit `endianness` (such as `Float64()`, or a data type parsed from `"float64"`) now use the host byte order instead of always little-endian. On big-endian hosts, arrays opened from V3 metadata now return native arrays, and `dtype="float64"` and `dtype=np.float64` produce the same data type. Nothing changes on little-endian hosts. Stored chunks are still little-endian by default.
- `ShardingCodec` now defaults its inner and index `bytes` codecs to little-endian, so sharded arrays written with default settings are byte-identical on every host.
- The `scale_offset` codec no longer fails on arrays whose byte order is not the host's, for example a `">f8"` or `">u8"` array on a little-endian host.
- Base64-encoded fill values of structured data types in Zarr V3 metadata are now read and written as little-endian bytes, whatever the byte order of the fields in memory.
- Tests no longer assume that the host is little-endian.
1 change: 1 addition & 0 deletions changes/4439.removal.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
On big-endian hosts, constructing `BytesCodec()` without an `endian` argument is deprecated and raises a `ZarrFutureWarning`. The omitted `endian` currently means the host byte order. A future release will change it to `"little"` on every host, so that stored bytes do not depend on the machine that wrote them. To keep the current behavior, pass `endian="big"`; to adopt the new default now, pass `endian="little"`. Nothing changes on little-endian hosts, where the default is already `"little"`. `BytesCodec.from_dict` and zarr's default serializer are unaffected.
26 changes: 26 additions & 0 deletions ci/big-endian/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
# Test image for a big-endian host (s390x, run under QEMU by .github/workflows/big-endian.yml).
# Debian packages supply numpy and numcodecs, so nothing heavy compiles under emulation;
# this is also the environment Debian tests zarr in. Debian does not package Google's
# crc32c library, so it is built here for the google-crc32c C extension.
FROM debian:sid

RUN apt-get update -qq && apt-get install -y -qq --no-install-recommends \
build-essential cmake git python3 python3-dev python3-pip python3-venv \
python3-crc32c python3-donfig python3-fsspec python3-hatch-vcs python3-hatchling \
python3-hypothesis python3-numcodecs python3-numpy python3-numpydoc python3-packaging \
python3-pytest python3-pytest-asyncio python3-pytest-benchmark python3-pytest-xdist \
python3-tomlkit python3-typer python3-typing-extensions \
&& rm -rf /var/lib/apt/lists/*

RUN git clone -q --depth 1 --branch 1.1.2 https://github.com/google/crc32c /tmp/crc32c \
&& cmake -S /tmp/crc32c -B /tmp/crc32c/build -DBUILD_SHARED_LIBS=1 \
-DCRC32C_BUILD_TESTS=0 -DCRC32C_BUILD_BENCHMARKS=0 -DCRC32C_USE_GLOG=0 \
-DCMAKE_POLICY_VERSION_MINIMUM=3.5 \
&& cmake --build /tmp/crc32c/build -j"$(nproc)" \
&& cmake --install /tmp/crc32c/build \
&& ldconfig \
&& rm -rf /tmp/crc32c

RUN python3 -m venv --system-site-packages /venv
ENV PATH=/venv/bin:$PATH
RUN pip install -q --no-binary google-crc32c google-crc32c msgspec pytest-accept
2 changes: 1 addition & 1 deletion src/zarr/api/asynchronous.py
Original file line number Diff line number Diff line change
Expand Up @@ -1024,7 +1024,7 @@ async def create(
Zarr format 3 only. Zarr format 2 arrays should use `filters` and `compressor` instead.

If no codecs are provided, default codecs will be used based on the data type of the array.
For most data types, the default codecs are the tuple `(BytesCodec(), ZstdCodec())`;
For most data types, the default codecs are the tuple `(BytesCodec(endian="little"), ZstdCodec())`;
data types that require a special [`zarr.abc.codec.ArrayBytesCodec`][], like variable-length strings or bytes,
will use the [`zarr.abc.codec.ArrayBytesCodec`][] required for the data type instead of [`zarr.codecs.BytesCodec`][].
dimension_names : Iterable[str | None] | None = None
Expand Down
2 changes: 1 addition & 1 deletion src/zarr/api/synchronous.py
Original file line number Diff line number Diff line change
Expand Up @@ -762,7 +762,7 @@ def create(
Zarr format 3 only. Zarr format 2 arrays should use `filters` and `compressor` instead.

If no codecs are provided, default codecs will be used based on the data type of the array.
For most data types, the default codecs are the tuple `(BytesCodec(), ZstdCodec())`;
For most data types, the default codecs are the tuple `(BytesCodec(endian="little"), ZstdCodec())`;
data types that require a special [`zarr.abc.codec.ArrayBytesCodec`][], like variable-length strings or bytes,
will use the [`zarr.abc.codec.ArrayBytesCodec`][] required for the data type instead of [`zarr.codecs.BytesCodec`][].
dimension_names : Iterable[str | None] | None = None
Expand Down
41 changes: 39 additions & 2 deletions src/zarr/codecs/bytes.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,13 +3,15 @@
import sys
import warnings
from dataclasses import dataclass, replace
from enum import Enum
from typing import TYPE_CHECKING, ClassVar, Final, Literal

from zarr.abc.codec import ArrayBytesCodec
from zarr.codecs._deprecated_enum import _coerce_enum_input, _DeprecatedStrEnumMeta
from zarr.core.common import JSON, parse_named_configuration
from zarr.core.dtype.common import HasEndianness
from zarr.core.dtype.npy.structured import Struct
from zarr.errors import ZarrFutureWarning

if TYPE_CHECKING:
from typing import Self
Expand All @@ -33,6 +35,30 @@ class Endian(metaclass=_DeprecatedStrEnumMeta):
_members: ClassVar[dict[str, str]] = {"little": "little", "big": "big"}


class _HostEndian(Enum):
"""Marks an omitted `endian` argument, which currently means the host byte order."""

token = 0


def _resolve_host_endian() -> EndianLiteral:
"""The byte order an omitted `endian` argument stands for, warning where it will change.

The default will become `"little"` on every host, so only big-endian hosts are affected.
"""
if sys.byteorder == "big":
warnings.warn(
"BytesCodec() without an `endian` argument stores chunks in the byte order of the "
"host, which is big-endian here. A future version of Zarr Python will default to "
"endian='little' on every host, so that stored bytes do not depend on the machine "
"that wrote them. Pass endian='big' to keep the current behavior, or "
"endian='little' to adopt the new default now.",
ZarrFutureWarning,
stacklevel=3,
)
return sys.byteorder


def _parse_endian(data: object) -> EndianLiteral:
if isinstance(data, str) and data in ENDIAN:
return data # type: ignore[return-value]
Expand All @@ -41,13 +67,24 @@ def _parse_endian(data: object) -> EndianLiteral:

@dataclass(frozen=True)
class BytesCodec(ArrayBytesCodec):
"""bytes codec"""
"""bytes codec

When `endian` is omitted it is the byte order of the host. That default is deprecated on
big-endian hosts: it will become `"little"` on every host, so that stored bytes do not depend
on the machine that wrote them.
"""

is_fixed_size = True

endian: EndianLiteral | None

def __init__(self, *, endian: Endian | EndianLiteral | None = sys.byteorder) -> None:
def __init__(
self,
*,
endian: Endian | EndianLiteral | Literal[_HostEndian.token] | None = _HostEndian.token,
) -> None:
if endian is _HostEndian.token:
endian = _resolve_host_endian()
if endian is None:
endian_parsed: EndianLiteral | None = None
else:
Expand Down
35 changes: 35 additions & 0 deletions src/zarr/codecs/scale_offset.py
Original file line number Diff line number Diff line change
Expand Up @@ -217,12 +217,33 @@ def _decode_float(
return result


def _in_native_byte_order(
arr: np.ndarray[tuple[Any, ...], np.dtype[Any]],
) -> np.ndarray[tuple[Any, ...], np.dtype[Any]]:
"""View-or-copy `arr` in the host byte order; a no-op for native input.

The arithmetic below must see native arrays: numpy returns native results, so a
byte-swapped input would trip the dtype-preservation check, and a byte-swapped dtype
compares unequal to its native twin (`>u8 != uint64` on a little-endian host).
"""
return arr.astype(arr.dtype.newbyteorder("="), copy=False)


def _encode(
arr: np.ndarray[tuple[Any, ...], np.dtype[Any]],
offset: np.generic,
scale: np.generic,
) -> np.ndarray[tuple[Any, ...], np.dtype[Any]]:
"""Compute ``(arr - offset) * scale`` without silent overflow, returning ``arr.dtype``."""
return _encode_native(_in_native_byte_order(arr), offset, scale).astype(arr.dtype, copy=False)


def _encode_native(
arr: np.ndarray[tuple[Any, ...], np.dtype[Any]],
offset: np.generic,
scale: np.generic,
) -> np.ndarray[tuple[Any, ...], np.dtype[Any]]:
"""`_encode` for an `arr` in the host byte order."""
# uint64 is split out first because its full range (up to 2**64-1) doesn't fit in int64,
# so the widening strategy used for every other integer dtype would itself overflow.
if arr.dtype == np.uint64:
Expand Down Expand Up @@ -253,6 +274,20 @@ def _decode(
scale_repr: object,
) -> np.ndarray[tuple[Any, ...], np.dtype[Any]]:
"""Compute ``arr / scale + offset`` without silent overflow, returning ``arr.dtype``."""
native = _in_native_byte_order(arr)
return _decode_native(native, offset, scale, scale_repr=scale_repr).astype(
arr.dtype, copy=False
)


def _decode_native(
arr: np.ndarray[tuple[Any, ...], np.dtype[Any]],
offset: np.generic,
scale: np.generic,
*,
scale_repr: object,
) -> np.ndarray[tuple[Any, ...], np.dtype[Any]]:
"""`_decode` for an `arr` in the host byte order."""
# uint64: same reasoning as _encode — its range exceeds int64, so the Python-int path is the
# only correct option. Exactness check runs first so non-divisible inputs fail before the
# slower object-dtype arithmetic.
Expand Down
7 changes: 5 additions & 2 deletions src/zarr/codecs/sharding.py
Original file line number Diff line number Diff line change
Expand Up @@ -460,8 +460,11 @@ def __init__(
self,
*,
chunk_shape: ShapeLike,
codecs: Iterable[Codec | dict[str, JSON]] = (BytesCodec(),),
index_codecs: Iterable[Codec | dict[str, JSON]] = (BytesCodec(), Crc32cCodec()),
codecs: Iterable[Codec | dict[str, JSON]] = (BytesCodec(endian="little"),),
index_codecs: Iterable[Codec | dict[str, JSON]] = (
BytesCodec(endian="little"),
Crc32cCodec(),
),
index_location: ShardingCodecIndexLocation | IndexLocation = "end",
subchunk_write_order: SubchunkWriteOrder = "morton",
) -> None:
Expand Down
4 changes: 2 additions & 2 deletions src/zarr/core/array.py
Original file line number Diff line number Diff line change
Expand Up @@ -1805,7 +1805,7 @@ def info(self) -> Any:
>>> arr.info
Type : Array
Zarr format : 3
Data type : Float64(endianness='little')
Data type : Float64(endianness=...)
Fill value : 0.0
Shape : (3, 4, 5)
Chunk shape : (2, 2, 2)
Expand Down Expand Up @@ -4023,7 +4023,7 @@ def info(self) -> Any:
>>> arr.info
Type : Array
Zarr format : 3
Data type : Float32(endianness='little')
Data type : Float32(endianness=...)
Fill value : 0.0
Shape : (10,)
Chunk shape : (2,)
Expand Down
9 changes: 7 additions & 2 deletions src/zarr/core/dtype/common.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
from __future__ import annotations

import sys
import warnings
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
Expand Down Expand Up @@ -200,10 +201,14 @@ class HasLength:
@dataclass(frozen=True, kw_only=True)
class HasEndianness:
"""
A mix-in class for data types with an endianness attribute
A mix-in class for data types with an endianness attribute.

The endianness is the byte order of the in-memory array, not the byte order of
stored chunks, which the `bytes` codec sets. Zarr V3 data type metadata carries
no byte order, so it defaults to the byte order of the host.
"""

endianness: EndiannessStr = "little"
endianness: EndiannessStr = sys.byteorder


@dataclass(frozen=True, kw_only=True)
Expand Down
28 changes: 22 additions & 6 deletions src/zarr/core/dtype/npy/structured.py
Original file line number Diff line number Diff line change
Expand Up @@ -583,6 +583,22 @@ def default_scalar(self) -> np.void:
values.append(value)
return self._cast_scalar_unchecked(tuple(values))

def _scalar_bytes_to_void(self, data: bytes, zarr_format: ZarrFormat) -> np.void:
"""Read the raw bytes of a base64 fill value as a scalar of this data type.

Zarr V2 metadata spells out the byte order of every field. Zarr V3 metadata has
none, so V3 fill value bytes are little-endian, like the default `bytes` codec.
"""
dtype = self.to_native_dtype()
stored = dtype if zarr_format == 2 else dtype.newbyteorder("<")
return cast("np.void", np.frombuffer(data, dtype=stored).astype(dtype)[0])

def _void_to_scalar_bytes(self, data: np.void, zarr_format: ZarrFormat) -> bytes:
"""Inverse of `_scalar_bytes_to_void`."""
dtype = self.to_native_dtype()
stored = dtype if zarr_format == 2 else dtype.newbyteorder("<")
return np.asarray(data).astype(stored).tobytes()

def from_json_scalar(self, data: JSON, *, zarr_format: ZarrFormat) -> np.void:
"""
Read a JSON-serializable value as a NumPy structured scalar.
Expand All @@ -606,8 +622,7 @@ def from_json_scalar(self, data: JSON, *, zarr_format: ZarrFormat) -> np.void:
"""
if check_json_str(data):
as_bytes = bytes_from_json(data, zarr_format=zarr_format)
dtype = self.to_native_dtype()
return cast("np.void", np.array([as_bytes]).view(dtype)[0])
return self._scalar_bytes_to_void(as_bytes, zarr_format)
raise TypeError(f"Invalid type: {data}. Expected a string.")

def to_json_scalar(self, data: object, *, zarr_format: ZarrFormat) -> str | dict[str, JSON]:
Expand All @@ -628,7 +643,9 @@ def to_json_scalar(self, data: object, *, zarr_format: ZarrFormat) -> str | dict
string of the bytes that make up the scalar. Subclasses may return
a dict for V3 format.
"""
return bytes_to_json(self.cast_scalar(data).tobytes(), zarr_format)
return bytes_to_json(
self._void_to_scalar_bytes(self.cast_scalar(data), zarr_format), zarr_format
)

@property
def item_size(self) -> int:
Expand Down Expand Up @@ -776,8 +793,7 @@ def from_json_scalar(self, data: JSON, *, zarr_format: ZarrFormat) -> np.void:
return self._cast_scalar_unchecked(tuple(field_values))
elif check_json_str(data):
as_bytes = bytes_from_json(data, zarr_format=zarr_format)
dtype = self.to_native_dtype()
return cast("np.void", np.array([as_bytes]).view(dtype)[0])
return self._scalar_bytes_to_void(as_bytes, zarr_format)
raise TypeError(f"Invalid type: {data}. Expected a dict or base64-encoded string.")

def to_json_scalar(self, data: object, *, zarr_format: ZarrFormat) -> str | dict[str, JSON]:
Expand All @@ -799,7 +815,7 @@ def to_json_scalar(self, data: object, *, zarr_format: ZarrFormat) -> str | dict
"""
scalar = self.cast_scalar(data)
if zarr_format == 2:
return bytes_to_json(scalar.tobytes(), zarr_format)
return bytes_to_json(self._void_to_scalar_bytes(scalar, zarr_format), zarr_format)
result: dict[str, JSON] = {}
for field_name, field_dtype in self.fields:
result[field_name] = field_dtype.to_json_scalar(
Expand Down
2 changes: 1 addition & 1 deletion src/zarr/testing/stateful.py
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,7 @@ def add_array(self, data: DataObject, name: str) -> None:
paths=st.just(parent),
array_names=st.just(name),
zarr_formats=st.just(3),
compressors=st.just(BytesCodec()),
compressors=st.just(BytesCodec(endian="little")),
open_mode="a",
),
label="generated array",
Expand Down
Loading
Loading