HDF5NDSource¶
HDF5NDSource exposes an HDF5 dataset through ProxyNDSource.
Local files use h5py directly. Remote files are scanned once with h5py
to build a native byte-range index, which can also be supplied through hdf5_index=.
Individual chunks are fetched on demand
and converted to Blosc2-compressed chunks stored in the surrounding
Proxy or RemoteArray cache.
The source is assumed immutable (assume_immutable=True). It supports fixed-size
boolean, integer, floating-point, complex, and fixed-length string arrays.
Uncompressed, gzip/deflate, shuffle, and Blosc2 chunks use direct range reads.
Other filter pipelines fall back to retained h5py dataset reads;
hdf5plugin enables its additional registered filters.
Local reads require h5py; hdf5plugin enables additional HDF5 filters.
Install the full HDF5 support with pip install "blosc2[hdf5]". Remote datasets also
need blosc2[fsspec] and the protocol driver, such as s3fs for S3.
For example, blosc2.open("hierarchy.h5::/d0/a2") uses the local reader.
Chunked datasets retain their HDF5 chunk shape; contiguous datasets use
automatically chosen Blosc2 cache chunks. The source remains read-only and
assumes the file is immutable. Its file handle is closed when the source is
garbage-collected.
- blosc2.available_datasets(url, storage_options: dict | None = None) list[str][source]¶
Return all dataset paths in an HDF5 file or native index.
- blosc2.scan_hdf5_index(urlpath, storage_options=None, *, dataset=None, path=None, unsupported=None, traffic=None, _filesystem=None, _lazy_allocations=False, _return_blob=False, _blob=None)[source]¶
Build a native byte-range index for a local or remote HDF5 source.
pathlimits discovery to one dataset, or one group subtree, plus its ancestor groups. The returned dictionary is JSON-compatible and can be supplied viahdf5_index=on later opens.- Parameters:
urlpath¶ (str or path-like) – Local path or fsspec URL of the immutable HDF5 source.
storage_options¶ (dict, optional) – Options passed to the fsspec filesystem.
path¶ (str, optional) – Build a scoped index for this dataset or group subtree. By default, index the complete container.
dataset¶ (str, optional) – Supported alias of
path; both must agree after stripping outer slashes.
- Returns:
A JSON-compatible native HDF5 index.
- Return type:
dict
- blosc2.validate_hdf5_index(index, urlpath=None, *, dataset=None, path=None)[source]¶
Validate and return a native HDF5 index.
urlpathchecks the recorded source URL.path(ordataset) additionally checks that a scoped index describes the requested dataset. Version-1 and version-2 native indexes are accepted.
- class blosc2.HDF5NDSource(urlpath, dataset: str | None = None, *, path: str | None = None, hdf5_index=None, storage_options=None, max_concurrency=8, blocks=None, cparams=None, _traffic: Traffic | None = None, _filesystem=None, _blob=None, _ensure_allocations=None, _source_cache_dir=None, _index_explicit=True)[source]¶
Read one immutable HDF5 dataset as Blosc2-compressed logical chunks.
Select it with keyword-only
pathor the supporteddatasetalias. If both are supplied they must agree after stripping outer slashes.- Attributes:
attrsThe user attributes of the source.
blocksThe block shape of the source.
chunksThe chunk shape of the source.
cparamsThe compression parameters of the source.
dtypeThe dtype of the source.
metaThe fixed-length metadata of the source.
shapeThe shape of the source.
vlmetaThe variable-length metadata of the source.
Methods
aget_chunk(nchunk)Return the compressed chunk in
selfasynchronously.get_chunk(nchunk)Return the compressed chunk in
self.close
- __init__(urlpath, dataset: str | None = None, *, path: str | None = None, hdf5_index=None, storage_options=None, max_concurrency=8, blocks=None, cparams=None, _traffic: Traffic | None = None, _filesystem=None, _blob=None, _ensure_allocations=None, _source_cache_dir=None, _index_explicit=True)[source]¶