RemoteStore

RemoteStore discovers a read-only B2Z, Zarr, HDF5 or Parquet source and returns RemoteArray and RemoteCTable leaves. Groups and leaves share one source session: a B2Z archive, a native HDF5 index, or a Zarr store. Zarr listing remains lazy.

The default CachePolicy.MEMORY shares a 256 MiB allowance across all leaves. Set max_cache_bytes to a positive integer to change it. CachePolicy.NONE retains no payload and rejects a limit. Passing cache_dir selects DISK when the policy is omitted; an explicit policy must agree with the cache location. DISK accepts max_cache_bytes=None for unbounded retention. Sources must be immutable. Generic blosc2.open(..., lazy=True, dataset=...) continues to open a single array.

A Parquet file has one CTable at its root. Pass allow_table_root=True to retain a store handle, then use store[""] to access the table. Parquet has no child selectors. Converted physical-column/row-group caches share the store’s policy, allowance, traffic counter, persistence, and refresh generation.

For HDF5, hdf5_index= accepts a native index dictionary, local JSON path, or remote fsspec URL. An explicit index skips hierarchy discovery and must match the source URL and selected scope. See Working with Remote Arrays for the sidecar-generation workflow.

with blosc2.RemoteStore(
    "https://host/data.h5", cache_policy=blosc2.CachePolicy.NONE
) as store:
    print(store.keys())  # immediate children
    info = store.get_info("experiment")  # metadata only
    group = store["experiment"]
    array = group["temperature"]
    values = array[:100]
    result = (array + 273.15).compute()

# Returned handles own their source lifetime independently.
values = array[:100]
group.close()
array.close()

Paths are relative to the selected group. store["a/b"] and store["a"]["b"] use the same source reader. Each lookup returns an independent handle. With NONE, repeated reads fetch again; no payload cache is retained. dataset="a" or a subgroup suffix in the URL selects a group at construction. An array root must be opened with RemoteArray instead.

keys() and get_info() do not construct leaf readers or payload caches. Discovery can read archive prefixes, attributes and small HDF5 inline values. get_info() returns a RemoteNode with a relative path, a kind (group, ndarray, ctable, remote_store or unsupported), known attributes and a diagnostic. Unknown array attributes are None; open the array to retrieve them. Unsupported nodes stay discoverable and raise NotImplementedError when selected. Missing paths raise KeyError.

Group attrs mappings are read-only. source returns the credential-free container descriptor and full group path. traffic is one shared source counter across all views, including discovery; do not add counts from aliases. It counts source reads, not connections or every HTTP request: HEAD requests and failed Zarr probes are excluded. cache_bytes counts retained compressed chunks and partial-block duplicates across the store, without double-counting aliases. The least recently used native chunk is evicted across all leaves when necessary. Oversized reads return their result before eviction; returned arrays and temporary buffers are outside the allowance. Closing a leaf preserves its warm cache while other session handles remain open. With NONE, cache_bytes is zero and max_cache_bytes is None.

Closing a handle, or exiting its context, releases its ownership. Existing child handles remain usable until closed or garbage-collected. The last handle closes the owned archive/store wrappers and private HTTP/S3 transport sessions. Operations on an explicitly closed handle raise RuntimeError. Standalone RemoteArray exports remain self-contained references, including the native HDF5 index when applicable.

Nested stores

Assigning a RemoteStore to a TreeStore persists a lazy, credential-free reference at the exact path supplied by the caller:

with blosc2.RemoteStore("s3://weather/europe.zarr", dataset="spain") as remote:
    with blosc2.TreeStore("catalog.b2z", mode="w") as tree:
        tree["/external/weather"] = remote

Remote B2Z discovery reports that object root as remote_store without opening the linked source. Lookup through the mount supports direct and chained paths. B2Z, HDF5, Zarr v2 and Zarr v3 targets and subgroup references are supported.

The outer RemoteStore overrides saved cache defaults and owns one policy, traffic counter and aggregate allowance across mounted sources. Pass nested_storage_options as a URL-to-options mapping or a callable accepting a source descriptor when mounted sources need different credentials. A local TreeStore can open one reference with runtime options through open_remote().

save() preserves nested references and includes only already-retained warm payload from mounts that were opened. It never opens an unvisited mount to save it. materialize(destination) follows all reachable mounts and writes one independent local .b2z or .b2d TreeStore. Materialization detects cycles, limits nesting to 64 levels, and publishes the destination atomically.

b2view uses RemoteStore for remote hierarchies with one 64 MiB MEMORY allowance, and RemoteArray for selected or directly opened leaves. Switching selection releases the UI handle while retaining the store’s warm chunks.

For persistent shared caching:

with blosc2.RemoteStore("https://host/data.h5", cache_dir="remote-cache") as store:
    with store["experiment/temperature"] as array:
        values = array[:100]
    print(store.cache_bytes, store.metadata_bytes)
    store.refresh()  # rebuild discovery; old child handles become stale

Each source and selected root has its own directory under cache_dir. One owner holds an exclusive operating-system lock until its last dependent handle closes. Conflicting opens raise RuntimeError, including in other processes; the operating system releases the lock after a process exits or crashes.

Reopening restores all previously created leaf caches and trims them against the new aggregate allowance before returning. The manifest preserves B2Z directory and bounded metadata reads, one native HDF5 index, and lazily discovered Zarr metadata. Metadata reads can contain small inline values or incidental bytes in bounded prefixes; they are separate from evictable payload. metadata_bytes is the encoded manifest size, and is zero without a disk manifest. Credentials and storage options must be supplied again at runtime.

The allowance measures compressed retained payload, not filesystem allocation. Container overhead and old generations are outside it; obsolete generation files are removed on the next exclusive reopen. Manifests survive payload eviction. Store backing files are private implementation details; use array save() or to_cframe() for standalone exports, optionally with include_cache=True.

Sources must remain immutable until explicit root refresh(). Refresh builds replacement discovery before publishing a new generation; failed discovery or publication leaves existing handles valid. Successful refresh makes existing child groups and arrays stale, requiring fresh lookups. Corrupt manifests and B2Z validator mismatches raise an error; use a fresh cache directory if the store cannot be opened. Fully offline reopening is not promised. The lifetime lock uses POSIX flock or Windows byte locking; Windows execution remains a CI check.

Shared sparse runtime caches

Services and multiple local processes can use blosc2.open(..., shared_cache=True) to keep simultaneous handles to the same private runtime cache:

with blosc2.open(
    "https://host/data.b2z",
    cache_dir="shared-runtime",
    shared_cache=True,
    max_cache_bytes=64 << 20,
) as store:
    with store["experiment/temperature"] as array:
        values = array[:100]
    hit, values = store.read_cached("experiment/temperature", slice(0, 100))

This mode stores leaf payload in sparse RemoteArray caches. Each operation acquires a store-wide OS lock, reloads discovery and leaf accounting, and applies one aggregate payload allowance. Handles may coexist across processes, while operations within a store serialize. All users of that directory must use the shared mode. A process-local memory cache or the ordinary exclusive cache_dir constructor must not write to it.

The default aggregate compressed-payload budget is 256 MiB. Explicitly pass max_cache_bytes=None to disable eviction. This is a post-operation payload bound, not a bound on metadata, total disk usage, or peak RAM. RemoteStore.with_sparse_cache() remains available for advanced attachment with manifests, seed carriers, and authorized filesystems; it uses the same default budget. Use a separate directory from ordinary exclusive caches.

Manifests and generation pointers are published atomically. A process that dies during an operation causes the next owner to discard the disposable payload generation; remote sources are not contacted by offline trimming or manifest recovery. Ordinary exceptions, such as a missing key, leave existing handles usable when operation cleanup succeeds. Refresh publishes a new generation and makes child handles in other processes stale. Reopen a store handle after another process refreshes it.

read_cached returns (False, None) on a miss without fetching missing payload. trim_sparse_cache trims leaves offline with a bounded chunk count. Its first implementation evicts in leaf order; the live aggregate coordinator handles eviction during ordinary reads. Allocated storage and old generations are separate from the compressed-payload allowance and belong to the service’s storage accounting and lifecycle management.

The shared constructor accepts an authorized _filesystem and validation callbacks for server use; these runtime objects are never persisted. A portable carrier archive can seed a new cache once, with source stamp and geometry checks. save exports ordinary portable warm/cold archives. Private sparse directories are not portable store artifacts. This protocol targets processes sharing a local filesystem, not distributed or network-filesystem ownership.

See Working with Remote Data for navigation, shared caching, traffic, credentials, and portable reference examples.

class blosc2.RemoteStore(urlpath, *, dataset=None, path=None, storage_options=None, cache_policy=<policy default>, max_cache_bytes=<policy default>, cache_dir=None, hdf5_index=None, _allow_array_root=False, allow_table_root=False, _filesystem=None, _manifest=None, _source_validator=None, _batch_validator=None, _filesystem_resolver=None, _manifest_validator=None, _max_nodes=None, _source_format=None, _hdf5_index=None, _hdf5_blob=None, _traffic=None, _transport=None, nested_storage_options=None, _b2z_blob=None, _allow_local_source=False, _parquet_conversion=None)[source]

Read-only remote B2Z, Zarr, HDF5, or Caterva2 API hierarchy.

Also used for local hierarchies opened with blosc2.open(..., cache_dir=...).

Discovery and returned array handles share source resources and traffic. MEMORY shares one bounded cache across all leaves; NONE retains no payload. DISK retains payload and discovery under an exclusively owned cache directory. Sources must be immutable until an explicit root refresh.

keys() lists immediate children; get_info() inspects metadata without creating an array cache. Paths are relative to this group. Closing a handle leaves its previously returned arrays and group handles usable.

hdf5_index accepts a native index dictionary or a local/remote JSON path for HDF5 sources. Supplying one skips HDF5 discovery.

path selects a group within the source. dataset remains a supported alias; when both are supplied they must agree after stripping outer slashes. None leaves selection unspecified; an empty string or slash selects the root. Selector keywords cannot be combined with a selector embedded in the URL. allow_table_root=True permits a table selection to remain a store handle for server integrations that manage one store operation across table reads.

Attributes:
attrs

Read-only group attributes, or None when discovery cannot decode them.

cache_bytes

Retained payload bytes; NONE retains no payload between reads.

cache_policy

Configured payload-retention policy.

info

Summary and hierarchy, discovering groups without opening leaf readers.

info_items

The fields shown by info.

is_cache_mutable

Whether the currently opened cache is writable.

max_cache_bytes

Configured compressed-payload allowance.

metadata_bytes

Encoded discovery manifest size, separate from retained payload.

mutable

The export default mutability for future exports.

source

Credential-free source descriptor, including this group’s full path.

traffic

Shared source counter, including discovery; aliases expose the same counter.

Methods

close()

Release this handle; the last dependent handle closes shared resources.

get_info([path])

Return node kind, known attributes and unsupported-node diagnostics.

keys()

Return sorted immediate child names, using only discovery metadata.

kind([path])

Return the discovered node kind.

materialize(destination, *[, overwrite])

Write this hierarchy and reachable store references as one local TreeStore.

read_cached(path[, item, nchunk])

Return (hit, result) atomically without fetching missing payload.

read_cached_table(operation)

Run a table operation from retained payload, returning (hit, value).

refresh()

Rebuild root discovery atomically; existing child handles become stale.

save(destination, *[, include_cache, ...])

Export the current store or subtree to a portable .b2z reference archive.

trim_sparse_cache(runtime_cache_path, ...[, ...])

Trim shared leaf payload without opening or contacting the source.

with_sparse_cache(urlpath, runtime_cache_path, *)

Attach an immutable remote hierarchy to a cache shared across processes.

property attrs

Read-only group attributes, or None when discovery cannot decode them.

property cache_bytes

Retained payload bytes; NONE retains no payload between reads.

property cache_policy

Configured payload-retention policy.

close()[source]

Release this handle; the last dependent handle closes shared resources.

get_info(path='')[source]

Return node kind, known attributes and unsupported-node diagnostics.

property info: InfoReporter

Summary and hierarchy, discovering groups without opening leaf readers.

Zarr discovery may issue metadata or listing requests. Linked remote stores are listed without following their references.

property info_items: list[tuple[str, object]]

The fields shown by info.

property is_cache_mutable: bool

Whether the currently opened cache is writable.

keys()[source]

Return sorted immediate child names, using only discovery metadata.

kind(path='')[source]

Return the discovered node kind.

materialize(destination, *, overwrite=False)[source]

Write this hierarchy and reachable store references as one local TreeStore.

property max_cache_bytes

Configured compressed-payload allowance.

property metadata_bytes

Encoded discovery manifest size, separate from retained payload.

property mutable: bool

The export default mutability for future exports.

read_cached(path, item=(), *, nchunk=None)[source]

Return (hit, result) atomically without fetching missing payload.

read_cached_table(operation)[source]

Run a table operation from retained payload, returning (hit, value).

refresh()[source]

Rebuild root discovery atomically; existing child handles become stale.

save(destination: str | PathLike, *, include_cache: bool = True, mutable: bool | None = None, overwrite: bool = False) → str[source]

Export the current store or subtree to a portable .b2z reference archive.

property source

Credential-free source descriptor, including this group’s full path.

property traffic

Shared source counter, including discovery; aliases expose the same counter.

Source reads are counted, not TCP connections or all HTTP requests. HEAD requests and failed Zarr probes are outside this counter.

static trim_sparse_cache(runtime_cache_path, source, target_bytes, *, max_chunks=64)[source]

Trim shared leaf payload without opening or contacting the source.

classmethod with_sparse_cache(urlpath, runtime_cache_path, *, dataset=None, path=None, manifest=None, max_cache_bytes=<policy default>, carrier=None, storage_options=None, _filesystem=None, _source_validator=None, _batch_validator=None, _filesystem_resolver=None, _manifest_validator=None, _max_nodes=None, _traffic=None, _transport=None, _source_format=None, _hdf5_index=None, _parquet_conversion=None)[source]

Attach an immutable remote hierarchy to a cache shared across processes.

All users of this private cache must use this constructor. Operations serialize per store, reload discovery, and enforce one aggregate payload allowance. The caller authorizes the supplied filesystem and manifest; no credentials or filesystem objects are persisted. Portable artifacts are exported with save rather than opened as mutable runtime storage. path and dataset select the group as in the ordinary constructor. The aggregate compressed-payload budget defaults to 256 MiB; pass max_cache_bytes=None for unlimited retention. For ordinary shared caching, prefer blosc2.open(url, cache_dir=..., shared_cache=True).

class blosc2.RemoteNode(path: str, kind: str, attrs: RemoteMetadataMapping | None, diagnostic: str | None = None)[source]

Discovery metadata; unknown attributes are None and require opening the array.

Attributes:
diagnostic