RemoteStore
RemoteStore discovers a read-only B2Z, Zarr, HDF5 or Parquet source and returns
RemoteArray and RemoteCTable leaves. Groups and leaves share one
source session: a B2Z archive, a native HDF5 index, or a Zarr store. Zarr listing
remains lazy.
The default CachePolicy.MEMORY shares a 256 MiB allowance across all leaves.
Set max_cache_bytes to a positive integer to change it. CachePolicy.NONE
retains no payload and rejects a limit. Passing cache_dir selects DISK when
the policy is omitted; an explicit policy must agree with the cache location.
DISK accepts max_cache_bytes=None for unbounded retention.
Sources must be immutable. Generic blosc2.open(..., lazy=True, dataset=...)
continues to open a single array.
A Parquet file has one CTable at its root. Pass allow_table_root=True to
retain a store handle, then use store[""] to access the table. Parquet has no
child selectors. Converted physical-column/row-group caches share the store’s
policy, allowance, traffic counter, persistence, and refresh generation.
For HDF5, hdf5_index= accepts a native index dictionary, local JSON path, or
remote fsspec URL. An explicit index skips hierarchy discovery and must match the
source URL and selected scope. See Working with Remote Arrays for the
sidecar-generation workflow.
with blosc2.RemoteStore(
"https://host/data.h5", cache_policy=blosc2.CachePolicy.NONE
) as store:
print(store.keys()) # immediate children
info = store.get_info("experiment") # metadata only
group = store["experiment"]
array = group["temperature"]
values = array[:100]
result = (array + 273.15).compute()
# Returned handles own their source lifetime independently.
values = array[:100]
group.close()
array.close()
Paths are relative to the selected group. store["a/b"] and
store["a"]["b"] use the same source reader. Each lookup returns an independent
handle. With NONE, repeated reads fetch again; no payload cache is retained.
dataset="a" or a subgroup suffix in the URL selects a group at construction.
An array root must be opened with RemoteArray instead.
keys() and get_info() do not construct leaf readers or payload caches.
Discovery can read archive prefixes, attributes and small HDF5 inline values.
get_info() returns a RemoteNode with a relative path, a kind (group,
ndarray, ctable, remote_store or unsupported), known attributes
and a diagnostic.
Unknown array attributes are None; open the array to retrieve them.
Unsupported nodes stay discoverable and raise NotImplementedError when
selected. Missing paths raise KeyError.
Group attrs mappings are read-only. source returns the credential-free
container descriptor and full group path. traffic is one shared source
counter across all views, including discovery; do not add counts from aliases.
It counts source reads, not connections or every HTTP request: HEAD requests and
failed Zarr probes are excluded. cache_bytes counts retained compressed chunks
and partial-block duplicates across the store, without double-counting aliases.
The least recently used native chunk is evicted across all leaves when necessary.
Oversized reads return their result before eviction; returned arrays and temporary
buffers are outside the allowance. Closing a leaf preserves its warm cache while
other session handles remain open. With NONE, cache_bytes is zero and
max_cache_bytes is None.
Closing a handle, or exiting its context, releases its ownership. Existing child
handles remain usable until closed or garbage-collected. The last handle closes
the owned archive/store wrappers and private HTTP/S3 transport sessions. Operations on an explicitly closed handle raise RuntimeError.
Standalone RemoteArray exports remain self-contained references, including
the native HDF5 index when applicable.
Nested stores
Assigning a RemoteStore to a TreeStore persists a lazy, credential-free
reference at the exact path supplied by the caller:
with blosc2.RemoteStore("s3://weather/europe.zarr", dataset="spain") as remote:
with blosc2.TreeStore("catalog.b2z", mode="w") as tree:
tree["/external/weather"] = remote
Remote B2Z discovery reports that object root as remote_store without opening
the linked source. Lookup through the mount supports direct and chained paths.
B2Z, HDF5, Zarr v2 and Zarr v3 targets and subgroup references are supported.
The outer RemoteStore overrides saved cache defaults and owns one policy,
traffic counter and aggregate allowance across mounted sources. Pass
nested_storage_options as a URL-to-options mapping or a callable accepting a
source descriptor when mounted sources need different credentials. A local
TreeStore can open one reference with runtime options through
open_remote().
save() preserves nested references and includes only already-retained warm
payload from mounts that were opened. It never opens an unvisited mount to save
it. materialize(destination) follows all reachable mounts and writes one
independent local .b2z or .b2d TreeStore. Materialization detects cycles,
limits nesting to 64 levels, and publishes the destination atomically.
b2view uses RemoteStore for remote hierarchies with one 64 MiB MEMORY
allowance, and RemoteArray for selected or directly opened leaves. Switching
selection releases the UI handle while retaining the store’s warm chunks.
For persistent shared caching:
with blosc2.RemoteStore("https://host/data.h5", cache_dir="remote-cache") as store:
with store["experiment/temperature"] as array:
values = array[:100]
print(store.cache_bytes, store.metadata_bytes)
store.refresh() # rebuild discovery; old child handles become stale
Each source and selected root has its own directory under cache_dir. One
owner holds an exclusive operating-system lock until its last dependent handle
closes. Conflicting opens raise RuntimeError, including in other processes;
the operating system releases the lock after a process exits or crashes.
Reopening restores all previously created leaf caches and trims them against the
new aggregate allowance before returning. The manifest preserves B2Z directory
and bounded metadata reads, one native HDF5 index, and lazily discovered Zarr
metadata. Metadata reads can contain small inline values or incidental bytes in
bounded prefixes; they are separate from evictable payload. metadata_bytes
is the encoded manifest size, and is zero without a disk manifest. Credentials
and storage options must be supplied again at runtime.
The allowance measures compressed retained payload, not filesystem allocation.
Container overhead and old generations are outside it; obsolete generation
files are removed on the next exclusive reopen. Manifests survive payload
eviction. Store backing files are private implementation details; use array
save() or to_cframe() for standalone exports, optionally with
include_cache=True.
Sources must remain immutable until explicit root refresh(). Refresh builds
replacement discovery before publishing a new generation; failed discovery or
publication leaves existing handles valid. Successful refresh makes existing
child groups and arrays stale, requiring fresh lookups. Corrupt manifests and
B2Z validator mismatches raise an error; use a fresh cache directory if the store
cannot be opened. Fully offline reopening is not promised. The lifetime lock
uses POSIX flock or Windows byte locking; Windows execution remains a CI check.
Shared sparse runtime caches
Services and multiple local processes can use blosc2.open(..., shared_cache=True)
to keep simultaneous handles to the same private runtime cache:
with blosc2.open(
"https://host/data.b2z",
cache_dir="shared-runtime",
shared_cache=True,
max_cache_bytes=64 << 20,
) as store:
with store["experiment/temperature"] as array:
values = array[:100]
hit, values = store.read_cached("experiment/temperature", slice(0, 100))
This mode stores leaf payload in sparse RemoteArray caches. Each operation
acquires a store-wide OS lock, reloads discovery and leaf accounting, and applies
one aggregate payload allowance. Handles may coexist across processes, while
operations within a store serialize. All users of that directory must use the
shared mode. A process-local memory cache or the ordinary exclusive
cache_dir constructor must not write to it.
The default aggregate compressed-payload budget is 256 MiB. Explicitly pass
max_cache_bytes=None to disable eviction. This is a post-operation payload
bound, not a bound on metadata, total disk usage, or peak RAM.
RemoteStore.with_sparse_cache() remains available for advanced attachment
with manifests, seed carriers, and authorized filesystems; it uses the same
default budget. Use a separate directory from ordinary exclusive caches.
Manifests and generation pointers are published atomically. A process that dies
during an operation causes the next owner to discard the disposable payload
generation; remote sources are not contacted by offline trimming or manifest recovery.
Ordinary exceptions, such as a missing key, leave existing handles usable when
operation cleanup succeeds.
Refresh publishes a new generation and makes child handles in other processes
stale. Reopen a store handle after another process refreshes it.
read_cached returns (False, None) on a miss without fetching missing
payload. trim_sparse_cache trims leaves offline with a bounded chunk count.
Its first implementation evicts in leaf order; the live aggregate coordinator
handles eviction during ordinary reads. Allocated storage and old generations
are separate from the compressed-payload allowance and belong to the service’s
storage accounting and lifecycle management.
The shared constructor accepts an authorized _filesystem and validation
callbacks for server use; these runtime objects are never persisted. A portable
carrier archive can seed a new cache once, with source stamp and geometry
checks. save exports ordinary portable warm/cold archives. Private sparse
directories are not portable store artifacts. This protocol targets processes
sharing a local filesystem, not distributed or network-filesystem ownership.
See Working with Remote Data for navigation,
shared caching, traffic, credentials, and portable reference examples.
-
class blosc2.RemoteStore(urlpath, *, dataset=None, path=None, storage_options=None, cache_policy=<policy default>, max_cache_bytes=<policy default>, cache_dir=None, hdf5_index=None, _allow_array_root=False, allow_table_root=False, _filesystem=None, _manifest=None, _source_validator=None, _batch_validator=None, _filesystem_resolver=None, _manifest_validator=None, _max_nodes=None, _source_format=None, _hdf5_index=None, _hdf5_blob=None, _traffic=None, _transport=None, nested_storage_options=None, _b2z_blob=None, _allow_local_source=False, _parquet_conversion=None)[source]
Read-only remote B2Z, Zarr, HDF5, or Caterva2 API hierarchy.
Also used for local hierarchies opened with blosc2.open(..., cache_dir=...).
Discovery and returned array handles share source resources and traffic.
MEMORY shares one bounded cache across all leaves; NONE retains no payload.
DISK retains payload and discovery under an exclusively owned cache directory.
Sources must be immutable until an explicit root refresh.
keys() lists immediate children; get_info() inspects metadata without
creating an array cache. Paths are relative to this group. Closing a handle
leaves its previously returned arrays and group handles usable.
hdf5_index accepts a native index dictionary or a local/remote JSON path
for HDF5 sources. Supplying one skips HDF5 discovery.
path selects a group within the source. dataset remains a supported
alias; when both are supplied they must agree after stripping outer slashes.
None leaves selection unspecified; an empty string or slash selects the root.
Selector keywords cannot be combined with a selector embedded in the URL.
allow_table_root=True permits a table selection to remain a store handle
for server integrations that manage one store operation across table reads.
- Attributes:
attrsRead-only group attributes, or None when discovery cannot decode them.
cache_bytesRetained payload bytes; NONE retains no payload between reads.
cache_policyConfigured payload-retention policy.
infoSummary and hierarchy, discovering groups without opening leaf readers.
info_itemsThe fields shown by info.
is_cache_mutableWhether the currently opened cache is writable.
max_cache_bytesConfigured compressed-payload allowance.
metadata_bytesEncoded discovery manifest size, separate from retained payload.
mutableThe export default mutability for future exports.
sourceCredential-free source descriptor, including this group’s full path.
trafficShared source counter, including discovery; aliases expose the same counter.
Methods
close()
|
Release this handle; the last dependent handle closes shared resources. |
get_info([path])
|
Return node kind, known attributes and unsupported-node diagnostics. |
keys()
|
Return sorted immediate child names, using only discovery metadata. |
kind([path])
|
Return the discovered node kind. |
materialize(destination, *[, overwrite])
|
Write this hierarchy and reachable store references as one local TreeStore. |
read_cached(path[, item, nchunk])
|
Return (hit, result) atomically without fetching missing payload. |
read_cached_table(operation)
|
Run a table operation from retained payload, returning (hit, value). |
refresh()
|
Rebuild root discovery atomically; existing child handles become stale. |
save(destination, *[, include_cache, ...])
|
Export the current store or subtree to a portable .b2z reference archive. |
trim_sparse_cache(runtime_cache_path, ...[, ...])
|
Trim shared leaf payload without opening or contacting the source. |
with_sparse_cache(urlpath, runtime_cache_path, *)
|
Attach an immutable remote hierarchy to a cache shared across processes. |
-
property attrs
Read-only group attributes, or None when discovery cannot decode them.
-
property cache_bytes
Retained payload bytes; NONE retains no payload between reads.
-
property cache_policy
Configured payload-retention policy.
-
close()[source]
Release this handle; the last dependent handle closes shared resources.
-
get_info(path='')[source]
Return node kind, known attributes and unsupported-node diagnostics.
-
property info: InfoReporter
Summary and hierarchy, discovering groups without opening leaf readers.
Zarr discovery may issue metadata or listing requests. Linked remote
stores are listed without following their references.
-
property info_items: list[tuple[str, object]]
The fields shown by info.
-
property is_cache_mutable: bool
Whether the currently opened cache is writable.
-
keys()[source]
Return sorted immediate child names, using only discovery metadata.
-
kind(path='')[source]
Return the discovered node kind.
-
materialize(destination, *, overwrite=False)[source]
Write this hierarchy and reachable store references as one local TreeStore.
-
property max_cache_bytes
Configured compressed-payload allowance.
-
property metadata_bytes
Encoded discovery manifest size, separate from retained payload.
-
property mutable: bool
The export default mutability for future exports.
-
read_cached(path, item=(), *, nchunk=None)[source]
Return (hit, result) atomically without fetching missing payload.
-
read_cached_table(operation)[source]
Run a table operation from retained payload, returning (hit, value).
-
refresh()[source]
Rebuild root discovery atomically; existing child handles become stale.
-
save(destination: str | PathLike, *, include_cache: bool = True, mutable: bool | None = None, overwrite: bool = False) → str[source]
Export the current store or subtree to a portable .b2z reference archive.
-
property source
Credential-free source descriptor, including this group’s full path.
-
property traffic
Shared source counter, including discovery; aliases expose the same counter.
Source reads are counted, not TCP connections or all HTTP requests.
HEAD requests and failed Zarr probes are outside this counter.
-
static trim_sparse_cache(runtime_cache_path, source, target_bytes, *, max_chunks=64)[source]
Trim shared leaf payload without opening or contacting the source.
-
classmethod with_sparse_cache(urlpath, runtime_cache_path, *, dataset=None, path=None, manifest=None, max_cache_bytes=<policy default>, carrier=None, storage_options=None, _filesystem=None, _source_validator=None, _batch_validator=None, _filesystem_resolver=None, _manifest_validator=None, _max_nodes=None, _traffic=None, _transport=None, _source_format=None, _hdf5_index=None, _parquet_conversion=None)[source]
Attach an immutable remote hierarchy to a cache shared across processes.
All users of this private cache must use this constructor. Operations
serialize per store, reload discovery, and enforce one aggregate payload
allowance. The caller authorizes the supplied filesystem and manifest;
no credentials or filesystem objects are persisted. Portable artifacts
are exported with save rather than opened as mutable runtime storage.
path and dataset select the group as in the ordinary constructor.
The aggregate compressed-payload budget defaults to 256 MiB; pass
max_cache_bytes=None for unlimited retention. For ordinary shared
caching, prefer blosc2.open(url, cache_dir=..., shared_cache=True).
-
class blosc2.RemoteNode(path: str, kind: str, attrs: RemoteMetadataMapping | None, diagnostic: str | None = None)[source]
Discovery metadata; unknown attributes are None and require opening the array.
- Attributes:
- diagnostic