Working with Remote Data¶
Use blosc2.open() to access remote arrays, tables, and hierarchies through
one API. Data is read on demand and cached locally; remote sources are read-only.
Selected source |
Result |
|---|---|
Standalone |
|
Parquet file, B2Z CTable, or PyTables table |
|
B2Z, Zarr, or HDF5 group |
All three inherit RemoteObject and expose source metadata, cache controls, traffic counters, reference saving, and context-manager support. See Working with Remote Arrays for array computations and Working with Remote Tables for table queries and format-specific behavior.
import blosc2
with blosc2.open(
"https://f001.backblazeb2.com/file/blosc2/hierarchy.h5",
cache_dir="cache-dir",
) as store:
print(store.info)
with store["d0/a0"] as array:
print(array.shape, array.dtype)
values = array[:10]
path="d0/a0" selects a node directly. dataset= is an equivalent alias; URL
selectors such as hierarchy.h5::d0/a0 also work. Use one selector spelling per
call. See Working with Remote Arrays for local-source caching and selector details.
Explore a hierarchy¶
store.keys()or iteration lists immediate children.store["group/array"]andstore["group"]["array"]select the same leaf.store.attrsexposes group attributes.store.get_info("group")returns aRemoteNodewithpath,kind,attrs, and an optional diagnostic, without opening a leaf reader.print(store.info)shows the source, cache policy, retained payload size, entry count, and hierarchy.DictStoreandTreeStoreprovide.infotoo.
Node kinds are group, ndarray, ctable, remote_store, or unsupported.
Unsupported nodes remain visible so you can inspect their diagnostics.
Hierarchy summaries read discovery metadata, but do not open leaf readers or
follow mounted remote stores. Zarr discovery can require listing requests.
HTTP Zarr hierarchies need directory listings or consolidated metadata. Before publishing a local Zarr hierarchy, run:
import zarr
zarr.consolidate_metadata("hierarchy.zarr")
Upload the updated zarr.json for Zarr v3 or .zmetadata for v2 alongside the
store. Without a usable listing, .info marks the hierarchy as incomplete;
selecting a known array path can still work.
Choose a cache policy¶
Setting |
Retained data |
|---|---|
Default: |
Compressed payload in RAM |
|
Persistent local disk cache ( |
|
No payload retained after the operation |
For one array, cache_path="array-cache.b2nd" selects an exact carrier filename.
Tables and stores use cache_dir.
url = "https://f001.backblazeb2.com/file/blosc2/readings-large.parquet"
with blosc2.open(url, cache_dir="cache-dir") as table:
print(table["temperature"][:10])
# Later opens reuse retained metadata and data.
with blosc2.open(url, cache_dir="cache-dir") as table:
print(table["temperature"][:10])
The default retained compressed-payload budget is 256 MiB, per standalone
array or shared across a table/store and its leaves. Set max_cache_bytes to
change it, or pass None for unlimited retention. Eviction happens after
operations. This budget does not bound peak RAM, metadata, result arrays,
or total disk use.
Reads follow the source layout: a small selection can require a whole compressed chunk, batch, or Parquet row group. Small B2Z and HDF5 containers may be fetched in full to reduce request overhead. See the format guides for details.
Storage options and credentials¶
HTTP/HTTPS and cloud URLs use fsspec. HTTP servers must support byte-range reads.
Install the appropriate driver for cloud protocols, such as s3fs for S3.
Pass filesystem settings through storage_options:
with blosc2.open(
"s3://bucket/readings.b2z",
storage_options={"profile": "blosc2", "endpoint_url": "https://s3.example.org"},
) as table:
print(table[:5])
For HTTP authentication, use
storage_options={"headers": {"Authorization": "Bearer <token>"}}.
Credentials and live filesystems are not serialized into saved references;
supply equivalent authentication when reopening. Authenticated users must use
separate cache directories. Caterva2 authentication uses blosc2.c2context().
Measure reads¶
traffic reports cumulative source-read requests and bytes. Reset it around an
operation to distinguish new reads from cache hits:
with blosc2.open(url, cache_dir="cache-dir") as table:
table.traffic.reset()
values = table["temperature"][:10]
print(table.traffic)
table.traffic.reset()
values = table["temperature"][:10]
print(table.traffic) # No source reads if the required data is still cached.
Stores and tables report their shared owner’s traffic and cache size, which can include sibling leaves. Standalone arrays report their own retained payload. These counters describe source reads, not every HTTP request made by a backend.
Save a reference or materialize data¶
save() writes a reference to the source plus data already retained in the
selected cache. It does not fetch missing data. materialize() reads the data
needed for an independent local object.
with blosc2.open(url, cache_dir="cache-dir") as table:
table.save("table-reference.b2z")
table.save("cold-reference.b2z", include_cache=False)
local = table.materialize(urlpath="local-table.b2z")
local.close()
with blosc2.open("table-reference.b2z") as restored:
print(restored[:5]) # Cached data is reused; missing data comes from the source.
Object |
Reference |
Independent data |
|---|---|---|
Array |
|
|
Table |
|
|
Store |
|
|
Table copy(), to_b2z(), and to_b2d() also produce independent local data.
A group view saves only its selected subtree. Existing reference destinations
require overwrite=True.
Reference mutability¶
mutable controls the cache of a saved reference; it never permits remote writes.
Export setting |
Behavior when reopened |
|---|---|
|
Cached data is read in place; misses are fetched transiently. Cache mutation and refresh are disallowed. |
|
A writable runtime cache can retain new reads. The reference archive stays unchanged. |
Pass mutable=True to save() or set the object’s .mutable property before
saving. .is_cache_mutable reports whether the current runtime cache is writable.
For standalone array carriers, mode="a" permits extending the on-disk cache;
mode="r" leaves it unchanged.
References still need their source for uncached data and the appropriate runtime
backends and credentials. Arbitrary custom ProxyNDSource instances cannot be
reconstructed automatically; recreate the source and attach its existing cache
with blosc2.Proxy(source, urlpath="cache.b2nd", mode="a").
Mount remote stores¶
A local TreeStore can hold references to remote hierarchies:
with blosc2.RemoteStore("s3://weather/europe.zarr", path="spain") as weather:
with blosc2.TreeStore("catalog.b2z", mode="w") as catalog:
catalog["/external/weather"] = weather
with blosc2.RemoteStore("catalog.b2z") as catalog:
with catalog["external/weather/temperature"] as array:
values = array[:100]
Targets open on demand. The outer store shares its cache policy, budget, and
traffic counter across mounts. For authenticated targets, pass
nested_storage_options as a URL-keyed mapping or a callable receiving the
credential-free descriptor. A local tree can use
tree.open_remote(path, storage_options=...) for one mount.
save() preserves mounts as references without opening unvisited targets.
RemoteStore.materialize() and TreeStore.materialize() expand reachable mounts
into one local tree, copying attributes and rebuilding table indexes. Cycles and
reference chains deeper than 64 raise ValueError; failed materialization leaves
an existing destination unchanged.
Refresh and close¶
Sources are assumed immutable. Warm caches can reuse discovery metadata without
checking whether the remote file changed. After replacing a source, call
refresh() on its root store or standalone array/table and obtain new child
handles. A failed discovery leaves the existing generation intact; a successful
refresh makes previously obtained child handles stale. Immutable reference
snapshots reject refresh.
Use context managers or close() to release resources. A store’s returned groups,
arrays, and tables can outlive its handle; connections and cache locks remain
owned until the last dependent handle closes. Table columns and views borrow
their root table and require it to remain open.
See also¶
Working with Remote Arrays — array formats, slicing, expressions, and prefetching.
Working with Remote Tables — table formats, queries, indexes, and materialization.
RemoteObject, RemoteStore, RemoteArray, and RemoteCTable — API reference.