Save and load¶
|
Save an array to a file. |
|
Open a persistent SChunk, NDArray, a remote C2Array, RemoteArray, Proxy, a DictStore, EmbedStore, or TreeStore, |
|
Load a persistent Blosc2 object into memory. |
|
Save a serialized NumPy array to a specified file path. |
|
Load a serialized NumPy array from a file. |
|
Save a serialized PyTorch or TensorFlow tensor or NumPy array to a specified file path. |
|
Load a serialized PyTorch or TensorFlow tensor or NumPy array from a file. |
|
Create a EmbedStore, NDArray, SChunk, BatchArray or ObjectArray instance from a contiguous frame buffer. |
- blosc2.save(array: NDArray, urlpath: str, contiguous=True, **kwargs: Any) None[source]¶
Save an array to a file.
- Parameters:
Examples
>>> import blosc2 >>> import numpy as np >>> # Create an array >>> array = blosc2.arange(0, 100, dtype=np.int64, shape=(10, 10)) >>> # Save the array to a file >>> blosc2.save(array, "array.b2", mode="w")
- blosc2.open(urlpath: str | Path | URLPath, mode: str = 'r', offset: int = 0, dataset: str | None = None, hdf5_index: dict | str | PathLike | None = None, *, path: str | None = None, shared_cache: bool = False, **kwargs: dict) SChunk | NDArray | BatchArray | ObjectArray | C2Array | RemoteArray | CTable | RemoteCTable | RemoteStore | LazyArray | Proxy | DictStore | TreeStore | EmbedStore[source]¶
Open a persistent SChunk, NDArray, a remote C2Array, RemoteArray, Proxy, a DictStore, EmbedStore, or TreeStore,
CTable,RemoteCTable, orRemoteStore.See the Notes section for more info on opening Proxy objects.
- Parameters:
urlpath¶ (str | pathlib.Path | URLPath class) – The path where the SChunk (or NDArray) is stored. URLPath class is exclusively a Caterva2 dataset reference, including when its
urlbaseis omitted and inherited fromc2context(); a server names its datasets by root and path rather than by URL. Any URL with a scheme (s3://,gs://,https://,zip://,memory://…) is opened through fsspec; see the Notes section for the limits.mode¶ (str, optional) –
Persistence mode: ‘r’ means read only (must exist); ‘a’ means read/write (create if it doesn’t exist); ‘w’ means create (overwrite if it exists). Defaults to ‘r’ ( read-only).
Open modes also define the allowed persistence side effects:
'r'never writes to the persistent object. It writes a local cache only whencache_dirorcache_pathexplicitly requests one; query acceleration and other implicit execution caches remain process-local only.'a'and'w'may persist explicit user-visible changes such as data, metadata, and index maintenance, but execution caches and query memoization still remain process-local only.
offset¶ (int, optional) – An offset in the file where super-chunk or array data is located (e.g. in a file containing several such objects). A nonzero offset in a local file opens the embedded Blosc2 frame directly, bypassing filename-based container format detection.
shared_cache¶ (bool, optional) – Share an on-demand disk cache between processes. Defaults to False. Requires
cache_dirand a remote source: a standalone.b2ndURL, a B2Z, HDF5, or Zarr container, or a Caterva2 URLPath class. Requires lazy access,CachePolicy.DISK, andassume_immutable=True. Enables lazy access for every remote source whenlazyis omitted or None; explicitlazy=Falseis rejected. Returns the selected table, group, or array using sparse cache storage and operation-scoped locks. Operations on the same cache serialize. All processes using this cache must enable sharing; use a separate directory from ordinary exclusive caches. The aggregate retained compressed-payload budget defaults to 256 MiB per standalone array or table/store owner;max_cache_bytes=Nonedisables eviction. This does not bound total disk usage or peak RAM. Authenticated Caterva2 users must use separate cache directories; tokens are not persisted.kwargs¶ (dict, optional) –
- lazy: bool or None, optional
None(the default) automatically selects the access mode.Truereturns a lazy remote object;Falserequests eager access. Known remote .b2nd arrays and .b2z archives default to lazy access, including whencache_dir=is supplied. Passlazy=Falseto download the whole container undercache_dirinstead.mmap_mode=or a nonzerooffset=forces the eager path. Dataset paths currently requirelazy=Trueand reject explicitFalse. For an fsspec URL or a Caterva2 URLPath class, return a RemoteArray over the remote array dataset and read the byte ranges a slice touches. B2Z and HDF5 table and group nodes returnRemoteCTableandRemoteStoreinstead. Selected local PyTables/HDF5 tables also returnRemoteCTable; local ordinary datasets remainRemoteArray. A slice landing in a small part of a large chunk costs only the blocks it touches when ranges are available; chunks small enough to be one cheap request are still fetched whole. What arrives is kept in memory (defaulting toCachePolicy.MEMORY), undercache_dir, or at the exactcache_path(asCachePolicy.DISK).- max_concurrency: int, optional
For lazy remote arrays and RemoteCTable: how many fetches to run at once, in a thread pool. A slice against an object store is almost entirely round-trip latency, so overlapping the requests is what makes a wide slice bearable. Defaults to 8; pass 1 for a protocol with no latency to hide, where the pool costs about 10 microseconds per chunk and saves nothing. Remote tables overlap independent requests across selected columns; their table-specific temporary buffer settings are available through
RemoteCTable, not through this general opener.- cache_dir: str | pathlib.Path, optional
Parent directory for a persistent RemoteArray,
RemoteCTable, orRemoteStorecache. For remote inputs it stores fetched chunks and metadata; for local Blosc2, HDF5, and Zarr inputs it stores accessed chunks and converted data. Supplying it selects on-demand, read-only access for local inputs. Local caches assume the source is immutable: callrefresh()on the cached handle after replacing it at the same path. No cache is created unless you name one.- cache_path: str | pathlib.Path, optional
Exact file for a persistent array cache (
CachePolicy.DISK), for local or remote standalone arrays and array leaves. Tables and groups requirecache_dir. Mutually exclusive withcache_dir.- cache_storage: str | pathlib.Path, optional
Deprecated alias for
cache_dir. Mutually exclusive withcache_dirandcache_path.- cache_policy: CachePolicy, optional
With
lazy=Trueon a remote source, return a RemoteArray using the requested retention policy (NONE,MEMORY, orDISK). When omitted, passingcache_dirorcache_pathdefaults toCachePolicy.DISK, while omitting them defaults toCachePolicy.MEMORY.- max_cache_bytes: int or None, optional
With
lazy=True, bound retained compressed cache payload for a RemoteArray after each operation. Defaults to 256 MiB for bothDISKandMEMORY. PassingNonewithDISKdisables cache eviction (unbounded cache). This does not bound the current operation’s working set or result.- mmap_mode: str, optional
If set, the file will be memory-mapped instead of using the default I/O functions and the mode argument will be ignored. For more info, see
blosc2.Storage. Please note that the w+ mode, which can be used to create new files, is not supported here since only existing files can be opened. You can useSChunk.__init__to create new files.- initial_mapping_size: int, optional
The initial size of the memory mapping. For more info, see
blosc2.Storage.- locking: bool, optional
Serialize accesses against other handles and other processes via a sidecar lock file. Enable it when several processes operate on the same container. The locking is advisory (every handle on the container must enable it) and cannot be combined with mmap_mode. The
BLOSC_LOCKINGenvironment variable enables it globally. For more info, seeblosc2.Storage.- cparams: dict
A dictionary with the compression parameters, which are the same that can be used in the
compress2()function. Typesize and blocksize cannot be changed.- dparams: dict
A dictionary with the decompression parameters, which are the same that can be used in the
decompress2()function.- storage_options: dict, optional
Parameters passed to the underlying
fsspecfilesystem when opening an fsspec URL (for instance credentials, endpoint URL, token, client_kwargs, etc.).- path: str, optional
Node path within HDF5, Zarr, or B2Z containers (e.g.
path="d0/d1/a2"). B2Z and HDF5 also support table and group paths in immutable containers. Requireslazy=True. None leaves selection unspecified;""and"/"select the root. Do not combine with a selector embedded in the URL.- dataset: str, optional
Supported alias of
path. Both may be supplied if they agree after stripping leading/trailing slashes; conflicting values raise ValueError.- hdf5_index: dict | str | PathLike, optional
Pre-computed native HDF5 index, or a local path or remote fsspec URL to its JSON encoding. It must match the source HDF5 URL and dataset scope. Arrays, PyTables tables, and hierarchy stores are supported.
- source_format: {None, “blosc2”, “zarr”, “hdf5”, “b2z”, “parquet”}, optional
Format of a lazy remote source. A
.zarrURL path component selects Zarr automatically; a.h5or.hdf5path selects HDF5 automatically; a.b2zpath selects B2Z automatically. An explicit value supports suffix-free paths. Zarr, HDF5, B2Z and Parquet sources automatically enablelazy=True.- parquet_options: dict, optional
PyArrow
ParquetFilereader options for a Parquet source. Conversion options such ascolumnsandmax_rowsare passed separately.- assume_immutable: bool, optional
With
lazy=True, skip remote identity checks before reads. Defaults toTrue. Local disk caches always assume immutable sources; set to a new path or explicitly rebuild the cache after changing a local source.
- Returns:
out – Proxy, DictStore, EmbedStore, TreeStore,
CTable,RemoteCTable, orRemoteStoreThe object found in the path.- Return type:
Notes
Returned objects can be used as context managers for API consistency. For objects with an explicit
close()implementation, exiting the context will close/flush them. Standalone HDF5RemoteArrayhandles close their HDF5 source, so the handle rejects further reads. Other logical handles such as regularSChunk,NDArray,C2Array, standalone non-HDF5RemoteArray,Proxy, andLazyArraycurrently treat context exit as a no-op. Store-derivedRemoteArrayhandles release their shared source ownership when closed; other handles from that store remain usable.If
urlpathis a URLPath class instance,modemust be ‘r’ andoffsetmust be 0. By default it returns a C2Array. Withlazy=Trueorshared_cache=True, it returns a RemoteArray (defaulting toCachePolicy.DISKwhencache_dirorcache_pathis provided, andCachePolicy.MEMORYotherwise). Authenticated users sharing a machine must use separate caches.fsspec URLs need the
fsspecextra (pip install "blosc2[fsspec]") and the driver for the protocol (s3fs,gcsfs…), which fsspec asks for by name when it is missing. Driver and protocol parameters (credentials, endpoint URL, region, etc.) can be passed directly viastorage_options.mode != 'r'always raises, as object stores have no rename and no locks. A plain URL read rebuilds the object from a cframe held in memory, so it covers.b2nd,.b2fand.b2eonly. Remote.b2zarchives default to lazy discovery: table nodes returnRemoteCTable, groups returnRemoteStore, and external array leaves return RemoteArray. Select a nested node withpath=or::path. Tables and groups usecache_dirfor DISK caching and default to MEMORY otherwise; array leaves additionally supportcache_path. Uselazy=False, cache_dir=...to download a complete archive and open its root locally. Local table archives returnCTable.Persistent data handling follows a no-hidden-writes rule except for an explicitly self-caching RemoteArray:
mode='r'is observational only and never mutates the opened object.mode='a'permits aDISKRemoteArray to retain remote chunks in its own carrier. Other execution caches are not serialized implicitly.mode='w'persists explicit mutations requested by the caller.
If the original object saved in
urlpathis a Proxy or a RemoteArray, this function reconstructs sources backed by a persistent local SChunk or NDArray, an fsspec URL, or a remote C2Array. Custom proxy sources must be recreated explicitly because their Python class and runtime state are not stored in the cache.When opening a LazyExpr keep in mind the note above regarding operands.
Examples
>>> import blosc2 >>> import numpy as np >>> import os >>> import tempfile >>> tmpdirname = tempfile.mkdtemp() >>> urlpath = os.path.join(tmpdirname, 'b2frame') >>> storage = blosc2.Storage(contiguous=True, urlpath=urlpath, mode="w") >>> nelem = 20 * 1000 >>> nchunks = 5 >>> chunksize = nelem * 4 // nchunks >>> data = np.arange(nelem, dtype="int32") >>> # Create SChunk and append data >>> schunk = blosc2.SChunk(chunksize=chunksize, data=data.tobytes(), storage=storage) >>> # Open SChunk >>> sc_open = blosc2.open(urlpath=urlpath, mode="r") >>> for i in range(nchunks): ... dest = np.empty(nelem // nchunks, dtype=data.dtype) ... schunk.decompress_chunk(i, dest) ... dest1 = np.empty(nelem // nchunks, dtype=data.dtype) ... sc_open.decompress_chunk(i, dest1) ... np.array_equal(dest, dest1) True True True True True
To open the same schunk memory-mapped, we simply need to pass the mmap_mode parameter:
>>> sc_open_mmap = blosc2.open(urlpath=urlpath, mmap_mode="r") >>> sc_open.nchunks == sc_open_mmap.nchunks True >>> all(sc_open.decompress_chunk(i, dest1) == sc_open_mmap.decompress_chunk(i, dest1) for i in range(nchunks)) True
- blosc2.load(urlpath: str | Path, offset: int = 0, **kwargs: dict)[source]¶
Load a persistent Blosc2 object into memory.
This is the in-memory counterpart to
open(). It opens urlpath in read-only mode and returns a standalone object that is not backed by the original file. ForCTable, this dispatches toCTable.load(); for array-like containers it returns an in-memory copy.- Parameters:
urlpath¶ (str | pathlib.Path) – Path to the persistent Blosc2 object.
offset¶ (int, optional) – Offset in the file where the object is located. This is mainly useful for SChunk/NDArray objects embedded in a larger file.
kwargs¶ (dict, optional) – Additional read-time keyword arguments passed to
open(), such asdparams.
- Returns:
A standalone in-memory Blosc2 object.
- Return type:
out
- Raises:
TypeError – If the opened object cannot be loaded as a standalone in-memory object.
Examples
>>> import blosc2 >>> import numpy as np >>> arr = blosc2.asarray(np.arange(10), urlpath="example.b2nd", mode="w") >>> loaded = blosc2.load("example.b2nd") >>> loaded.urlpath is None True >>> np.array_equal(loaded[:], arr[:]) True >>> blosc2.remove_urlpath("example.b2nd")
- blosc2.save_array(arr: ndarray, urlpath: str, chunksize: int | None = None, **kwargs: dict) int[source]¶
Save a serialized NumPy array to a specified file path.
- Parameters:
arr¶ (np.ndarray) – The NumPy array to be saved.
urlpath¶ (str) – The path for the file where the array will be saved. An fsspec URL (
s3://,gs://,memory://…) writes the whole container in one shot; it needs thefsspecextra and the protocol driver installed.chunksize¶ (int) – The size (in bytes) for the chunks during compression. If not provided, it is computed automatically.
kwargs¶ (dict, optional) – These are the same as the kwargs in
SChunk.__init__.
- Returns:
out – The number of bytes of the saved array.
- Return type:
int
Examples
>>> import numpy as np >>> a = np.arange(1e6) >>> serial_size = blosc2.save_array(a, "test.bl2", mode="w") >>> serial_size < a.size * a.itemsize True
See also
- blosc2.load_array(urlpath: str, dparams: dict | None = None, *, storage_options: dict | None = None) ndarray[source]¶
Load a serialized NumPy array from a file.
- Parameters:
urlpath¶ (str) – The path to the file containing the serialized array.
dparams¶ (dict, optional) – A dictionary with the decompression parameters, which can be used in the
decompress2()function.storage_options¶ (dict, optional) – Parameters passed to the underlying
fsspecfilesystem when opening an fsspec URL.
- Returns:
out – The deserialized NumPy array.
- Return type:
np.ndarray
- Raises:
TypeError – If
urlpathis not in cframe formatRunTimeError – If any other error is detected.
Examples
>>> import numpy as np >>> a = np.arange(1e6) >>> serial_size = blosc2.save_array(a, "test.bl2", mode="w") >>> serial_size < a.size * a.itemsize True >>> a2 = blosc2.load_array("test.bl2") >>> np.array_equal(a, a2) True
See also
- blosc2.save_tensor(tensor: tensorflow.Tensor | torch.Tensor | np.ndarray, urlpath: str, chunksize: int | None = None, **kwargs: dict) int[source]¶
Save a serialized PyTorch or TensorFlow tensor or NumPy array to a specified file path.
- Parameters:
tensor¶ (tensorflow.Tensor, torch.Tensor, or np.ndarray) – The tensor or array to be saved.
urlpath¶ (str) – The file path where the tensor or array will be saved. An fsspec URL (
s3://,gs://,memory://…) writes the whole container in one shot; it needs thefsspecextra and the protocol driver installed.chunksize¶ (int) – The size (in bytes) for the chunks during compression. If not provided, it is computed automatically.
kwargs¶ (dict, optional) – These are the same as the kwargs in
SChunk.__init__.
- Returns:
out – The number of bytes of the saved tensor or array.
- Return type:
int
Examples
>>> import numpy as np >>> th = np.arange(1e6, dtype=np.float32) >>> serial_size = blosc2.save_tensor(th, "test.bl2", mode="w") >>> if not os.getenv("BTUNE_TRADEOFF"): ... assert serial_size < th.size * th.itemsize ...
See also
- blosc2.load_tensor(urlpath: str, dparams: dict | None = None, *, storage_options: dict | None = None) tensorflow.Tensor | torch.Tensor | np.ndarray[source]¶
Load a serialized PyTorch or TensorFlow tensor or NumPy array from a file.
- Parameters:
urlpath¶ (str) – The path to the file where the tensor or array is stored.
dparams¶ (dict, optional) – A dictionary with the decompression parameters, which are the same as those used in the
decompress2()function.storage_options¶ (dict, optional) – Parameters passed to the underlying
fsspecfilesystem when opening an fsspec URL.
- Returns:
out – The unpacked PyTorch or TensorFlow tensor or NumPy array.
- Return type:
tensor or ndarray
- Raises:
TypeError – If
urlpathis not in cframe formatRunTimeError – If some other problem is detected.
Examples
>>> import numpy as np >>> th = np.arange(1e6, dtype=np.float32) >>> size = blosc2.save_tensor(th, "test.bl2", mode="w") >>> if not os.getenv("BTUNE_TRADEOFF"): ... assert size < th.size * th.itemsize ... >>> th2 = blosc2.load_tensor("test.bl2") >>> np.array_equal(th, th2) True
See also
- blosc2.from_cframe(cframe: bytes | str, copy: bool = True) EmbedStore | NDArray | SChunk | ListArray | BatchArray | ObjectArray | C2Array | RemoteArray[source]¶
Create a EmbedStore, NDArray, SChunk, BatchArray or ObjectArray instance from a contiguous frame buffer.
- Parameters:
cframe¶ (bytes or str) – The bytes object containing the in-memory cframe.
copy¶ (bool) – Whether to internally make a copy. Default is True. With False the returned object points into cframe’s buffer and keeps a reference to it (the buffer lives for as long as the object does), saving time/memory at the cost of the buffer staying pinned.
- Returns:
- out – BatchArray, ObjectArray, or
A new instance of the appropriate type containing the data passed.
- Return type:
See also
from_cframe(),from_cframe(),from_cframe()