RemoteCTable

RemoteCTable is a read-only blosc2.CTable backed by a remote B2Z archive, a local or remote PyTables/HDF5 table, or a Parquet file. Fixed-width, shaped, nullable, UTF-8, batch-backed variable-length, batch-backed list, struct/object, and dictionary columns are fetched on demand. Open local PyTables tables through blosc2.open() with path= or a ::table selector. Standalone tables can be opened directly; tables inside a hierarchy can be selected with dataset= or through blosc2.RemoteStore. PyTables/HDF5 sources may supply hdf5_index= as a native index dictionary, local JSON path, or remote fsspec URL. This skips source discovery and does not modify the HDF5 file; see Working with Remote Arrays.

Batch-backed reads transfer and decode whole compressed batches. Dictionary codes remain selective, while the full vocabulary is loaded on first use. ListArray batches contain 2048 list cells by default. Smaller batches reduce overfetch for sparse reads; larger batches usually improve scans and compression. The row limit is not a byte limit, so one unusually large list can still require a large transfer. See ListArray for batching controls.

Nullable list elements, nested lists, and contains/overlaps predicates have the same semantics as local tables. Without an index a predicate scans the remote list batches. A persisted kind="membership" index on a flat scalar list fetches only the compressed posting batches for requested values. If the result projects only other columns, the list payload remains unopened. Indexes on nested lists and structs are not supported. ListArray storage="vl" also remains unavailable through RemoteCTable.

Scalar queries automatically use persisted SUMMARY, FULL, PARTIAL, OPSI, and BUCKET indexes. Their sidecars are opened lazily and participate in the outer table’s cache budget and traffic accounting. SUMMARY reads compact min/max records before fetching candidate data blocks; positional indexes use their navigation data to fetch selected value and row-position ranges. Queries retain a correct scan fallback when an index layout or expression is unsupported.

Saving and materializing have different meanings:

with blosc2.RemoteCTable(url) as table:
    table["temperature"][:100]  # warm part of one column
    table.save("reference.b2z")
    table.save("cold-reference.b2z", include_cache=False)
    local = table.materialize(urlpath="complete.b2z")

save() writes the source descriptor, table metadata, and by default only payload already retained in the cache. Missing data is still read from the original source after reopening. materialize() returns an independent local table and reads all data needed for it. copy(), to_b2z(), and to_b2d() remain local materialization operations inherited from blosc2.CTable.

Table cache bytes, limits, and traffic are scoped to the shared remote owner and may include sibling leaves. This also applies when a table column is itself a RemoteArray reference to another fsspec or Caterva2 URL: the outer table policy overrides the carrier policy for that handle, and all such columns share one budget. Standalone instances of those arrays keep their original policies. A table selected from a store owns an independent handle, but its columns and views remain borrowed from that table. Refresh a nested table through its root store; a standalone table can call refresh().

Server processes can share one bounded sparse disk cache with blosc2.open(url, cache_dir="shared-cache", shared_cache=True). All processes using that directory must enable sharing. Operations serialize per store, and the aggregate compressed-payload budget defaults to 256 MiB; explicitly pass max_cache_bytes=None for unlimited retention. Use a separate directory from ordinary exclusive caches. The budget does not bound total disk usage or peak RAM.

RemoteCTable.with_sparse_cache(url, runtime_cache_path) remains available for advanced attachment with manifests or seed carriers, with the same default budget and shared-cache implementation.

Parquet cache details

Parquet cache_dir stores each accessed physical field and row group as a native CTable directory under the RemoteStore generation’s parquet-groups directory. The common manifest retains the footer, schema, conversion options, source marker, and row-group boundaries. Warm opens reuse this discovery and payload without contacting the source. Cached sources are assumed immutable until refresh() is called. New portable .b2z archives use the common RemoteStore manifest. Local Parquet sources use the same cache layout and may export .b2z references tied to the local source path.

HTTP Parquet sources need byte-range support. A persistent cache also needs a source size and version marker, such as an ETag, modification time, or Backblaze B2 file ID. HTTP read-ahead is disabled by default to avoid fetching unused bytes; an explicit storage_options={"cache_type": "bytes"} restores fsspec buffering. The traffic counter counts fetched ranges; explicitly enabled transport buffering may make actual HTTP transfer different.

Multi-column selections, including table previews, fetch independent Parquet column chunks concurrently using max_concurrency (8 by default). Decoding and cache writes stay on the calling thread. row_buffer_bytes bounds each wave of compressed responses; one physical field larger than the budget runs alone. A physical field with nested leaves is fetched as one unit. Reads still load whole column chunks for the selected row groups, even for a few rows. Use max_concurrency=1 for serial transport.

See Working with Remote Tables for column access, filtering, buffering, reference saving, and materialization examples.

class blosc2.RemoteCTable(urlpath=None, *, dataset=None, path=None, storage_options=None, cache_policy=<policy default>, max_cache_bytes=<policy default>, cache_dir=None, shared_cache=False, hdf5_index=None, max_concurrency=8, metadata_buffer_bytes=8388608, row_buffer_bytes=67108864, source_format=None, parquet_options=None, columns=None, max_rows=None, string_max_length=None, null_storage=None, auto_null_sentinels=True, separate_nested_cols=True, list_serializer='msgpack', blosc2_batch_size=2048, blosc2_items_per_block=None, batch_size=2048, cparams=None, dparams=None, validate=False, _filesystem=None, _filesystem_resolver=None, _batch_validator=None)[source]

A read-only CTable whose columns are fetched on demand.

Supported columns include fixed-width, UTF-8, batch-backed variable-length, batch-backed list, struct/object and dictionary columns. Batch reads transfer one whole compressed batch; dictionary decoding loads the full vocabulary on first use.

Independent column requests overlap by default. max_concurrency defaults to 8; use 1 for serial reads. metadata_buffer_bytes (8 MiB) and row_buffer_bytes (64 MiB) bound temporary transport batches, not retained caches or total RAM. An indivisible oversized unit is read alone. These positive-integer settings can also be changed on an open table; views use their base table’s settings. There is no automatic CPU/RAM-based tuning.

blosc2.open accepts max_concurrency but not the table-specific buffer keywords. Use this constructor or the returned table’s settings to tune buffers. Cache policies and max_cache_bytes remain independent.

A RemoteCTable’s cache policy applies to every column read through it, including columns backed by independent RemoteArray carriers. Those columns share the table owner’s cache budget and traffic accounting; their persisted standalone policies are not used or modified by the table.

hdf5_index accepts a native index dictionary or a local/remote JSON path for PyTables/HDF5 sources. Supplying one skips HDF5 discovery.

path selects the table within the source. dataset remains a supported alias; when both are supplied they must agree after stripping outer slashes. None leaves selection unspecified; an empty string or slash selects the root. Selector keywords cannot be combined with a selector embedded in the URL.

Attributes:
attrs

Read-only user attributes.

blocks

Block shape shared by the table’s aligned fixed-size columns.

cache_bytes

Compressed payload currently retained at this object’s scope.

cache_policy

Configured payload-retention policy.

cbytes

Total compressed size in bytes (all columns + valid_rows mask).

chunks

Chunk shape shared by the table’s aligned fixed-size columns.

computed_columns

Read-only view of the computed-column definitions.

cratio

Compression ratio for the whole table payload.

indexes

Return a list of blosc2.Index handles for all active indexes.

info

Get information about this table.

info_items

Structured summary items used by info().

is_cache_mutable

Whether the current local cache is writable, not the remote table.

max_cache_bytes

Configured compressed-payload allowance.

max_concurrency
metadata_buffer_bytes
metadata_bytes
mutable

The export default mutability for future reference exports.

nbytes

Total uncompressed size in bytes (all columns + valid_rows mask).

ncols

Total number of columns, including computed (virtual) columns.

nrows
row_buffer_bytes
schema

The compiled schema that drives this table’s columns and validation.

source

Credential-free descriptor for the selected remote object.

traffic

Source-read accounting for this object or its shared owner.

vlmeta

Variable-length metadata attached to this table.

Methods

add_column(name, spec[, values])

Add a new column filled from values, or from the default declared in spec.

add_computed_column(name, expr, *[, dtype, ...])

Add a read-only virtual column computed from stored columns.

add_generated_column(name, *, values[, ...])

Add a stored generated column maintained by the table.

append(data)

Append a single row to the table.

apply(func, *[, columns, dtype, engine])

Run a row-batch UDF over live column values and materialize the result.

assign(**named_exprs)

Return a view with additional computed columns, without copying data.

close()

Release resources owned by this handle.

column_schema(name)

Return the CompiledColumn descriptor for name.

compact()

Physically rewrite every column array keeping only live rows.

compact_index([col_name, expression, name])

Compact an index, merging any incremental append runs.

convert_nulls([columns, to, null_value, inplace])

Convert nullable columns between sentinel and validity-mask storage.

copy([compact, urlpath, overwrite, chunks, ...])

Return a new standalone copy of this table.

cov()

Return the covariance matrix as a numpy array.

create_index([col_name, field, expression, ...])

Build and register an index for a stored column or table expression.

delete(ind)

Mark one or more rows as deleted (tombstone deletion).

describe()

Print a per-column statistical summary.

drop_column(name)

Remove a column from the table.

drop_computed_column(name)

Remove a computed column from the table.

drop_index([col_name, expression, name])

Remove an index and delete any sidecar files.

dropna([subset])

Return a view excluding rows where any column in subset is null.

extend(data, *[, validate])

Append multiple rows at once.

from_arrow(schema[, batches, urlpath, mode, ...])

Build a CTable from an Arrow schema and iterable of record batches.

from_csv(path, row_cls, *[, header, sep, ...])

Build a CTable from a CSV file.

from_pandas(df, row_cls)

Build a CTable from a pandas DataFrame.

from_parquet(path, *[, columns, batch_size, ...])

Read a Parquet file into a CTable.

group_by(keys, *[, sort, dropna, engine, ...])

Return a deferred group-by object for this table.

head([N])

Return a view of the first N live rows (default 5).

index([col_name, expression, name])

Return the index handle for a stored-column or expression target.

iter_arrow_batches(*[, columns, batch_size, ...])

Yield live rows as Arrow batches (65,536 rows per batch by default).

iter_sorted(cols[, ascending, start, stop, ...])

Iterate rows in sorted order without materializing a full copy.

load(urlpath)

Load a persistent table from urlpath into RAM.

materialize(*[, urlpath, overwrite, ...])

Return an independent local copy of this table or view.

materialize_computed_column(name, *[, ...])

Materialize a computed column into a new stored snapshot column.

open(urlpath, *[, mode, mmap_mode])

Open a persistent CTable from urlpath.

open_reference(path, *[, storage_options, ...])

Reopen a saved Parquet RemoteStore archive.

rebuild_index([col_name, expression, name])

Drop and recreate an index with the same parameters.

refresh()

Reload a standalone table and invalidate its old columns and views.

refresh_generated_column(name)

Recompute a stored generated/materialized column from its source columns.

refresh_generated_columns(*[, source])

Refresh all generated columns, optionally only those depending on source.

rename_column(old, new)

Rename a column.

sample(n, *[, seed])

Return a read-only view of n randomly chosen live rows.

save([destination, urlpath, include_cache, ...])

Export this table as a portable remote-reference archive.

schema_dict()

Return a JSON-compatible dict describing this table's schema.

select(cols)

Return a column-projection view exposing only cols.

slice(start[, stop, copy])

Return a contiguous range of live (non-deleted) rows.

sort_by(cols[, ascending, inplace, view])

Return the table sorted by one or more columns.

sorted_slice(col, key, *[, ascending])

Return rows key in col-sorted order, reading only the slice window.

tail([N])

Return a view of the last N live rows (default 5).

take(indices, /)

Return a compact table containing rows at the requested positions.

to_arrow()

Convert all live rows to a pyarrow.Table.

to_b2d(urlpath, *[, overwrite, compact, ...])

Write this table to a directory-backed store.

to_b2z(urlpath, *[, overwrite, compact, ...])

Write this table to a compact .b2z container.

to_cframe(*[, preserve_sources])

Serialize this table to a bytes buffer (a CFrame).

to_csv([path, header, sep, encoding])

Write all live rows to CSV.

to_pandas()

Convert to a pandas DataFrame.

to_parquet(path, *[, columns, batch_size, ...])

Write Parquet in batches of 65,536 rows by default.

to_string(*[, max_rows, max_width, ...])

Return a tabular string representation of the table.

trim_capacity()

Shrink fixed-width physical storage to the last live row position.

view(new_valid_rows)

Return a row-filter view backed by a boolean mask array without copying data.

where(expr_result, *[, columns])

Return a row-filtered view matching a boolean predicate.

with_sparse_cache(urlpath, runtime_cache_path, *)

Attach a remote CTable to a sparse disk cache shared across processes.

property attrs

Read-only user attributes.

property cache_bytes

Compressed payload currently retained at this object’s scope.

property cache_policy

Configured payload-retention policy.

close() → None[source]

Release resources owned by this handle.

property is_cache_mutable: bool

Whether the current local cache is writable, not the remote table.

property max_cache_bytes

Configured compressed-payload allowance.

property mutable: bool

The export default mutability for future reference exports.

classmethod open_reference(path, *, storage_options=None, parquet_options=None)[source]

Reopen a saved Parquet RemoteStore archive.

refresh() → None[source]

Reload a standalone table and invalidate its old columns and views.

Preserve the table and its cache on discovery/initialization failure. For tables obtained from a RemoteStore, refresh the root store instead.

save(destination: str | PathLike | None = None, *, urlpath: str | PathLike | None = None, include_cache: bool = True, mutable: bool | None = None, overwrite: bool = False) → str[source]

Export this table as a portable remote-reference archive.

property source

Credential-free descriptor for the selected remote object.

property traffic

Source-read accounting for this object or its shared owner.

property vlmeta

Variable-length metadata attached to this table.

Returns a mapping-like proxy that supports item access, iteration, and the [:] bulk getter. Values are serialised via msgpack, so all standard types (int, float, str, bool, list, dict) are supported. The metadata is stored separately from the internal schema metadata and persists through close() / reopen for disk-backed tables.

Examples

>>> import blosc2
>>> import dataclasses
>>> @dataclasses.dataclass
... class Row:
...     x: int = 0
>>> t = blosc2.CTable(Row)
>>> t.attrs["author"] = "Alice"
>>> t.attrs["tags"] = ["alpha", "beta"]
>>> t.attrs["count"] = 42
>>> print(t.attrs["author"])
Alice
>>> print(t.attrs[:])
{'author': 'Alice', 'tags': ['alpha', 'beta'], 'count': 42}
>>> del t.attrs["count"]
>>> for name in t.attrs:
...     print(name, t.attrs[name])
...
author Alice
tags ['alpha', 'beta']
classmethod with_sparse_cache(urlpath, runtime_cache_path, *, dataset=None, path=None, manifest=None, max_cache_bytes=<policy default>, carrier=None, storage_options=None, max_concurrency=8, metadata_buffer_bytes=8388608, row_buffer_bytes=67108864, _filesystem=None, _filesystem_resolver=None, _batch_validator=None, _source_validator=None, _manifest_validator=None, _max_nodes=None, source_format=None)[source]

Attach a remote CTable to a sparse disk cache shared across processes.

path and dataset select the table as in the ordinary constructor. The aggregate compressed-payload budget defaults to 256 MiB; pass max_cache_bytes=None for unlimited retention. For ordinary shared caching, prefer blosc2.open(url, cache_dir=..., shared_cache=True).