RemoteCTable¶
RemoteCTable is a read-only blosc2.CTable backed by a remote B2Z
archive, a local or remote PyTables/HDF5 table, or a Parquet file. Fixed-width,
shaped, nullable, UTF-8, batch-backed variable-length,
batch-backed list, struct/object, and dictionary columns are fetched on demand.
Open local PyTables tables through blosc2.open() with path= or a
::table selector.
Standalone tables can be opened directly; tables inside a hierarchy can be
selected with dataset= or through blosc2.RemoteStore.
PyTables/HDF5 sources may supply hdf5_index= as a native index dictionary,
local JSON path, or remote fsspec URL. This skips source discovery and does not
modify the HDF5 file; see Working with Remote Arrays.
Batch-backed reads transfer and decode whole compressed batches. Dictionary codes remain selective, while the full vocabulary is loaded on first use. ListArray batches contain 2048 list cells by default. Smaller batches reduce overfetch for sparse reads; larger batches usually improve scans and compression. The row limit is not a byte limit, so one unusually large list can still require a large transfer. See ListArray for batching controls.
Nullable list elements, nested lists, and contains/overlaps predicates
have the same semantics as local tables. Without an index a predicate scans the
remote list batches. A persisted kind="membership" index on a flat scalar
list fetches only the compressed posting batches for requested values. If the
result projects only other columns, the list payload remains unopened. Indexes
on nested lists and structs are not supported. ListArray storage="vl" also
remains unavailable through RemoteCTable.
Scalar queries automatically use persisted SUMMARY, FULL, PARTIAL,
OPSI, and BUCKET indexes. Their sidecars are opened lazily and participate
in the outer table’s cache budget and traffic accounting. SUMMARY reads compact
min/max records before fetching candidate data blocks; positional indexes use
their navigation data to fetch selected value and row-position ranges. Queries
retain a correct scan fallback when an index layout or expression is unsupported.
Saving and materializing have different meanings:
with blosc2.RemoteCTable(url) as table:
table["temperature"][:100] # warm part of one column
table.save("reference.b2z")
table.save("cold-reference.b2z", include_cache=False)
local = table.materialize(urlpath="complete.b2z")
save() writes the source descriptor, table metadata, and by default only
payload already retained in the cache. Missing data is still read from the
original source after reopening. materialize() returns an independent local
table and reads all data needed for it. copy(), to_b2z(), and to_b2d()
remain local materialization operations inherited from blosc2.CTable.
Table cache bytes, limits, and traffic are scoped to the shared remote owner and
may include sibling leaves. This also applies when a table column is itself a
RemoteArray reference to another fsspec or Caterva2 URL: the outer table
policy overrides the carrier policy for that handle, and all such columns share
one budget. Standalone instances of those arrays keep their original policies.
A table selected from a store owns an independent handle, but its columns and
views remain borrowed from that table. Refresh a nested table through its root
store; a standalone table can call refresh().
Server processes can share one bounded sparse disk cache with
blosc2.open(url, cache_dir="shared-cache", shared_cache=True). All processes
using that directory must enable sharing. Operations serialize per store,
and the aggregate compressed-payload budget defaults to 256 MiB; explicitly
pass max_cache_bytes=None for unlimited retention. Use a separate directory
from ordinary exclusive caches. The budget does not bound total disk usage or
peak RAM.
RemoteCTable.with_sparse_cache(url, runtime_cache_path) remains available
for advanced attachment with manifests or seed carriers, with the same default
budget and shared-cache implementation.
Parquet cache details¶
Parquet cache_dir stores each accessed physical field and row group as a
native CTable directory under the RemoteStore generation’s parquet-groups
directory. The common manifest retains the footer, schema, conversion options,
source marker, and row-group boundaries. Warm opens reuse this discovery and
payload without contacting the source. Cached sources are assumed immutable
until refresh() is called. New portable .b2z archives use the common
RemoteStore manifest. Local Parquet sources use the same cache layout and may
export .b2z references tied to the local source path.
HTTP Parquet sources need byte-range support. A persistent cache also needs a
source size and version marker, such as an ETag, modification time, or
Backblaze B2 file ID. HTTP read-ahead is disabled by default to avoid fetching
unused bytes; an explicit storage_options={"cache_type": "bytes"} restores
fsspec buffering. The traffic counter counts fetched ranges; explicitly
enabled transport buffering may make actual HTTP transfer different.
Multi-column selections, including table previews, fetch independent Parquet
column chunks concurrently using max_concurrency (8 by default). Decoding
and cache writes stay on the calling thread. row_buffer_bytes bounds each
wave of compressed responses; one physical field larger than the budget runs
alone. A physical field with nested leaves is fetched as one unit. Reads still
load whole column chunks for the selected row groups, even for a few rows.
Use max_concurrency=1 for serial transport.
See Working with Remote Tables for column access, filtering, buffering, reference saving, and materialization examples.
- class blosc2.RemoteCTable(urlpath=None, *, dataset=None, path=None, storage_options=None, cache_policy=<policy default>, max_cache_bytes=<policy default>, cache_dir=None, shared_cache=False, hdf5_index=None, max_concurrency=8, metadata_buffer_bytes=8388608, row_buffer_bytes=67108864, source_format=None, parquet_options=None, columns=None, max_rows=None, string_max_length=None, null_storage=None, auto_null_sentinels=True, separate_nested_cols=True, list_serializer='msgpack', blosc2_batch_size=2048, blosc2_items_per_block=None, batch_size=2048, cparams=None, dparams=None, validate=False, _filesystem=None, _filesystem_resolver=None, _batch_validator=None)[source]¶
A read-only CTable whose columns are fetched on demand.
Supported columns include fixed-width, UTF-8, batch-backed variable-length, batch-backed list, struct/object and dictionary columns. Batch reads transfer one whole compressed batch; dictionary decoding loads the full vocabulary on first use.
Independent column requests overlap by default.
max_concurrencydefaults to 8; use 1 for serial reads.metadata_buffer_bytes(8 MiB) androw_buffer_bytes(64 MiB) bound temporary transport batches, not retained caches or total RAM. An indivisible oversized unit is read alone. These positive-integer settings can also be changed on an open table; views use their base table’s settings. There is no automatic CPU/RAM-based tuning.blosc2.openacceptsmax_concurrencybut not the table-specific buffer keywords. Use this constructor or the returned table’s settings to tune buffers. Cache policies andmax_cache_bytesremain independent.A RemoteCTable’s cache policy applies to every column read through it, including columns backed by independent RemoteArray carriers. Those columns share the table owner’s cache budget and traffic accounting; their persisted standalone policies are not used or modified by the table.
hdf5_indexaccepts a native index dictionary or a local/remote JSON path for PyTables/HDF5 sources. Supplying one skips HDF5 discovery.pathselects the table within the source.datasetremains a supported alias; when both are supplied they must agree after stripping outer slashes. None leaves selection unspecified; an empty string or slash selects the root. Selector keywords cannot be combined with a selector embedded in the URL.- Attributes:
attrsRead-only user attributes.
blocksBlock shape shared by the table’s aligned fixed-size columns.
cache_bytesCompressed payload currently retained at this object’s scope.
cache_policyConfigured payload-retention policy.
cbytesTotal compressed size in bytes (all columns + valid_rows mask).
chunksChunk shape shared by the table’s aligned fixed-size columns.
computed_columnsRead-only view of the computed-column definitions.
cratioCompression ratio for the whole table payload.
indexesReturn a list of
blosc2.Indexhandles for all active indexes.infoGet information about this table.
info_itemsStructured summary items used by
info().is_cache_mutableWhether the current local cache is writable, not the remote table.
max_cache_bytesConfigured compressed-payload allowance.
- max_concurrency
- metadata_buffer_bytes
- metadata_bytes
mutableThe export default mutability for future reference exports.
nbytesTotal uncompressed size in bytes (all columns + valid_rows mask).
ncolsTotal number of columns, including computed (virtual) columns.
- nrows
- row_buffer_bytes
schemaThe compiled schema that drives this table’s columns and validation.
sourceCredential-free descriptor for the selected remote object.
trafficSource-read accounting for this object or its shared owner.
vlmetaVariable-length metadata attached to this table.
Methods
add_column(name, spec[, values])Add a new column filled from values, or from the default declared in spec.
add_computed_column(name, expr, *[, dtype, ...])Add a read-only virtual column computed from stored columns.
add_generated_column(name, *, values[, ...])Add a stored generated column maintained by the table.
append(data)Append a single row to the table.
apply(func, *[, columns, dtype, engine])Run a row-batch UDF over live column values and materialize the result.
assign(**named_exprs)Return a view with additional computed columns, without copying data.
close()Release resources owned by this handle.
column_schema(name)Return the
CompiledColumndescriptor for name.compact()Physically rewrite every column array keeping only live rows.
compact_index([col_name, expression, name])Compact an index, merging any incremental append runs.
convert_nulls([columns, to, null_value, inplace])Convert nullable columns between sentinel and validity-mask storage.
copy([compact, urlpath, overwrite, chunks, ...])Return a new standalone copy of this table.
cov()Return the covariance matrix as a numpy array.
create_index([col_name, field, expression, ...])Build and register an index for a stored column or table expression.
delete(ind)Mark one or more rows as deleted (tombstone deletion).
describe()Print a per-column statistical summary.
drop_column(name)Remove a column from the table.
drop_computed_column(name)Remove a computed column from the table.
drop_index([col_name, expression, name])Remove an index and delete any sidecar files.
dropna([subset])Return a view excluding rows where any column in subset is null.
extend(data, *[, validate])Append multiple rows at once.
from_arrow(schema[, batches, urlpath, mode, ...])Build a
CTablefrom an Arrow schema and iterable of record batches.from_csv(path, row_cls, *[, header, sep, ...])Build a
CTablefrom a CSV file.from_pandas(df, row_cls)Build a
CTablefrom a pandas DataFrame.from_parquet(path, *[, columns, batch_size, ...])Read a Parquet file into a
CTable.group_by(keys, *[, sort, dropna, engine, ...])Return a deferred group-by object for this table.
head([N])Return a view of the first N live rows (default 5).
index([col_name, expression, name])Return the index handle for a stored-column or expression target.
iter_arrow_batches(*[, columns, batch_size, ...])Yield live rows as Arrow batches (65,536 rows per batch by default).
iter_sorted(cols[, ascending, start, stop, ...])Iterate rows in sorted order without materializing a full copy.
load(urlpath)Load a persistent table from urlpath into RAM.
materialize(*[, urlpath, overwrite, ...])Return an independent local copy of this table or view.
materialize_computed_column(name, *[, ...])Materialize a computed column into a new stored snapshot column.
open(urlpath, *[, mode, mmap_mode])Open a persistent CTable from urlpath.
open_reference(path, *[, storage_options, ...])Reopen a saved Parquet RemoteStore archive.
rebuild_index([col_name, expression, name])Drop and recreate an index with the same parameters.
refresh()Reload a standalone table and invalidate its old columns and views.
refresh_generated_column(name)Recompute a stored generated/materialized column from its source columns.
refresh_generated_columns(*[, source])Refresh all generated columns, optionally only those depending on source.
rename_column(old, new)Rename a column.
sample(n, *[, seed])Return a read-only view of n randomly chosen live rows.
save([destination, urlpath, include_cache, ...])Export this table as a portable remote-reference archive.
schema_dict()Return a JSON-compatible dict describing this table's schema.
select(cols)Return a column-projection view exposing only cols.
slice(start[, stop, copy])Return a contiguous range of live (non-deleted) rows.
sort_by(cols[, ascending, inplace, view])Return the table sorted by one or more columns.
sorted_slice(col, key, *[, ascending])Return rows
keyincol-sorted order, reading only the slice window.tail([N])Return a view of the last N live rows (default 5).
take(indices, /)Return a compact table containing rows at the requested positions.
to_arrow()Convert all live rows to a
pyarrow.Table.to_b2d(urlpath, *[, overwrite, compact, ...])Write this table to a directory-backed store.
to_b2z(urlpath, *[, overwrite, compact, ...])Write this table to a compact
.b2zcontainer.to_cframe(*[, preserve_sources])Serialize this table to a bytes buffer (a CFrame).
to_csv([path, header, sep, encoding])Write all live rows to CSV.
to_pandas()Convert to a pandas DataFrame.
to_parquet(path, *[, columns, batch_size, ...])Write Parquet in batches of 65,536 rows by default.
to_string(*[, max_rows, max_width, ...])Return a tabular string representation of the table.
trim_capacity()Shrink fixed-width physical storage to the last live row position.
view(new_valid_rows)Return a row-filter view backed by a boolean mask array without copying data.
where(expr_result, *[, columns])Return a row-filtered view matching a boolean predicate.
with_sparse_cache(urlpath, runtime_cache_path, *)Attach a remote CTable to a sparse disk cache shared across processes.
- property attrs¶
Read-only user attributes.
- property cache_bytes¶
Compressed payload currently retained at this object’s scope.
- property cache_policy¶
Configured payload-retention policy.
- property max_cache_bytes¶
Configured compressed-payload allowance.
- classmethod open_reference(path, *, storage_options=None, parquet_options=None)[source]¶
Reopen a saved Parquet RemoteStore archive.
- refresh() None[source]¶
Reload a standalone table and invalidate its old columns and views.
Preserve the table and its cache on discovery/initialization failure. For tables obtained from a RemoteStore, refresh the root store instead.
- save(destination: str | PathLike | None = None, *, urlpath: str | PathLike | None = None, include_cache: bool = True, mutable: bool | None = None, overwrite: bool = False) str[source]¶
Export this table as a portable remote-reference archive.
- property source¶
Credential-free descriptor for the selected remote object.
- property traffic¶
Source-read accounting for this object or its shared owner.
- property vlmeta¶
Variable-length metadata attached to this table.
Returns a mapping-like proxy that supports item access, iteration, and the
[:]bulk getter. Values are serialised via msgpack, so all standard types (int, float, str, bool, list, dict) are supported. The metadata is stored separately from the internal schema metadata and persists throughclose()/ reopen for disk-backed tables.Examples
>>> import blosc2 >>> import dataclasses >>> @dataclasses.dataclass ... class Row: ... x: int = 0 >>> t = blosc2.CTable(Row) >>> t.attrs["author"] = "Alice" >>> t.attrs["tags"] = ["alpha", "beta"] >>> t.attrs["count"] = 42 >>> print(t.attrs["author"]) Alice >>> print(t.attrs[:]) {'author': 'Alice', 'tags': ['alpha', 'beta'], 'count': 42} >>> del t.attrs["count"] >>> for name in t.attrs: ... print(name, t.attrs[name]) ... author Alice tags ['alpha', 'beta']
- classmethod with_sparse_cache(urlpath, runtime_cache_path, *, dataset=None, path=None, manifest=None, max_cache_bytes=<policy default>, carrier=None, storage_options=None, max_concurrency=8, metadata_buffer_bytes=8388608, row_buffer_bytes=67108864, _filesystem=None, _filesystem_resolver=None, _batch_validator=None, _source_validator=None, _manifest_validator=None, _max_nodes=None, source_format=None)[source]¶
Attach a remote CTable to a sparse disk cache shared across processes.
pathanddatasetselect the table as in the ordinary constructor. The aggregate compressed-payload budget defaults to 256 MiB; passmax_cache_bytes=Nonefor unlimited retention. For ordinary shared caching, preferblosc2.open(url, cache_dir=..., shared_cache=True).