<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet type="text/xsl" href="../assets/xml/rss.xsl" media="all"?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Blosc Home Page  (Posts about arrow)</title><link>https://blosc.org/</link><description></description><atom:link href="https://blosc.org/categories/arrow.xml" rel="self" type="application/rss+xml"></atom:link><language>en</language><copyright>Contents © 2026 &lt;a href="mailto:blosc@blosc.org"&gt;The Blosc Developers&lt;/a&gt; </copyright><lastBuildDate>Wed, 30 Sep 2026 11:03:43 GMT</lastBuildDate><generator>Nikola (getnikola.com)</generator><docs>http://blogs.law.harvard.edu/tech/rss</docs><item><title>Common Interface for Remote and Local Datasets</title><link>https://blosc.org/posts/remote-local-objects/</link><dc:creator>Francesc Alted</dc:creator><description>&lt;div&gt;&lt;p&gt;I am happy to announce that the new &lt;strong&gt;Python-Blosc2 4.14&lt;/strong&gt; brings a common interface to remote Parquet tables, Zarr and HDF5 arrays, PyTables tables, and indeed, native Blosc2 containers.
Most interestingly, a local cache keeps fetched data available for the next request/query; moreover, if cache is persistent, you can reuse the downloaded data throughout different sessions.&lt;/p&gt;
&lt;h3&gt;TL;DR&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One opener for everything&lt;/strong&gt;: &lt;code&gt;blosc2.open()&lt;/code&gt; provides a unified interface across remote Parquet, Zarr, HDF5, PyTables, and native Blosc2 datasets over HTTP, S3, or any &lt;code&gt;fsspec&lt;/code&gt; filesystem.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zero full downloads&lt;/strong&gt;: Stream only the exact chunks, row groups, or physical columns your queries touch—no need to pull gigabytes of data to inspect a slice.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Smart persistent caching&lt;/strong&gt;: Passing &lt;code&gt;cache_dir=...&lt;/code&gt; automatically caches fetched data locally across sessions and processes, eliminating redundant network roundtrips.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Index-accelerated remote queries&lt;/strong&gt;: Remote queries automatically exploit persisted PyTables and Blosc2 indexes without fetching entire columns or rebuilding indexes locally.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ecosystem interoperability&lt;/strong&gt;: Combine remote arrays across different formats directly in lazy expressions (&lt;code&gt;b + z * h5&lt;/code&gt;), or stream Arrow batches into DuckDB and Polars.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To whet your appetite, here is a practical tour using public datasets on Backblaze B2 (which is fully compatible with S3).
Here we are going to exercise the HTTP path, but all the protocols supported by &lt;a href="https://filesystem-spec.readthedocs.io/en/latest/index.html"&gt;fsspec&lt;/a&gt; (like S3, GDrive, GCS, SFTP and many others) should be supported too.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://blosc.org/posts/remote-local-objects/"&gt;Read more…&lt;/a&gt; (4 min remaining to read)&lt;/p&gt;&lt;/div&gt;</description><category>arrow</category><category>Blosc2</category><category>cache</category><category>ctable</category><category>hdf5</category><category>parquet</category><category>pytables</category><category>remote</category><category>zarr</category><guid>https://blosc.org/posts/remote-local-objects/</guid><pubDate>Wed, 30 Sep 2026 10:45:00 GMT</pubDate></item></channel></rss>