scottstanie
Hi, I'm not sure how often these get updated, but I saw that in the [Cloud tutorial](https://github.com/HDFGroup/hdf5-tutorial/blob/main/07-S3-and-the-Cloud.ipynb), it simply says that "the native library is inefficient, use `h5coro` instead". Perhaps the section could link to [Aleksandar's presentation](https://hdfeos.org/workshops/ws25/presentations/axj.pdf) describing how to write files that are efficient for remote access, and maybe some kind of `h5py` demo where you've added the `fs_page_size` option: ```python from typing import Iterator from contextlib import contextmanager import h5py @contextmanager def get_remote_h5( url: str, page_size: int = 4 * 1024 * 1024, rdcc_nbytes: int = 1024 * 1024 * 100, ) -> Iterator[h5py.File]: """Open a remote HDF5 file using the ROS3 driver. Parameters ---------- url : str S3 URL to the HDF5 file. page_size : int, optional File system page size in bytes. Default is 4 MB. rdcc_nbytes : int, optional Raw data chunk cache size in bytes. Default is 100 MB. Yields ------ h5py.File Opened HDF5 file. """ import boto3 session = boto3.Session() creds = session.get_credentials() frozen_creds = creds.get_frozen_credentials() ros3_kwargs = { "aws_region": b"us-west-2", # Hard-coded or detect from environment "secret_id": frozen_creds.access_key.encode(), "secret_key": frozen_creds.secret_key.encode(), } if frozen_creds.token: ros3_kwargs["session_token"] = frozen_creds.token.encode() # Set page size for cloud-optimized HDF5 cloud_kwargs = dict(fs_page_size=page_size, rdcc_nbytes=rdcc_nbytes) with h5py.File(url, "r", driver="ros3", **ros3_kwargs, **cloud_kwargs) as hf: yield hf ```