CLOUD-NES (pilot workshop)
Hands-on workshop: transform geospatial data for the cloud, test formats and chunking, and benchmark real-world remote access.
Hands-on workshop: transform geospatial data for the cloud, test formats and chunking, and benchmark real-world remote access.
Cloud-native technologies are changing how research data is stored, accessed and shared. But "cloud-native" isn't just "put it in a bucket." Formats well suited to cloud environments such as Zarr, Cloud-Optimized GeoTIFF and GeoParquet are increasingly adopted, yet practical questions remain: how should a dataset be structured for efficient remote access, which format suits which use case, and how do you measure whether it actually helps?
In this hands-on workshop, we take real geospatial datasets from legacy formats to cloud-native representations, experiment with chunking and other parameters, and benchmark the results under realistic access patterns.
Audience
This workshop is for researchers and research support staff who work with (or plan to publish) large geospatial datasets: PhD candidates, postdocs, data stewards and research software engineers.
Prerequisites
The level is intermediate. You should be comfortable with: Writing and running Python in Jupyter notebooks Using xarray or (geo)pandas for basic loading, selecting and plotting Working with common geospatial concepts: raster versus vector data, coordinate reference systems, and formats such as GeoTIFF and NetCDF ⠀No prior experience with cloud computing, object storage or Zarr/Parquet is needed.
Syllabus
Why cloud-native ? HTTP range requests, chunking and tiling. Reading from object storage: S3 access in Python. Converting to cloud-optimized formats: COG, Zarr and GeoParquet, and how chunking, compression and overviews affect them. Benchmarking: time, requests and bytes for typical access patterns. Virtual datasets: making legacy formats cloud-accessible. Publishing with STAC : making datasets findable.
