Cloud-Native Architectures and Modern Data Formats for Operational Oceanography, Climate, and Meteorology at Multi-Terabyte Scale.

September 1-3, 2026

9:00 - 17:30 UTC-3

Instructor: Tobias Ferreira

REGISTRATION CLOSED ON JULY 3. SELECTED CANDIDATES HAVE ALREADY BEEN CONTACTED.

Participants were selected based on institution, availability, and candidate profile, aiming to support diversity, inclusion, and the best use of the course.

General Information

ABOUT THIS COURSE: This course covers the complete workflow for modernizing n-dimensional scientific data used in oceanography, climate, and meteorology. Participants will learn how to migrate data from traditional formats, such as NetCDF, to cloud-native formats such as Zarr and VirtualiZarr, enabling efficient and scalable access to multi-terabyte datasets. The course covers the full data pipeline, including object storage strategies, data conversion, parallel processing using tools such as Dask, and efficient cloud-based data access and visualization. By the end of the course, participants will have a practical understanding of how to build modern, scalable, and reproducible workflows for analyzing large volumes of environmental data.

COURSE NOTES: Course materials, including examples and notebooks, will be available on a page dedicated to this course in the project repository on GitHub. Participants will receive a link to access the material before the course begins.

WHO CAN ATTEND?: This course is intended for researchers, data scientists, and engineers who work with oceanography, climate, meteorology, or related fields. It is especially relevant for those who need to handle large volumes of data (gigabytes to terabytes) and want to modernize their workflows using cloud-native technologies. The course is open to participants from different institutions, subject to seat availability.

WHERE:

In person: R. Aloisio Teixeira, 278 - Cidade Universitaria, Rio de Janeiro. Follow the instructions below to get to the course venue: OpenStreetMap or Google Maps.

Online participants: For online participants, a link to the virtual meeting will be provided before the course starts.

WHEN: September 1-3, 2026; 9:00 - 17:30 UTC-3 Add to your Google Calendar.

PREREQUISITES: Previous experience with cloud computing or cloud-native formats is not required. However, the course will be taught in Python, and participants are expected to have good familiarity with the language (loops, functions, use of NumPy and xarray libraries). Participants should also be comfortable using a command-line terminal and have prior experience with scientific data analysis and traditional formats such as NetCDF or GRIB.

REQUIREMENTS:

In person: Participants must bring a laptop with internet access. All required course setup will be done in a cloud environment, so no specific software installation is required locally.

Online participants: Unlike in-person participants, online participants will use their own computers during the course. It will be necessary to install development tools in advance, such as Python, Jupyter, and VS Code (or equivalent), and to ensure a stable internet connection. Detailed instructions will be sent by email before the course begins. To ensure that everyone is ready, we will hold two setup sessions before the course starts (dates will be announced soon). During these sessions, we will be available to guide installation and solve potential issues. Minimum hardware requirements for online participants include: Intel Core i5 CPU or better, with at least 4 cores; 16 GB RAM as a comfortable minimum for data science (32 GB is excellent if you already work with large volumes); and 512 GB or more of storage, ensuring enough room for conda or venv environments, packages, caches, and some local datasets.

REGISTRATION: REGISTRATION CLOSED. Participants were selected based on institution, availability, and candidate profile, aiming to support diversity, inclusion, and the best use of the course. Selected participants received a confirmation email on July 17.

CONTACT: Please send an email to tobias.ferreira@noc.ac.uk for more information.

WORKSHOP ROLES: To learn more about workshop roles (who will do what), see our Workshop FAQ.

THE CARPENTRIES: The Carpentries project includes a community of instructors, helpers, maintainers, and others. Each of these roles has specific responsibilities to ensure the workshop is a positive experience for all participants.

ACCESSIBILITY: We are committed to making this course accessible to everyone.

We are dedicated to providing a positive and accessible learning environment for everyone. We do not require participants to provide documentation of disabilities or disclose any unnecessary personal information. However, we want to help create an inclusive and accessible experience for all participants. We encourage you to share any information that may be helpful for making your Carpentries experience accessible.


Schedule

September 1 - Foundations and Formats

Before Pre-workshop survey
09:00 Opening: INPO
09:15 Course introduction, goals, and setup
09:45 Introduction to Data Formats, Metadata, and Vocabulary
10:20 Challenges of n-dimensional data
10:40 BREAK
11:00 What are cloud-native formats?
11:40 Zarr - Data Model, Metadata, and Chunked Storage
12:40 Python tools for working with Zarr
13:00 LUNCH
14:00 Python tools for working with Zarr
15:00 How to Choose Chunks for Analysis and Processing at Scale
15:30 BREAK
15:50 Parallel Processing for Zarr
16:50 Daily wrap-up and discussion
17:00 Close

September 2 - Storage, Conversion, and Versioning

09:00 Opening and recap
09:10 Hands-on: Zarr datasets
10:40 BREAK
11:10 Hands-on: Zarr datasets
11:30 Object storage: concepts, remote access, and organizing data in the cloud
12:20 Conversion Workflow of Traditional Formats to Zarr
13:00 LUNCH
14:00 Conversion Workflow of Traditional Formats to Zarr
15:30 BREAK
15:50 Data versioning with Icechunk
16:50 Daily wrap-up and discussion
17:00 Close

September 3 - Virtualization, Interoperability, and Visualization

09:00 Opening and recap
09:10 Creating virtual Zarr stores with VirtualiZarr
10:30 BREAK
10:50 Talk session: how are other institutions using Zarr?
11:50 Zarr Data Organization in the Cloud with STAC
13:00 LUNCH
14:00 Visualization - Multiscale Zarr and GeoZarr
15:30 BREAK
15:50 Discussion: Architectures and Best Practices
16:40 Closing and post-workshop survey
17:00 End

The material to be taught in this workshop is being piloted and a precise schedule has not yet been established. The workshop will include regular breaks. Please contact the workshop organizers if you would like more information about the planned schedule.


Setup

For in-person participants, there is no need to install any specific software locally, as all the required setup for the course will be provided through a cloud-based environment. However, participants are encouraged to bring a laptop with internet access so they can follow along with the hands-on activities.

For online participants, you will need to install several development tools in advance, including Python, Jupyter, and VS Code (or an equivalent editor), and ensure that you have a stable internet connection. Detailed installation instructions will be sent by email before the course begins. To help everyone get ready, we will hold two setup sessions before the start of the course (dates will be announced soon). During these sessions, we will be available to guide you through the installation process and help resolve any technical issues.

Note for Online Participants

We would like to be transparent about the course format.
Since the in-person sessions will involve a high level of interaction, hands-on exercises, and one-on-one support, we unfortunately will not be able to provide the same level of real-time assistance to participants attending remotely.
That said, we will do our best to include online participants throughout all course activities. I will be doing extensive live coding during the course, explaining each step of the development process, and I will make every effort to answer questions live. However, I may not be able to respond to every question in real time, especially if there are many questions at once.
In addition, I will also be available before the course during the setup sessions, and after the course, to help online participants whenever possible.
Our goal is for everyone to be able to follow the course content and get the most out of the experience.

Code of Conduct

All participants in Carpentries activities are required to follow the Code of Conduct. This document also explains how to report an incident if necessary.