REGISTRATION CLOSED ON JULY 3.
Participants were selected based on institution, availability, and candidate profile, aiming to support diversity, inclusion, and the best use of the course. The selected participants received a confirmation email on July 17.
The Arq-Cloud course covers the complete workflow for modernizing n-dimensional scientific data used in oceanography, climate, and meteorology. Participants will learn how to migrate data from traditional formats, such as NetCDF, to cloud-native formats such as Zarr and VirtualiZarr, enabling efficient and scalable access to multi-terabyte datasets. The course covers the full data pipeline, including object storage strategies, data conversion, parallel processing using tools such as Dask, and efficient cloud-based data access and visualization. By the end of the course, participants will have a practical understanding of how to build modern, scalable, and reproducible workflows for analyzing large volumes of environmental data.
Course materials, including examples and notebooks, will be available on a page dedicated to this course in the project repository on GitHub. Participants will receive a link to access the material before the course begins.
This course is intended for researchers, data scientists, and engineers who work with oceanography, climate, meteorology, or related fields. It is especially relevant for those who need to handle large volumes of data (gigabytes to terabytes) and want to modernize their workflows using cloud-native technologies. The course is open to participants from different institutions, subject to seat availability.
In person: R. Aloísio Teixeira, 278 - Cidade Universitária, Rio de Janeiro. Follow the instructions below to get to the course venue: OpenStreetMap or Google Maps.
Online participants: For online participants, a link to the virtual meeting will be provided before the course starts.
September 1-3, 2026; 9:00 - 17:30 UTC-3 Add to your Google Calendar.
Previous experience with cloud computing or cloud-native formats is not required. However, the course will be taught in Python, and participants are expected to have good familiarity with the language (loops, functions, use of NumPy and xarray libraries). Participants should also be comfortable using a command-line terminal and have prior experience with scientific data analysis and traditional formats such as NetCDF or GRIB.
Participants must bring a laptop with internet access. All setup required for the course will be carried out in a cloud environment, so no specific software needs to be installed locally.
During the course, we will use the computational services and resources provided by JASMIN. JASMIN is the UK system that provides computing resources for data analysis in environmental sciences.
JASMIN ACCOUNT CREATION: You received an email with instructions to create a JASMIN account on the 21st and the 24th of August. It is essential that you complete this setup before the beginning of the course.
SETUP MATERIAL: We have prepared setup material with instructions for configuring the environment on your own machines after the course: https://noc-oi.github.io/prep-work-cloud-native-geoscience-course/. It includes instructions for Windows, Linux, and macOS. Please note that, during the course, in-person participants should use JASMIN, as this will facilitate the learning process and help instructors provide support with any potential issues.
INTRODUCTORY MATERIAL: The setup material also includes introductory content on:
We highly recommend that in-person participants read this material before the course to align their prior knowledge and allow everyone to make the most of the workshop.
SETUP SUPPORT SESSIONS: We will hold three online sessions to assist participants in setting up their environments. For in-person participants, since they will be using JASMIN during the course, these sessions will focus on helping them configure their local environments to work with the course material after the activities conclude. During these sessions, we will provide a brief introduction to the course, demonstrate the installation process, and be available to help troubleshoot any issues. Please refer to the SETUP SUPPORT SESSIONS section in the requirements for online participantes for more information.
PRE-COURSE QUESTIONNAIRE: Before the workshop, we would like to better understand your familiarity with the topics to be covered so that we can adapt the content and the time dedicated to each topic, making the workshop more suitable for participants' needs. The questionnaire also includes practical questions about course organization, such as dietary restrictions during the event and consent for photography and image use. We have therefore prepared this questionnaire: https://forms.gle/DJf1WaMkb1RaukpZ7. Please complete it by 21 August.
PRACTICAL INFORMATION: We have prepared a welcome document for all in-person participants, with information about the course, logistics, transport, meals, hotels, and other important guidance. Please read this document carefully before the course begins.
Online participants will use their own computers throughout the course. They must install the tools described in the setup material in advance and ensure a stable internet connection.
MINIMUM SPECIFICATIONS: To ensure a good course experience, we recommend that online participants use a computer with at least an Intel Core i5 processor (or equivalent) with 8 cores, 16 GB of RAM (32 GB recommended for larger datasets), and approximately 35 GB of available storage.
SETUP MATERIAL: The setup material is available at https://noc-oi.github.io/prep-work-cloud-native-geoscience-course/. It contains instructions for Windows, Linux, and macOS, including Python installation, example dataset downloads, and installation verification. You should begin the course (01 September, at 9:00 AM) with your local environment configured and tested and the example datasets downloaded.
INTRODUCTORY MATERIAL: The setup material also includes introductory content on:
We highly recommend that online participants read this introductory material before the course to align their prior knowledge and allow everyone to make the most of the workshop.
SETUP SUPPORT SESSIONS: Three online sessions will be held to assist participants with configuring their environment. During these sessions, we will provide a brief introduction to the course, demonstrate the installation process, and be available to help resolve any issues.
The Meeting ID and Passcode information for these sessions has already been sent to participants. The sessions are optional, and participants may attend whichever session is most convenient.
PRE-COURSE QUESTIONNAIRE: Before the workshop, we would like to better understand your familiarity with the topics to be covered so that we can adapt the content and the time dedicated to each topic, making the workshop more suitable for participants' needs. The questionnaire also includes practical questions about course organization, such as consent for photography and image use. We have therefore prepared this questionnaire: https://forms.gle/DJf1WaMkb1RaukpZ7. Please complete it by 21 August.
</div>
We would like to be transparent about the course format.
Since the in-person activities will involve substantial interaction, exercises, and individual support, we unfortunately cannot provide the same level of real-time assistance to remote participants.
Even so, we will do our best to include online participants in all activities. I will do extensive live coding during the course, explaining each development step, and will make every effort to answer questions live. However, I may not be able to answer online questions during the course, especially if many questions arise simultaneously.
We will also be available before the course, during the setup sessions, and after the course to help online participants whenever possible. Our goal is for everyone to follow the content and make the most of the experience.
REGISTRATION CLOSED ON 3 JULY.
Participants were selected based on their institution, availability, and candidate profile, with the aim of ensuring diversity, inclusion, and the best possible use of the course. Selected participants received a confirmation email on 17 July.
Please send an email to tobias.ferreira@noc.ac.uk for more information.
| Before | Pre-workshop survey |
| 09:00 | Opening: INPO |
| 09:15 | Course introduction, goals, and setup |
| 09:45 | Data Formats, Metadata, and Vocabulary |
| 10:30 | BREAK |
| 10:50 | Challenges of N-Dimensional Data |
| 11:35 | Cloud-Native Formats |
| 12:00 | Zarr Data Model and Chunked Storage |
| 13:00 | LUNCH |
| 14:00 | Python for Zarr |
| 14:50 | Choosing Chunks at Scale |
| 15:30 | BREAK |
| 15:50 | Parallel Processing with Zarr |
| 17:00 | Close |
| 09:00 | Opening and recap |
| 09:10 | Reading Real-World Zarr Datasets in Python |
| 10:30 | BREAK |
| 10:50 | Object Storage and Cloud Data Organization |
| 11:35 | Converting Traditional Formats to Zarr |
| 13:00 | LUNCH |
| 14:00 | Converting Traditional Formats to Zarr (continued) |
| 15:10 | Versioning Data with Icechunk |
| 15:30 | BREAK |
| 15:50 | Versioning Data with Icechunk (continued) |
| 17:00 | Close |
| 09:00 | Opening and recap |
| 09:30 | Lightning talks: How are other institutions using Zarr? |
| 10:40 | BREAK |
| 11:00 | Virtual Zarr with Virtualizarr |
| 11:50 | Organizing Cloud Zarr Data with STAC |
| 12:40 | Visualizing Multiscale Zarr and GeoZarr |
| 13:00 | LUNCH |
| 14:00 | Visualizing Multiscale Zarr and GeoZarr (continued) |
| 15:30 | BREAK |
| 15:50 | Architecture and Best Practices |
| 16:40 | Closing and post-workshop survey |
| 17:00 | End |
Find quick answers here about participation, preparation, and how the course will run.
No. The course will introduce the concepts of cloud computing and cloud-native formats from the beginning. However, we recommend familiarity with Python, especially functions, loops, NumPy, and xarray, as well as basic experience using the terminal and working with scientific data in formats such as NetCDF or GRIB.
Approximately 95% of the course will be taught in Portuguese, including the lectures, explanations, demonstrations, and most discussions with participants.
On the third day, we will have a special session of approximately one hour with presentations from professionals at UK institutions who have experience using Zarr and other technologies related to cloud-native data. These presentations will be delivered in English.
All course materials are written in English. This decision was made because the materials will be open and reused in future editions of the course, including activities aimed at the UK and European communities.
Therefore, fluency in English is not required to follow most of the course, but familiarity with reading technical documentation in English is recommended.
In-person participants will take part in the activities at INPO in Rio de Janeiro and will use JASMIN computing resources during the course. Online participants will follow the course remotely via Zoom and will use their own computers, with their environment configured in advance according to the setup instructions provided.
The content presented will be the same for both participation formats. However, because the course will include many practical activities, live coding, and direct interaction with participants attending in person, the experience and level of support available to online participants will be different.
Online participants will follow the course live via Zoom during all three days. A significant part of the course will be delivered through live coding, allowing participants to follow the examples, commands, and workflows demonstrated by the instructor step by step.
We would, however, like to be transparent about the format. Because the in-person activities will involve considerable interaction, exercises, and individual support, unfortunately we will not be able to provide online participants with the same level of real-time support available to those attending in person.
Questions can be submitted through the Zoom chat and through a collaborative document that will be used during the course. We will do our best to answer questions from online participants, but depending on the number of questions and the in-person activities taking place, some questions may not be answered immediately during the session.
We will also provide support before the course through the environment setup sessions, and we intend to maintain a communication channel after the course so that participants can continue sharing questions and experiences.
It depends on how you are participating. In-person participants will mainly work on JASMIN during the course and will need to create their account by following the instructions sent by email.
Online participants will use their own computers and therefore need to configure and test their local environment in advance by following the setup materials, as well as downloading the example datasets before the first day of the course.
We recommend reviewing the introductory material on the shell, Git, and working with multidimensional data using xarray, available in the course preparation materials.
We will also hold optional online support sessions to help participants configure their environment. During these sessions, we will demonstrate the installation process and will be available to help resolve any issues before the course begins.
Yes. All teaching materials were developed specifically for this course, including the theoretical explanations, examples, notebooks, exercises, and practical activities.
The content has been organised to progressively introduce concepts and tools related to cloud-native architectures and modern formats for geoscientific data, combining theoretical foundations with practical examples.
All course materials will be made openly and freely available on a dedicated GitHub page. This includes presentations, notebooks, examples, code, exercises, and, whenever possible, the datasets used during the activities.
The link to the materials will be shared with participants before the course begins and will remain available after the course has finished.
Unfortunately, the course will not be recorded. Recording the sessions would significantly increase the complexity and logistics involved in running the course, particularly for the instructor, who will be focused on delivering the content, live coding, and supporting participants.
However, all materials used during the course will be made openly and freely available. Participants will therefore be able to access and reuse the notebooks, examples, exercises, code, and datasets used during the activities after the course.
Before the course: communication will take place primarily by email, using the contact details provided on this page. Important information about preparation, access to the course, and any updates will be sent through this channel.
During the course: for online participants, communication will mainly take place through Zoom and the platform's chat. We will also have an editable collaborative document that participants and the instructor can use to share questions, comments, links, references, and other useful information throughout the course.
After the course: we intend to create a permanent communication space, possibly through an email group or a Slack channel. The aim is to allow participants to continue exchanging experiences, questions, materials, and examples of how they are using the technologies introduced during the course.
We also intend to invite professionals and researchers from the United Kingdom to participate in this space, helping to strengthen links between the Brazilian and UK communities working with geoscientific data and cloud-native technologies. Further details about this initiative will be provided during the course.
Photos and other records of activities may be taken during the course.
These images may be used to publicize and document the course institutionally, including in reports, websites, social media, and presentations.
We are committed to making this course accessible to everyone.
We are dedicated to providing a positive and accessible learning environment for everyone. We do not require participants to provide documentation of disabilities or disclose any unnecessary personal information. However, we want to help create an inclusive and accessible experience for all participants. We encourage you to share any information that may be helpful for making your Carpentries experience accessible.
The Carpentries project comprises a community of instructors, helpers, maintainers, and others. Each of these roles has specific responsibilities to ensure the course is a positive experience for all participants.
To learn more about course roles (who will do what), see our Workshop FAQ.
All participants in Carpentries activities must follow the Code of Conduct. This document also describes how to report an incident, if necessary.