Dynamic Workflows and Federated Deep Learning

Part 2: HPC Algorithms for Science & Tech

This course introduces dynamic workflows and federated learning in the context of High Performance Computing. Students first explore the principles of workflow design, learning how to build, execute, and optimize complex data analysis pipelines using tools such as Dask and the Common Workflow Language (CWL).
Emphasis is placed on task parallelism, reproducibility, and efficient data handling, with practical exercises on normalization, dimensionality reduction, and clustering workflows. The course then shifts to federated learning, where participants implement
client–server architectures for distributed training, develop aggregation strategies, and apply algorithms such as FedAvg. Advanced topics include task-based parallelization of federated training and the integration of workflow systems to manage distributed experiments. Through hands-on laboratory sessions, students gain the skills to design dynamic workflows,
orchestrate machine learning pipelines, and implement federated learning approaches that combine scalability, reproducibility, and efficiency.