FONDA PhD Defense: Fabian Lehmann on “Adaptive Scheduling of Dynamic Workflows”

Fabian Lehmann defended his dissertation, “Adaptive Scheduling of Dynamic Workflows” with distinction on June 29th, 2026. Fabian was a member of subproject B5 in FONDA Phase I. His work examines the tradeoff between workflow portability and efficiency, and introduces WONDERS, a best of both worlds optimization strategy for Nextflow workflows.

In order for scientific workflows to be portable, workflow management systems such as Nextflow cannot rely on hard-coded assumptions about the dataset or underlying infrastructure to improve efficiency. WONDERS utilizes the Common Workflow Scheduling Interface (CWSI), developed by Fabian, to bring workflow context to the resource manager, then combines his three novel approaches for workflow optimization: WOW, PONDER, and SCALE.

WOW is a scheduling approach which reduces network congestion within a cluster by basing scheduling decisions primarily on the co-location of tasks and data. PONDER and SCALE are tools which predict and minimize the ongoing use of memory (PONDER) and CPU (SCALE) over the course of a workflow execution. The combination of the three techniques in WONDERS results in super-additive improvements in workflow makespan (average reduction ~50% compared to nf-core default settings) when tested on a variety of Nextflow workflows using real-world data from remote sensing and bioinformatics.

Congratulations Fabian!

FONDA PhD Defense: Sebastian Müller on “Metamorphic Testing for Scientific Software Using Geometrically Representable Input Data”

On June 18th, 2026 Sebastian Müller from FONDA Phase I, subproject A3, successfully defended his dissertation “Metamorphic Testing for Scientific Software Using Geometrically Representable Input Data”. His work focused on testing research software, which can become “untestable” with oracle-based testing methods as a result of its complex and exploratory nature.

Sebastian demonstrated that research software that uses geometrically representable input data (such as images, coordinate grids, or volumetric models) can be tested by applying invariant geometric transformations to the input data. If the software is doing what it is supposed to, the output from the transformed data should then be extremely similar to the original output. He also developed GeoMetaMorph, a tool for detecting likely invariant transformations to use in such tests.

Congratulations Sebastian!

Our book is published!

FONDA members, along with over 100 co-authors from 54 organizations in 17 countries, are proud to present Workflow Systems for Large-Scale Scientific Data Analysis!

The book’s 25 chapters are divided into four main areas (1) Introduction, (2) Systems, (3) Applications, and (4) Technologies, and cover the current state of the art in workflow research for scientific data analysis. The book is open access, and digital copies can be downloaded for free. Hard copies can also be ordered (at the same link) for € 42 each.

Book Cover for Workflow systems for large scale scientific data analysis

We would especially like to thank the editors of the book, Rafael Ferreira da Silva, Sean Wilkinson, Marcus Hilbrich, and Ulf Leser. Funding for the first edition, along with much of the underlying research was provided by the German Research Foundation (DFG). The book is published by Berlin Universities Publishing. Finally, thanks to all the co-authors for their contributions and ongoing support!

Links to the individual chapters along with a comprehensive author list for each can be found here: The Workflow Book

FONDA Winter Lecture Series Continues into 2026!

On January 20th at 15:00, FONDA will continue our Winter Lecture Series. We will meet at HU-Main Campus, Unter den Linden 6, Room 2249a for three exciting talks by:

  • Laura Koesten – University of Vienna: A Human-Centered Perspective on Data-Centric Sensemaking
  • Gabin An – Korea University: Finding Bugs at Scale: LLM-Based Fault Localization in Real-World Software
  • Karsten Peters-von Gehlen – Deutsches Klimarechenzentrum: Bridging Petabyte-Scale Climate Data and Analysis Workflows with FAIR Digital Objects

The complete schedule can be found here: FONDA Winter Lecture Series and you can also follow the lectures online.

We are looking forward to seeing you there!

November 25th-26th: Workshop on reducing carbon emissions in large-scale computational data analysis

On Tuesday November 25th, in lieu of our normal Lecture Series talks, we will begin our Workshop on reducing carbon emissions in large-scale computational data analysis! The workshop will continue on Wednesday with more speakers and opportunities for discussion.

The workshop was planned in collaboration with the Quantitative Biology Center (QBiC), the bioinformatics core facility from the University of Tübingen. We have invited 14 speakers from around Europe and the UK to discuss current research into the environmental impacts of large scale scientific computing. More information, including the schedule, can be found here.

FONDA Winter 2025/2026 Lecture Series Begins Oct. 28

On selected Tuesday afternoons from October 28th 2025 to mid-February 2026, FONDA will host 2-3 short scientific talks on a variety of topics related to large scale data analysis and workflows in natural science. The complete schedule (subject to change) can be found here: FONDA Winter Lecture Series.

On October 28th, we will meet in the Humboldt-Kabinett (first floor seminar room of Rudower Chaussee 25, 12489 Berlin-Adlershof) at 15:15 for the following two talks:

  • Nikos Tsakiridis – University of Thessaloniki: From Petabytes to Pedons: Cloud-Native Earth Analytics for Soil Mapping
  • Matthes Rieke – 52 Degrees North: Biodiversity monitoring with openEO – a look at scalability and reproducibility

You can also follow our lecture series online.

We are looking forward to seeing you there!

Data Aware Scheduling Method Now Available for Nextflow

With the latest release of the nf-cws Nextflow-Plugin, Fabian Lehmann and Friedrich Tschirpke introduce the WOW scheduling method as a production-ready feature. This release marks the transition of the WOW approach from a research prototype to a usable software component, fully integrated with the official Nextflow versions (v24.04.0 up to v25.02.3-edge).

Key Features:

  • WOW Scheduling for Nextflow: The Workflow-Aware data movement and task scheduling (WOW) method is now available as part of nf-cws. This enables dynamic coordination of data transfers and task execution, reducing network congestion and workflow runtime.
  • Seamless Integration: The nf-cws plugin can be used directly with Nextflow’s Kubernetes executor, requiring no experimental patches or custom setups.
  • Production Use: The improvements demonstrated in the original publication can now be leveraged by all Nextflow users in real-world scenarios.

-Fabian Lehmann

PI-Lecture Series part 2

Our second set of PI-Lectures will be today at 15:00 at Einstein Center Digital Future. The following PIs will be presenting their research areas:

  • Patrick Hostert – Satellite Remote Sensing
  • Matthias Boehm – System Infrastructure for Data-centric ML Pipelines
  • Tillman Rabl – Carbon-efficient Data Systems
  • Odej Kao – LLMOps for Reliability and Availability of Massive AI Infrastructures

FONDA contributed the first non-Bioinformatics Workflow to nf-core

We have successfully ported and contributed our Rangeland workflow [1], [2] to the nf-core workflow repository. This milestone marks the release of the first non-bioinformatics workflow on nf-core, which now serves as a blueprint for workflows in remote sensing and other domains. Being a part of the nf-core community confirms that our workflow uses best practices and ensures accessibility to researchers worldwide.


The Rangeland workflow analyzes trends in grassland changes, providing valuable insights for environmental research.

This achievement was made possible by subproject B5: Felix Kummer, Katarzyna Ewa Lewińska, Fabian Lehmann, and David Frantz.

-Fabian Lehmann