Galaxy Tutorial with Wolfgang Maier

On July 7th IRTG members and other interested FONDA folks attended a tutorial on the workflow and data analysis platform Galaxy with Wolfgang Maier from the University of Freiburg. The tutorial covered fundamentals of Galaxy, and included hands-on exercises in running, editing, and automating workflows. The tutorial highlighted how feature-rich and widely adopted Galaxy is and inspired many in-depth conversations about sustainability and usability along with thoughtful comparisons between Galaxy and other common workflow systems like Snakemake and NextFlow.

The following day, members of subprojects A2, B1, C1, C3, C2 and A7 met with Wolfgang to discuss potential areas for collaboration going forward.

FONDA PhD Defense: Fabian Lehmann on “Adaptive Scheduling of Dynamic Workflows”

Fabian Lehmann defended his dissertation, “Adaptive Scheduling of Dynamic Workflows” with distinction on June 29th, 2026. Fabian was a member of subproject B5 in FONDA Phase I. His work examines the tradeoff between workflow portability and efficiency, and introduces WONDERS, a best of both worlds optimization strategy for Nextflow workflows.

In order for scientific workflows to be portable, workflow management systems such as Nextflow cannot rely on hard-coded assumptions about the dataset or underlying infrastructure to improve efficiency. WONDERS utilizes the Common Workflow Scheduling Interface (CWSI), developed by Fabian, to bring workflow context to the resource manager, then combines his three novel approaches for workflow optimization: WOW, PONDER, and SCALE.

WOW is a scheduling approach which reduces network congestion within a cluster by basing scheduling decisions primarily on the co-location of tasks and data. PONDER and SCALE are tools which predict and minimize the ongoing use of memory (PONDER) and CPU (SCALE) over the course of a workflow execution. The combination of the three techniques in WONDERS results in super-additive improvements in workflow makespan (average reduction ~50% compared to nf-core default settings) when tested on a variety of Nextflow workflows using real-world data from remote sensing and bioinformatics.

Congratulations Fabian!

FONDA PhD Defense: Sebastian Müller on “Metamorphic Testing for Scientific Software Using Geometrically Representable Input Data”

On June 18th, 2026 Sebastian Müller from FONDA Phase I, subproject A3, successfully defended his dissertation “Metamorphic Testing for Scientific Software Using Geometrically Representable Input Data”. His work focused on testing research software, which can become “untestable” with oracle-based testing methods as a result of its complex and exploratory nature.

Sebastian demonstrated that research software that uses geometrically representable input data (such as images, coordinate grids, or volumetric models) can be tested by applying invariant geometric transformations to the input data. If the software is doing what it is supposed to, the output from the transformed data should then be extremely similar to the original output. He also developed GeoMetaMorph, a tool for detecting likely invariant transformations to use in such tests.

Congratulations Sebastian!

Our book is published!

FONDA members, along with over 100 co-authors from 54 organizations in 17 countries, are proud to present Workflow Systems for Large-Scale Scientific Data Analysis!

The book’s 25 chapters are divided into four main areas (1) Introduction, (2) Systems, (3) Applications, and (4) Technologies, and cover the current state of the art in workflow research for scientific data analysis. The book is open access, and digital copies can be downloaded for free. Hard copies can also be ordered (at the same link) for € 42 each.

Book Cover for Workflow systems for large scale scientific data analysis

We would especially like to thank the editors of the book, Rafael Ferreira da Silva, Sean Wilkinson, Marcus Hilbrich, and Ulf Leser. Funding for the first edition, along with much of the underlying research was provided by the German Research Foundation (DFG). The book is published by Berlin Universities Publishing. Finally, thanks to all the co-authors for their contributions and ongoing support!

Links to the individual chapters along with a comprehensive author list for each can be found here: The Workflow Book

November 25th-26th: Workshop on reducing carbon emissions in large-scale computational data analysis

On Tuesday November 25th, in lieu of our normal Lecture Series talks, we will begin our Workshop on reducing carbon emissions in large-scale computational data analysis! The workshop will continue on Wednesday with more speakers and opportunities for discussion.

The workshop was planned in collaboration with the Quantitative Biology Center (QBiC), the bioinformatics core facility from the University of Tübingen. We have invited 14 speakers from around Europe and the UK to discuss current research into the environmental impacts of large scale scientific computing. More information, including the schedule, can be found here.

Snakemake Tutorial with Johannes Köster

FONDA has invited Professor Johannes Köster to give a tutorial on Snakemake, a widely used, python based workflow management system. Snakemake allows users to create scalable, human readable, reproducible workflows for scientific data analysis.

Professor Köster the leader of the Bioinformatics and Computational Oncology group at the Institute for AI in Medicine at the University of Duisburg-Essen, where his work focuses on reproducibility and bioinformatics workflows. He is the author and lead developer of Snakemake.

The full-day tutorial will be on July 02, 2025 starting at 9 am. Please contact Tobias Price if you are interested in attending.

FONDA Integrated Research Training Group Courses

The Integrated Research Training Group (IRTG) has been very active in the past month providing courses and workshops for FONDA’s doctoral researchers. On April 28th and 29th Seqera, the company which supports the open source workflow engine “Nextflow” provided a tutorial using Nextflow to write and evaluate bioinformatics workflows. On May 7th, the IRTG members participated in a training course in Good Scientific Practice, especially geared towards computational sciences.

Future courses include a tutorial on GitLab with scientists from the German Aerospace Center on May 20th and 21st, a course on gender bias awareness on June 17, along with courses on workflow simulator WfCommons, and python-based workflow execution engine “Snakemake”.

PI-Lecture Series Part 8

The final installment of FONDA’s PI-Lecture Series will take place on March 17th from 15:00-17:30 in Adlershof (Humboldt-Kabinett, Rudower Chaussee 25). The following PIs will give talks on their ongoing research:

  • Henning Meyerhenke – Workflow Scheduling and (Other) Graph Algorithms for Parallel & Distributed Systems
  • Thomas Kosch – TBA
  • Ulf Leser – Knowledge Management in Bioinformatics
  • Björn Scheuermann – Modern Web Transport Protocols (online)

We have had a lot of excellent talks over the last few months. The purpose of this lecture series was to introduce all of our new FONDA members to the research areas of the PIs. Based on the quality of questions and conversations, this has been very successful!

I’m looking forward to more conversations about science with everyone at our upcoming spring retreat.

PI Lecture Series Part 5

Our fifth set of PI-Lectures will be Monday Feb 24th starting at 15:00 in Adlershof – Rudower Chaussee 25, 12489 Berlin. We will meet in the Humboldt-Kabinett for talks regarding ongoing research by the following PIs:

  • Nicole Schweikardt – Logic in Computer Science
  • Matthias Weidlich – On Events and Processes
  • Claudia Draxl – From science to data and back
  • Knut Reinert – Hierarchical Interleaved Bloom Filter: Enabling ultrafast, approximate sequence queries

FONDA PhD student Mario Sänger successfully defends his PhD thesis on “Representation Learning for Biomedical Text Mining”

Mario Sänger, a member of the group “Human-computer interaction for Scientific Software”, successfully defended his PhD thesis on November 25, 2024. His work focuses on using representation learning to extract meaningful connections between biomedical entities, such as genes, diseases, proteins, and pharmaceuticals from a corpus of PubMed abstracts, as well as biomedical knowledge bases. In addition to demonstrating the feasibility of this corpus-wide approach, he also benchmarked and tested existing pre-trained language models (PLMs) for sentence-level relation prediction. His results show that additional context from biomedical knowledge databases does not enhance the most robust carefully tuned PLMs.

In FONDA, he collaborated with Prof. Dr. Thomas Kosch, exploring the use of ChatGPT as a tool to support users in designing and implementing scientific workflows.

Congratulations Mario, and all the best!