Skip to main content

SYNTHEMA | AI-Driven Synthetic Health Data for Haematology

From Synthetic Data Generation to Reusable Research Infrastructure

As SYNTHEMA approaches its final stage, Andoni Beristain Iraola and Imanol Isasa Reinoso from Vicomtech reflect on the technical advances developed within the project and how they could support future research well beyond rare haematological diseases.

Rare disease research faces a fundamental challenge. Progress depends on access to sufficiently large and diverse datasets, yet patient populations are small and their data are often distributed across different hospitals and countries.

Within SYNTHEMA, Vicomtech has led Work Package 3, focused on the development of advanced anonymisation and synthetic data generation pipelines, with all other SYNTHEMA partners contributing to it. Working closely with technical and clinical partners, the team has addressed challenges spanning machine learning, privacy, federated learning and the generation of complex multimodal health data.

From unrestricted generation to governed synthetic data

One of the main advances has been the development of a two-step validation and filtering pipeline for synthetic data.

Generative models can potentially produce an unlimited number of samples, making it difficult to maintain strict privacy guarantees when data are generated dynamically. SYNTHEMA therefore moved away from unconstrained, on-demand generation towards a governed model acceptance process.

Following federated training, a large pool of synthetic data is generated and assessed for statistical fidelity, clinical utility and privacy risks. This process also includes expert clinical review, adding a human-in-the-loop component before synthetic datasets are approved for use.

Once validated, these datasets can be securely queried through GraphQL filtering rather than generated again from scratch. This approach reduces the risk of uncontrolled changes in generated data and helps ensure that distributed datasets remain within the privacy conditions established during validation.

Generating synthetic data in extremely small cohorts

The characteristics of rare haematological diseases create an additional technical challenge.

Conditions such as Sickle Cell Disease and Acute Myeloid Leukaemia often involve datasets with many complex variables but relatively few patients. These “wide and low-sample” datasets are particularly difficult for conventional generative approaches.

To address this, SYNTHEMA adapted several advanced synthetic data generation methods to horizontal federated learning environments. These include federated CTGAN, Variational Autoencoders combined with Bayesian Gaussian Mixture Models, federated Bayesian Networks and Conditional Flow Matching.

Using the Flower federated learning ecosystem, participating institutions can collaboratively train these models without transferring raw patient data outside their own infrastructures.

Extending synthetic data beyond clinical tables

The work also goes beyond structured clinical records.

For Sickle Cell Disease, the project has developed lesion-aware, mask-conditioned diffusion models for brain MRI synthesis, designed to reproduce Silent Cerebral Infarcts.

For Acute Myeloid Leukaemia, cytology and histopathology images can be generated through a two-stage VAE-Diff approach, separating tissue architecture from finer textural characteristics.

Another important area is multimodal synthetic data generation. Using the Partial Information Decomposition framework, SYNTHEMA has worked to preserve meaningful relationships between clinical variables, genomic profiles and imaging features. The objective is to generate synthetic patients whose different data modalities remain biologically coherent and preserve relevant survival patterns.

Adding causal modelling and federated anonymisation

Synthetic data generation is only one part of the infrastructure developed within Work Package 3.

The project has also implemented causal inference architectures, including SA-TEDVAE, to support counterfactual treatment-effect estimation. These models can be used to investigate “what-if” scenarios, such as estimating how bone marrow transplantation could affect patient survival.

At the same time, standalone and federated anonymisation tools, including FedAn, developed by Netcompany (leader of Task 3.2 on data anonymisation), provide automated privacy assessments based on measures such as k-anonymity, l-diversity and information loss. These privacy metrics can then be linked to the corresponding data catalogues, one of which is the Anonymized Data Catalogue, developed by Netcompany, along with dataset ingestion and annotation services following the HealthDCAT-AP cataloguing standard to align with the EHDS (European Health Data Space) directions for cross border data sharing and secondary use of health datasets.

Building an infrastructure that can move beyond haematology

A central principle behind the SYNTHEMA architecture has been modularity.

Many of its technical components are not specific to Sickle Cell Disease or Acute Myeloid Leukaemia and could therefore be adapted to other medical fields, including oncology, neurology, rare paediatrics and cardiovascular research.

The first step would be to adapt the underlying data models and metadata. SYNTHEMA relies on standardised structures including OMOP CDM and HealthDCAT-AP, supporting alignment with the European Health Data Space. Moving into another clinical domain would therefore primarily require mapping the relevant medical data into these common standards, while much of the ingestion, storage and catalogue infrastructure could remain unchanged.

The same principle applies to validation. SYNTHEMA’s Synthetic Validation Framework evaluates generated datasets across three main areas: fidelity, privacy and utility. Statistical fidelity measures and privacy tests, including membership inference and attribute disclosure assessments, can be reused across medical domains. The part that requires clinical adaptation is the definition of disease-specific utility endpoints, such as tumour response criteria, event-free survival or organ-specific segmentation.

The generative models themselves are also designed to be portable. Flow-matching approaches for complex tabular data, mask-conditioned diffusion models for medical imaging and multimodal coherence frameworks can be retrained for other clinical domains where data remain scarce and fragmented across institutions.

Finally, the infrastructure has been designed for reproducible deployment, this has occurred in WP2, led by UPM with strong contribution from Netcompany. Its microservices-based Kubernetes architecture and standardised Flower interfaces allow new clinical consortia to deploy local and central components while reusing existing deployment blueprints and CI/CD processes, delivered by Netcompany.

From a technical concept to reusable research infrastructure

The work carried out within SYNTHEMA shows how federated synthetic data generation can move from a theoretical privacy-preserving approach towards a practical research infrastructure.

The project has established a framework for research environments where access to patient data is both limited and highly sensitive. It integrates federated learning, synthetic data generation, anonymisation, multimodal modelling, validation and interoperable data standards within a single technical approach.

The next step is not necessarily to redesign this infrastructure for every new disease. Much of the technical foundation can remain in place, while clinical data models, validation endpoints and training datasets are adapted to the requirements of each new medical domain.