Background: Nodular lymphocyte-predominant Hodgkin lymphoma (NLPHL) is a rare B-cell lymphoma with generally favorable survival but risk of relapse and transformation, and limited evidence for standardized treatment due to small cohorts. Rituximab is a key therapeutic option, but data scarcity restricts robust clinical conclusions. In this context, synthetic data generation is investigated as a strategy to address data scarcity by expanding data availability while preserving the statistical and clinically relevant properties of real-world cohorts, enabling more reliable survival and treatment-response analyses in rare disease settings.
Methods: Real-world anonymized NLPHL patient data derived from a multicenter data collection coordinated by IRCCS Policlinico San Matteo in Pavia were used, including longitudinal clinical, laboratory, treatment, and outcome variables. Synthetic data were generated using a diffusion-based architecture for mixed-type tabular data, with generation conditioned on treatment type (rituximab-based vs conventional therapy). Synthetic data quality was assessed through replication of Kaplan–Meier progression-free survival (PFS) analyses compared with the real cohort, using disease progression as the event of interest and time from treatment initiation to progression as survival time.
Results: The real-world cohort comprised 292 patients, while the synthetic dataset included 1000 patients after post-processing removal of anomalous samples. Preliminary survival analyses compared PFS between real and synthetic cohorts. Kaplan–Meier PFS curves and corresponding confidence intervals were evaluated; log-rank test showed no significant difference between curves (p=0.9).
Conclusion: This study evaluated a conditional diffusion model for synthetic data generation in NLPHL, a rare disease with limited cohort sizes. Conditioning on treatment allowed generation of synthetic cohorts preserving clinically relevant distributions, with utility supported by agreement between real and synthetic Kaplan–Meier survival curves. However, results are limited by the retrospective single-cohort design and small sample size, and synthetic data should be considered complementary to real data. Future work should explore outcome-conditioned generation and multi-center validation to improve time-to-event modeling and robustness.
Gabriele Santangelo, Eleonora Fresi, Virginia Ferretti, Martha Berliner, Sofia Pedrali, Alessandro Mazzacane, Arianna Dagliati, Manuel Gotti