Synthetic Data from Domain Knowledge: Pretraining Medical Deep Networks under Data Scarcity
- Date
- May 18 (Mon) 14:00 - 15:00, 2026 (JST)
- Speaker
-
- Naoki Nonaka (Senior Research Scientist, Medical Science Deep Learning Team, Division of Applied Mathematical Science, RIKEN Center for Interdisciplinary Theoretical and Mathematical Sciences (iTHEMS))
- Venue
- Seminar Room #359 (Main Venue)
- via Zoom
- Language
- English
- Host
- Naoki Nonaka
Training deep learning models typically requires large-scale data, yet in the medical domain such data are often difficult to obtain due to privacy constraints, the rarity of certain diseases, and the high cost of acquisition. In this talk, I present one approach to this challenge: pretraining with synthetic data generated from domain knowledge. As concrete examples, I introduce the synthesis of electrocardiograms (ECG) and phonocardiograms (PCG). For ECG, each waveform component (P, Q, R, S, and T) is modeled with Gaussian functions; for PCG, synthetic signals are generated by combining S1 and S2 heart sounds with modulated noise. I show that pretraining a model on such synthetic data and then fine-tuning on a small amount of real data substantially improves classification performance compared to training on real data alone, and that this improvement becomes more pronounced as the size of the real dataset decreases. I will also touch on extensions such as self-supervised learning with synthetic data and a comparison between knowledge-driven simulators and learned generative models, and discuss the broader potential of domain knowledge as a data source for medical applications where real data are limited.
This is a closed event for scientists. Non-scientists are not allowed to attend. If you are not a member or related person and would like to attend, please contact us using the inquiry form. Please note that the event organizer or speaker must authorize your request to attend.