日時
2026年5月18日(月)14:00 - 15:00 (JST)
講演者
  • 野中 尚輝 (理化学研究所 数理創造研究センター (iTHEMS) 数理展開部門 医科学深層学習チーム 上級研究員)
言語
英語
ホスト
Naoki Nonaka

Training deep learning models typically requires large-scale data, yet in the medical domain such data are often difficult to obtain due to privacy constraints, the rarity of certain diseases, and the high cost of acquisition. In this talk, I present one approach to this challenge: pretraining with synthetic data generated from domain knowledge. As concrete examples, I introduce the synthesis of electrocardiograms (ECG) and phonocardiograms (PCG). For ECG, each waveform component (P, Q, R, S, and T) is modeled with Gaussian functions; for PCG, synthetic signals are generated by combining S1 and S2 heart sounds with modulated noise. I show that pretraining a model on such synthetic data and then fine-tuning on a small amount of real data substantially improves classification performance compared to training on real data alone, and that this improvement becomes more pronounced as the size of the real dataset decreases. I will also touch on extensions such as self-supervised learning with synthetic data and a comparison between knowledge-driven simulators and learned generative models, and discuss the broader potential of domain knowledge as a data source for medical applications where real data are limited.

このイベントは研究者向けのクローズドイベントです。一般の方はご参加頂けません。メンバーや関係者以外の方で参加ご希望の方は、フォームよりお問い合わせ下さい。講演者やホストの意向により、ご参加頂けない場合もありますので、ご了承下さい。

このイベントについて問い合わせる