Synthetic EEG generation using diffusion models for motor imagery tasks (2025)
Open in webOpen in zoteroOpen pdf
1 Abstract
Electroencephalography (EEG) is a widely used, non-invasive method for capturing brain activity, and is particularly relevant for applications in Brain-Computer Interfaces (BCI). However, collecting high-quality EEG data remains a major challenge due to sensor costs, acquisition time, and inter-subject variability. To address these limitations, this study proposes a methodology for generating synthetic EEG signals associated with motor imagery brain tasks using Diffusion Probabilistic Models (DDPM). The approach involves preprocessing real EEG data, training a diffusion model to reconstruct EEG channels from noise, and evaluating the quality of the generated signals through both signal-level and task-level metrics. For validation, we employed classifiers such as K-Nearest Neighbors (KNN), Convolutional Neural Networks (CNN), and U-Net to compare the performance of synthetic data against real data in classification tasks. The generated data achieved classification accuracies above 95%, with low mean squared error and high correlation with real signals. Our results demonstrate that synthetic EEG signals produced by diffusion models can effectively complement datasets, improving classification performance in EEG-based BCIs and addressing data scarcity.
2 NOTES
In this paper the authors propose to use Diffusion Probabilistic Models (DDPM) to generate synthetic EEG data. However, they opted to generate data for only a couple of channels, and even more interesting, only for the non relevant channels in motor imager (Fp1, Fp2, AF3, AF4, F7, F8, T7, and T8). At least that is what I understood, because they say the reconstruct them based on adjacent channels. This is interesting, maybe you don’t need a single model for the whole thing, maybe a few for specific channels are more important, since then you can pass to them only the data from a couple channels, this way they are gonna be more specialized, instead of trying to be a general model. In theory this is interesting, but the way they did is simply weird. Furthermore, they don’t test the data against EEGNet or any EEG specific model, instead the chose Linear Regression and KNN. I don’t know, they did plenty of experiments, their results are on par with a cWGAN-GP, but not better than simply using the original data. They should also have used TT-TS or TS-TT (train on test-test on train, and etc). It is a good initial paper.