EEG-GAN: A generative EEG augmentation toolkit for enhancing neural classification (2025)

Open in webOpen in zoteroOpen pdf

1 Abstract

Electroencephalography (EEG) is a widely applied method for decoding neural activity, offering insights into cognitive function and driving advancements in neurotechnology. However, decoding EEG data remains challenging, as classification algorithms typically require large datasets that are expensive and time-consuming to collect. Recent advances in generative artificial intelligence have enabled the creation of realistic synthetic EEG data, yet no method has consistently demonstrated that such synthetic data can lead to improvements in EEG decodability across diverse datasets. Here, we introduce EEG-GAN, an open-source generative adversarial network (GAN) designed to augment EEG data. In the most comprehensive evaluation study to date, we assessed its capacity to generate realistic EEG samples and enhance classification performance across four datasets, five classifiers, and seven sample sizes, while benchmarking it against six established augmentation techniques. We found that EEG-GAN, when trained to generate raw single-trial EEG signals, produced signals that reproduce grand-averaged waveforms and time-frequency patterns of the original data. Furthermore, training classifiers on additional synthetic data improved their ability to decode held-out empirical data. EEG-GAN achieved up to a 16% improvement in decoding accuracy, with enhancements consistent across datasets but varying among classifiers. Data augmentations were particularly effective for smaller sample sizes (30 and below), significantly improving 70% of these classification analyses and only significantly impairing 4% of analyses. Moreover, EEG-GAN significantly outperformed all benchmark techniques in 69% of the comparisons across datasets, classifiers, and sample sizes and was only significantly outperformed in 3% of comparisons. These findings establish EEG-GAN as a robust toolkit for generating realistic EEG data, which can effectively reduce the costs associated with real-world EEG data collection for neural decoding tasks.

2 NOTES

This paper proposes a GAN architecture to generate EEG data. The main difference from usual applications is that they include an Encoder before the Discriminator, so instead of passing the time-series it passes and encoding of the EEG data. They mention that this structure results in quicker and more stable training as it decouples feature extraction from generation, simplifying the training process.

Now that I think about it it is actually quite nice, because then the Generator won’t be generating a time-series. Instead, it will generate the equivalent to the encoded features of the time-series, which is obviously simpler. The biggest problem is that the Decoder becomes the clear bottleneck. If it can’t produce the decent samples then it doesn’t matter at all. And most of all, the decoded generated sample is then evaluated where? Maybe it is trained as an usual VAE…

The paper is awfully structured, but it also has another two great points: their code, which is a beautiful framework, and the fact that they used Train-Synthetic-Test-Real (TSTR) and Train-Real-Test-Real (TRTR), the TSTR is an absolutely basic GAN evaluation and I simply can’t stand papers that don’t use it.

Their results are good, but they just compared to the VAE so nothing much can be said about it. Overall, a good paper.