The more, the better? Evaluating the role of EEG preprocessing for deep learning applications (2024)
Open in webOpen in zoteroOpen pdf
1 Abstract
The last decade has witnessed a notable surge in deep learning applications for the analysis of electroencephalography (EEG) data, thanks to its demonstrated superiority over conventional statistical techniques. However, even deep learning models can underperform if trained with bad processed data. While preprocessing is essential to the analysis of EEG data, there is a need of research examining its precise impact on model performance. This causes uncertainty about whether and to what extent EEG data should be preprocessed in a deep learning scenario. This study aims at investigating the role of EEG preprocessing in deep learning applications, drafting guidelines for future research. It evaluates the impact of different levels of preprocessing, from raw and minimally filtered data to complex pipelines with automated artifact removal algorithms. Six classification tasks (eye blinking, motor imagery, Parkinson’s and Alzheimer’s disease, sleep deprivation, and first episode psychosis) and four different architectures commonly used in the EEG domain were considered for the evaluation. The analysis of 4800 different trainings revealed statistical differences between the preprocessing pipelines at the intra-task level, for each of the investigated models, and at the inter-task level, for the largest one. Raw data generally leads to underperforming models, always ranking last in averaged score. In addition, models seem to benefit more from minimal pipelines without artifact handling methods, suggesting that EEG artifacts may contribute to the performance of deep neural networks.
2 NOTES
In this paper the authors evaluate four different preprocessing pipelines for deep learning application. The pipelines are the ones shown in the table bellow, and the models are the EEGNet, ShallowNet, DeepConvNet and FBCNet. They evaluated it in a nested leave-N-subjects out cross-evaluation on many paradigms (Eye, MMI, Parkinson, Alzheimer, Sleep and FEP), looked at the accuracy but most important, they used the Frieman’s test to rank the pipelines and the Nemenyi’s test to compare each one against the others. Results showed that for all models the Filt pipeline had the best overall results, and Raw tended to be the worst.
