EEG Conformer: Convolutional Transformer for EEG Decoding and Visualization (2022)
Open in webOpen in zoteroOpen pdf
1 Abstract
Due to the limited perceptual field, convolutional neural networks (CNN) only extract local temporal features and may fail to capture long-term dependencies for EEG decoding. In this paper, we propose a compact Convolutional Transformer, named EEG Conformer, to encapsulate local and global features in a unified EEG classification framework. Specifically, the convolution module learns the low-level local features throughout the one-dimensional temporal and spatial convolution layers. The self-attention module is straightforwardly connected to extract the global correlation within the local temporal features. Subsequently, the simple classifier module based on fully-connected layers is followed to predict the categories for EEG signals. To enhance interpretability, we also devise a visualization strategy to project the class activation mapping onto the brain topography. Finally, we have conducted extensive experiments to evaluate our method on three public datasets in EEG-based motor imagery and emotion recognition paradigms. The experimental results show that our method achieves state-of-the-art performance and has great potential to be a new baseline for general EEG decoding. The code has been released in https://github.com/eeyhsong/EEG-Conformer.
2 NOTES
This paper has one of the first transformers network applied to EEG data. It’s first layers are built to be like in the EEGNet, first a temporal then a spatial convolution, followed by an average pooling. However, from the pooling onward the architecture is totally different, first, the pooling is meant to reduce the size of the features, so that the token, created by taking a row of the output of the pooling, is of the right size. If it is the kernel is too big then the temporal features are too smoothed and lose useful details, and if too small then the self-attention’s performance may easily be affected by local noise. After evaluating they chose (1,75) for the kernel size, which I believe is the same as in the EEGNet. After this it is then applied self-attention layers on the token, each with heads. After some experiments they set and . This is essentially how their architecture is built. However, something of high interest is that they used data-augmentation, in particular, the same segmentation and reconstruction (S&R) as that used in many other papers related to EEG and Transformers. They also mention developing a gradient based activation algorithm to visualize the impact of the transformer on the class activation over the topography, but don’t seem that interesting.
