Physics-Informed Attention Temporal Convolutional Network for EEG-Based Motor Imagery Classification (2022)
Open in webOpen in zoteroOpen pdf
1 Abstract
The brain-computer interface (BCI) is a cutting-edge technology that has the potential to change the world. Electroencephalogram (EEG) motor imagery (MI) signal has been used extensively in many BCI applications to assist disabled people, control devices or environments, and even augment human capabilities. However, the limited performance of brain signal decoding is restricting the broad growth of the BCI industry. In this paper, we propose an attention-based temporal convolutional network (ATCNet) for EEG-based motor imagery classification. The ATCNet model utilizes multiple techniques to boost the performance of MI classification with a relatively small number of parameters. ATCNet employs scientific machine learning to design a domain-specific DL model with interpretable and explainable features, multi-head self-attention to highlight the most valuable features in MI-EEG data, temporal convolutional network (TCN) to extract high-level temporal features, and convolutional-based sliding window to augment the MI-EEG data efficiently. The proposed model outperforms the current state-of-the-art techniques in the BCI Competition IV-2a dataset with an accuracy of 85.38% and 70.97% for the subject-dependent and subject-independent modes, respectively.
2 NOTES
This paper presents an architecture for motor imagery classification, where the authors use three main components: a convolutional (CV) block, an attention (AT) block and a temporal convolutional (TC) block. Their use of the CV block is different from the one in the EEGNet because they chose to use 2D convolution instead of separable convolution. Bellow we can see this block, beautiful figure, which shows that the idea is to reduce the dimension of the input as features are extracted. This makes the final layer have shape (, ) which is then reduced using sliding window, so becomes . In their case they set 5 sliding windows, so in size. Seems quite small, but makes sense since the following is an AT block, so it it is better to input multiple smaller tokens instead of a simple larger one. They chose to use only depth () 2 for the attention, 2 head attentions, each with size 8, which seems to be quite few, I expected at least 8… weird, however they did an ablation study showing that 1 or 2 head attentions is much better than 3 or more (8 being the worst). They also did an experiment change and the number of windows, which led to and being chosen. Something that I did not expect was the use of the TC block in the attention layer, but that makes sense. In previous paper I did see the linear layer being replaced by a regular convolution, so it is clear that a temporal one would be more effective in this type of signal.



