A cross-session non-stationary attention-based motor imagery classification method with critic-free domain adaptation (2025)
Open in webOpen in zoteroOpen pdf
1 Abstract
Recent studies increasingly employ deep learning to decode electroencephalogram (EEG) signals. While deep learning has improved the performance of motor imagery (MI) classification to some extent, challenges remain due to significant variances in EEG data across sessions and the limitations of convolutional neural networks (CNNs). EEG signals are inherently non-stationary, traditional multi-head attention typically uses normalization methods to reduce non-stationarity and improve performance. However, non-stationary factors are crucial inherent properties of EEG signals and provide valuable guidance for decoding temporal dependencies in EEG signals. In this paper, we introduce a novel CNN combined with the Non-stationary Attention (NSA) and Critic-free Domain Adaptation Network (NSDANet), tailored for decoding MI signals. This network starts with temporal–spatial convolution devised to extract spatial–temporal features from EEG signals. It then obtains multi-modal information from average and variance perspectives. We devise a new self-attention module, the Non-stationary Attention (NSA), to capture the non-stationary temporal dependencies of MI-EEG signals. Furthermore, to align feature distributions between the source and target domains, we propose a critic-free domain adaptation network that uses the Nuclear-norm Wasserstein discrepancy (NWD) to minimize the interdomain differences. NWD complements the original classifier by acting as a critic without a gradient penalty. This integration leverages discriminative information for feature alignment, thus enhancing EEG decoding performance. We conducted extensive cross-session experiments on both BCIC IV 2a and BCIC IV 2b dataset. Results demonstrate that the proposed method outperforms some existing approaches.
2 NOTES
This paper is a direct improvement over the EEG-TransNet model. The main difference is that the Self-Attention Module has some modifications, being now called Non-Stationary Attention (NSA) as the authors added an MLP-Projector which is supposed to handle the non-stationarity of EEG signals, since according to them the usual Attention Models only handle stationary data. This MLP-Projector consists of an one-dimensional convolution and a MLP Layer, where the convolution is applied over the signal and then non-stationary factors are obtained with and . These are used to rescale the data, as shown in the second figure bellow, suchat that the attention becomes . The seconds main difference is the use of a critic-free (since the classifier is also the critic) domain adaptor which uses the Nuclear-norm Wasserstein distance, that should align the features between the source and target domain using adversarial training. Their results are very good, showing higher accuracy than the other literature models. One interesting part is the ablation results, in which they show that the domain adaptation was the most important method on their architecture, followed by data augmentation (simply splitting the segment into 8 segments and randomly re-combining it, which they took from the EEG-TransNet paper). Great explanation and figures, but do require some deeper knowledge on domain adaptation.



