A multiscale convolutional fusion riemannian manifold learning method for EEG decoding (2025)
Open in webOpen in zoteroOpen pdf
1 Abstract
Deep learning-based methods have been widely adopted for decoding motor imagery (MI) EEG signals. However, due to the low signal-to-noise ratio, non-stationary characteristics, and spatial complexity of EEG data, conventional deep learning approaches operating in Euclidean space often struggle to capture the intrinsic spatial coupling structures of EEG signals. Moreover, the limited size of EEG datasets frequently leads to overfitting. To address these challenges, this study proposes a learning framework that integrates multi-scale convolution with Riemannian manifold learning. Specifically, multiple convolutional kernels of varying sizes are employed to extract detailed features from both low-frequency and high-frequency components of EEG signals. These features are then stacked into tensors to enable comprehensive multi-granularity representation from limited data. The stacked multi-scale features are subsequently projected onto a Riemannian manifold to extract global inter-channel coupling information. In addition, a manifold-based encoder-decoder architecture is employed to promote feature reuse, thereby enhancing the model’s ability to represent nonlinear data structures. The proposed model was evaluated on the BCIC-IV-2a and BCIC-IV-2b datasets, achieving average decoding accuracies of 82.32% and 82.03%, respectively. Experimental results demonstrate that the proposed method outperforms traditional approaches in terms of classification accuracy and robustness.
2 NOTES
Serious problem: the authors do not use either within-session nor within-subject evaluation. Instead, they concatenate all the session and do a k-fold cross-validation (but they do not mention how much folds). This is problematic because it makes the problem much easier, and the comparison against some models (such as EEGNet and FBCNet) are probably not correct, since I don’t think they do it this way.
On the @maneFBCNetMultiviewConvolutional2021 paper they say: “We did not use the inter-session data in the CV analysis to avoid the confounding influence of inter-session variability which is a known problem in the BCI domain.” and “In HO analysis, the complete data from session 1 for the given subject was used for the training purpose, and the resulting model was tested on the session 2 data.”.
Now, about the architecture itself, it seems interesting. The idea of using encoder-decoder for SPDNet is interesting, maybe I will try something like this. The paper is not good thought.
