MAtt: A Manifold Attention Network for EEG Decoding (2022)

Open in webOpen in zoteroOpen pdf

1 Abstract

Recognition of electroencephalographic (EEG) signals highly affect the efficiency of non-invasive brain-computer interfaces (BCIs). While recent advances of deep-learning (DL)-based EEG decoders offer improved performances, the development of geometric learning (GL) has attracted much attention for offering exceptional robustness in decoding noisy EEG data. However, there is a lack of studies on the merged use of deep neural networks (DNNs) and geometric learning for EEG decoding. We herein propose a manifold attention network (mAtt), a novel geometric deep learning (GDL)-based model, featuring a manifold attention mechanism that characterizes spatiotemporal representations of EEG data fully on a Riemannian symmetric positive definite (SPD) manifold. The evaluation of the proposed MAtt on both time-synchronous and -asyncronous EEG datasets suggests its superiority over other leading DL methods for general EEG decoding. Furthermore, analysis of model interpretation reveals the capability of MAtt in capturing informative EEG features and handling the non-stationarity of brain dynamics.

2 NOTES

In this paper the authors propose an attention module for use with SPD matrices. This follows papers that have been re-purposing modules from usual neural networks into the context of Riemannian geometry (or manifolds other than Euclidean). Instead of relying on the linear layer, present in the attention module, the authors instead use the BiMap, for the ReLU they use ReEig. However, the construction of the attention matrix, which is done using the dot product can’t be directly introduced here, since we have SPD matrices. Therefore, the authors propose to use a similarity based on the Log-Euclidean distance between query and key, defined as , resulting in the attention matrix . They also use the Softmax to shrink the range along the row direction, making values in a row have convexity constraint property. The backward pass also has some details, but I think if using something like geoopt it should be sufficient to restrict the matrices to being SPD and using Riemannian Gradient Descent. They achieved accuracy superior to FBCNet (and many other models). Overall it is a very good paper and even has the code in github.