1 MAtt

MAtt (manifold attention network) is a geometric deep-learning EEG decoder that performs self-attention directly on SPD matrices via a Log-Euclidean metric, decoding time-asynchronous (MI) and time-synchronous (SSVEP, ERN) EEG with one architecture (NeurIPS 2022).

Ref: @panMAttManifoldAttention2022

Source code: https://github.com/CECNL/MAtt

Overview

  • Problem: existing GDL-for-EEG studies map spatial statistics onto the SPD manifold but cannot map the temporal information of EEG onto the manifold, or still rely on Euclidean tools to handle EEG features.
  • Idea: a manifold attention module that treats temporally segmented SPD representations as query/key/value and aggregates them with a softmax-weighted Log-Euclidean mean; parameters live on the Stiefel manifold and are updated with Riemannian gradient descent + QR retraction.
  • Contributions: (1) a GDL framework for general EEG decoding; (2) lightweight, interpretable, efficient hybrid Euclidean/Riemannian feature extraction; (3) empirical superiority over leading DL decoders on MI, SSVEP, and ERN; (4) neuroscientific interpretation (contralateral C3/C4, midline CPz, mu-band ~10 Hz, first-epoch-dominant attention).
  • Datasets (exactly three): BCI Competition IV Dataset 2a (4-class MI, 9 subjects, 22 ch, 250 Hz), MAMEM-SSVEP-II (5-class SSVEP, 11 subjects, 8 occipital channels), and BCI-ERN (binary imbalanced error-related negativity, 16 subjects, 56 ch).
  • Preprocessing: MI: downsampled “from 256 Hz to 128 Hz” (internally inconsistent — 2a is recorded at 250 Hz), band-pass 4–38 Hz, epochs 0.5–4 s after cue onset (438 timepoints). SSVEP: band-pass 1–50 Hz, 8 channels (PO7, PO3, POz, PO4, PO8, O1, Oz, O2), each trial cut into four 1-s segments from 1–5 s. ERN: downsample 600→128 Hz, band-pass 1–40 Hz, 56 ch × 160 timepoints.
  • Splits (subject-specific, no cross-subject): MI: session 1 train (1/8 validation), session 2 test; SSVEP: sessions 1–4 train (1/4 validation), session 5 test; ERN: first 4 sessions train (1/4 validation), 5th test. Ten repeats per subject for MI and SSVEP; AUC for ERN.

Architecture

Pipeline: FE → E2R → manifold attention → R2E → FC + softmax.

  1. Feature extraction (FE): two convolutional layers — a spatial-filtering conv and a spatiotemporal conv — with settings following SCCNet.
  2. E2R: split embedding into epochs , compute the sample covariance matrix (SCM) of each, then trace-normalize and add with :

yielding the SPD sequence (a TraceNorm-style operation).
3. Manifold attention: per-position BiMap-style projections with (, full row rank so outputs stay SPD):

Similarity uses the Log-Euclidean distance :

and outputs are weighted Log-Euclidean means over values:

  1. R2E + classifier: ReEig nonlinearity, then LogEig (, upper triangle → ), fully connected layer + softmax; cross-entropy loss.

Optimization: Riemannian gradient descent on the Stiefel manifold for ; Euclidean gradient minus its normal-space projection, then retraction via QR decomposition. Batched Q/K/V computation reduces per-iteration complexity from to a constant.

Model Parameters

  • Number of epochs : 3 for MI (2a) and ERN; 7 for SSVEP.
  • Attention dims: input SPD ; , ; numerical , , and total parameter count: not reported.
  • added on the diagonal of every SCM.
  • Batch size, learning rate value, number of training epochs: not reported.

Training Parameters

  • Loss: cross-entropy; ERN uses AUC as the evaluation criterion.
  • Optimizer: Riemannian gradient descent on the Stiefel manifold (learning-rate value not reported), QR-based retraction.
  • Model selection: lowest validation loss within 350 iterations (MI, ), 180 iterations (SSVEP, ), 130 iterations (ERN, ); then test on the held-out session.
  • Repeat protocol: mean accuracy over 10 repeats per subject for 2a and MAMEM-SSVEP-II.
  • Hardware (timing only): two cores of an Intel Xeon W-2133 CPU.
  • Statistics: Wilcoxon signed-rank test with Bonferroni correction across models.

Results

Performance comparison (MI/SSVEP accuracy %, ERN AUC %; mean ± SD):

ModelMI (BCIC-IV-2a)SSVEP (MAMEM-II)ERN (BCI-ERN)
ShallowConvNet61.84 ± 6.3956.93 ± 6.9771.86 ± 2.64
EEGNet57.43 ± 6.2553.72 ± 7.2374.28 ± 2.47
SCCNet71.95 ± 5.0562.11 ± 7.7070.93 ± 2.31
EEG-TCNet67.09 ± 4.6655.45 ± 7.6677.05 ± 2.46
TCNet-Fusion56.52 ± 3.0745.00 ± 6.4570.46 ± 2.94
FBCNet71.45 ± 4.4553.09 ± 5.6760.47 ± 3.06
MBEEGSE64.58 ± 6.0756.45 ± 7.2775.46 ± 2.34
MAtt74.71 ± 5.0165.50 ± 8.2076.01 ± 2.28

MAtt beats all baselines on MI and SSVEP; on ERN it trails EEG-TCNet by ~1 pp.

Ablation (FE = feature extractor, MA = manifold attention, SA = Euclidean self-attention):

PartsMISSVEPERN
FE26.08 ± 0.7020.18 ± 1.1173.40 ± 2.27
MA60.73 ± 5.8030.51 ± 2.5759.47 ± 3.56
FE + SA49.19 ± 2.7222.91 ± 2.0063.77 ± 1.71
FE + MA (MAtt)74.71 ± 5.0165.50 ± 8.2076.01 ± 2.28

FE+MA significantly outperforms FE+SA on all datasets; every component is necessary.

Training time per iteration (s): MAtt 0.96 ± 0.084 (MI), 2.26 ± 0.160 (SSVEP), 0.52 ± 0.017 (ERN); slowest overall on MI/SSVEP among compared models.

Statistics: MI — MAtt vs all baselines non-significant (smallest p = 0.11). SSVEP — MAtt significantly better than EEG-TCNet (p = 0.03). ERN — EEG-TCNet significantly better than MAtt; MAtt significantly better than FBCNet (p = 0.01).

Code

https://github.com/CECNL/MAtt