1 MAtt
MAtt (manifold attention network) is a geometric deep-learning EEG decoder that performs self-attention directly on SPD matrices via a Log-Euclidean metric, decoding time-asynchronous (MI) and time-synchronous (SSVEP, ERN) EEG with one architecture (NeurIPS 2022).
Ref: @panMAttManifoldAttention2022
Source code: https://github.com/CECNL/MAtt
Overview
- Problem: existing GDL-for-EEG studies map spatial statistics onto the SPD manifold but cannot map the temporal information of EEG onto the manifold, or still rely on Euclidean tools to handle EEG features.
- Idea: a manifold attention module that treats temporally segmented SPD representations as query/key/value and aggregates them with a softmax-weighted Log-Euclidean mean; parameters live on the Stiefel manifold and are updated with Riemannian gradient descent + QR retraction.
- Contributions: (1) a GDL framework for general EEG decoding; (2) lightweight, interpretable, efficient hybrid Euclidean/Riemannian feature extraction; (3) empirical superiority over leading DL decoders on MI, SSVEP, and ERN; (4) neuroscientific interpretation (contralateral C3/C4, midline CPz, mu-band ~10 Hz, first-epoch-dominant attention).
- Datasets (exactly three): BCI Competition IV Dataset 2a (4-class MI, 9 subjects, 22 ch, 250 Hz), MAMEM-SSVEP-II (5-class SSVEP, 11 subjects, 8 occipital channels), and BCI-ERN (binary imbalanced error-related negativity, 16 subjects, 56 ch).
- Preprocessing: MI: downsampled “from 256 Hz to 128 Hz” (internally inconsistent — 2a is recorded at 250 Hz), band-pass 4–38 Hz, epochs 0.5–4 s after cue onset (438 timepoints). SSVEP: band-pass 1–50 Hz, 8 channels (PO7, PO3, POz, PO4, PO8, O1, Oz, O2), each trial cut into four 1-s segments from 1–5 s. ERN: downsample 600→128 Hz, band-pass 1–40 Hz, 56 ch × 160 timepoints.
- Splits (subject-specific, no cross-subject): MI: session 1 train (1/8 validation), session 2 test; SSVEP: sessions 1–4 train (1/4 validation), session 5 test; ERN: first 4 sessions train (1/4 validation), 5th test. Ten repeats per subject for MI and SSVEP; AUC for ERN.
Architecture
Pipeline: FE → E2R → manifold attention → R2E → FC + softmax.
- Feature extraction (FE): two convolutional layers — a spatial-filtering conv and a spatiotemporal conv — with settings following SCCNet.
- E2R: split embedding into epochs , compute the sample covariance matrix (SCM) of each, then trace-normalize and add with :
yielding the SPD sequence (a TraceNorm-style operation).
3. Manifold attention: per-position BiMap-style projections with (, full row rank so outputs stay SPD):
Similarity uses the Log-Euclidean distance :
and outputs are weighted Log-Euclidean means over values:
- R2E + classifier: ReEig nonlinearity, then LogEig (, upper triangle → ), fully connected layer + softmax; cross-entropy loss.
Optimization: Riemannian gradient descent on the Stiefel manifold for ; Euclidean gradient minus its normal-space projection, then retraction via QR decomposition. Batched Q/K/V computation reduces per-iteration complexity from to a constant.
Model Parameters
- Number of epochs : 3 for MI (2a) and ERN; 7 for SSVEP.
- Attention dims: input SPD ; , ; numerical , , and total parameter count: not reported.
- added on the diagonal of every SCM.
- Batch size, learning rate value, number of training epochs: not reported.
Training Parameters
- Loss: cross-entropy; ERN uses AUC as the evaluation criterion.
- Optimizer: Riemannian gradient descent on the Stiefel manifold (learning-rate value not reported), QR-based retraction.
- Model selection: lowest validation loss within 350 iterations (MI, ), 180 iterations (SSVEP, ), 130 iterations (ERN, ); then test on the held-out session.
- Repeat protocol: mean accuracy over 10 repeats per subject for 2a and MAMEM-SSVEP-II.
- Hardware (timing only): two cores of an Intel Xeon W-2133 CPU.
- Statistics: Wilcoxon signed-rank test with Bonferroni correction across models.
Results
Performance comparison (MI/SSVEP accuracy %, ERN AUC %; mean ± SD):
| Model | MI (BCIC-IV-2a) | SSVEP (MAMEM-II) | ERN (BCI-ERN) |
|---|---|---|---|
| ShallowConvNet | 61.84 ± 6.39 | 56.93 ± 6.97 | 71.86 ± 2.64 |
| EEGNet | 57.43 ± 6.25 | 53.72 ± 7.23 | 74.28 ± 2.47 |
| SCCNet | 71.95 ± 5.05 | 62.11 ± 7.70 | 70.93 ± 2.31 |
| EEG-TCNet | 67.09 ± 4.66 | 55.45 ± 7.66 | 77.05 ± 2.46 |
| TCNet-Fusion | 56.52 ± 3.07 | 45.00 ± 6.45 | 70.46 ± 2.94 |
| FBCNet | 71.45 ± 4.45 | 53.09 ± 5.67 | 60.47 ± 3.06 |
| MBEEGSE | 64.58 ± 6.07 | 56.45 ± 7.27 | 75.46 ± 2.34 |
| MAtt | 74.71 ± 5.01 | 65.50 ± 8.20 | 76.01 ± 2.28 |
MAtt beats all baselines on MI and SSVEP; on ERN it trails EEG-TCNet by ~1 pp.
Ablation (FE = feature extractor, MA = manifold attention, SA = Euclidean self-attention):
| Parts | MI | SSVEP | ERN |
|---|---|---|---|
| FE | 26.08 ± 0.70 | 20.18 ± 1.11 | 73.40 ± 2.27 |
| MA | 60.73 ± 5.80 | 30.51 ± 2.57 | 59.47 ± 3.56 |
| FE + SA | 49.19 ± 2.72 | 22.91 ± 2.00 | 63.77 ± 1.71 |
| FE + MA (MAtt) | 74.71 ± 5.01 | 65.50 ± 8.20 | 76.01 ± 2.28 |
FE+MA significantly outperforms FE+SA on all datasets; every component is necessary.
Training time per iteration (s): MAtt 0.96 ± 0.084 (MI), 2.26 ± 0.160 (SSVEP), 0.52 ± 0.017 (ERN); slowest overall on MI/SSVEP among compared models.
Statistics: MI — MAtt vs all baselines non-significant (smallest p = 0.11). SSVEP — MAtt significantly better than EEG-TCNet (p = 0.03). ERN — EEG-TCNet significantly better than MAtt; MAtt significantly better than FBCNet (p = 0.01).