1 SPD-Mani-Net
SPD-Mani-Net is a Siamese manifold-to-manifold network that learns discriminative low-dimensional SPD covariance representations for motor-imagery BCI classification.
Date of publication: 23/03/2023
Ref: @pengReducingDimensionalitySPD2023
Code at: https://github.com/pxxyyz/SPD-Manifold-Network
Preprocessing code referenced by the paper: https://sites.google.com/site/fabienlotte/research/code-and-softwares
Overview
The model addresses the quadratic growth of covariance dimensionality with the number of EEG channels and the instability of high-dimensional covariance estimates under small-sample conditions. An EEG trial with channels and samples is represented by
SPD-Mani-Net maps to a more discriminative matrix in , , while preserving positive definiteness. Unlike the original SPDNet classification pipeline, its output remains on an SPD manifold and is classified directly by minimum distance to Riemannian mean (MDRM), rather than being vectorized for an end-to-end Euclidean classifier.
Datasets (the paper evaluates exactly two public BCI datasets):
- BCI Competition III, Dataset IIIa: 3 subjects, 60 EEG channels, four MI classes, 250 Hz.
- BCI Competition IV Dataset 2a: 9 subjects, 22 EEG channels, four MI classes, 250 Hz.

Architecture
A branch of the Siamese network alternates dimensionality-reducing BiMap layers and nonlinear Shrinkage layers:
The shrinkage layer preserves the dimension and eigenvectors of its input while moving every eigenvalue toward the mean eigenvalue . It compensates for the systematic overestimation of large eigenvalues and underestimation of small eigenvalues when high-dimensional covariance matrices are estimated from few samples, and prevents near-zero eigenvalues from making the logarithm in the affine-invariant Riemannian distance ill-conditioned.
After eigenvalue decomposition , the shrunk covariance matrix can be rewritten as follows:
It is worth noting that and have the same eigenvector matrix , and is an SPD matrix with the same dimension of . More importantly, the extreme eigenvalues of are modified towards the average value , which means that the largest and smallest eigenvalues are shrunk and elongated, respectively. Therefore, proper shrinkage is more effective and robust for small sample training sets. Compared with ReEig, shrinkage modifies all eigenvalues rather than hard-clipping only those below - soft thresholding vs hard thresholding. The direct forward expression does not require SVD; the ablation reports shorter training time than ReEig.
Two identical branches share all . Given a pair , the network computes
and minimizes the contrastive loss
where denotes a similar pair, a dissimilar pair, and is the contrastive margin. Similar pairs are pulled together and dissimilar pairs closer than are pushed apart. The value of and the pair-sampling procedure are not reported.
The trained mapping is followed by MDRM on the reduced SPD manifold. LogEig is discussed as an SPDNet primitive but is not part of the reported SPD-Mani-Net branch.
For multi-subject learning, SPD-Mani-Net+Reg introduces a regularization layer. Subject is aligned through toward the composite target mean
where the weights depend on inter-subject Riemannian distances. This reuses covariance geometry from other subjects to reduce calibration requirements.
Model Parameters
| Experiment | BiMap dimensions | Nonlinear layer |
|---|---|---|
| Synthetic | , , (3 BiMap + 2 nonlinear) | Shrinkage ; SPDNet ReEig |
| 2-class EEG, IIIa | , , , , (5 BiMap + 5 nonlinear) | Shrinkage |
| 2-class EEG, BCI Competition IV Dataset 2a | , , , , | Shrinkage |
| 4-class multi-subject, IIIa | , (2 BiMap + 2 nonlinear) | ; Reg |
| 4-class multi-subject, BCI Competition IV Dataset 2a | , | ; Reg |
Training Parameters
- Optimization: SGD with backpropagation; BiMap weights updated on the Stiefel manifold through a Riemannian gradient and retraction.
- Learning rate: .
- Batch size: 50.
- BiMap initialization: random semi-orthogonal matrices.
- Epochs, stopping criterion, momentum, weight decay, scheduler, and hardware: not reported.
EEG preprocessing and splitting:
- Both datasets sampled at 250 Hz and band-pass filtered at 8–30 Hz (same preprocessing as reference [45]).
- Two-class experiment retains only left- and right-hand MI.
- Two-class uses the original competition train/test split, which for BCI Competition IV Dataset 2a corresponds to training session T → evaluation session E (within-subject cross-session); the dataset audit labels this protocol “Within-Subject / Cross-Session” for both datasets (see BCI Competition III Dataset IIIa).
- Dataset IIIa: 45 trials/class in each set for B1, 30 trials/class for B2 and B3.
- Dataset IIa (2a): 72 trials/class in each train/test set for subjects C1–C9.
- Four-class multi-subject experiment pools all subjects (multi-subject, cross-subject per the dataset audits); all signals of each subject are merged as training and testing set, with the precise split construction not stated.
- Exact epoch interval, baseline correction, artifact rejection, and covariance regularization are not reported.
- Caution: the means ± SD reported in the tables below pool all 12 B/C subject columns (IIIa + 2a together). The dataset audit summaries in BCI Competition III Dataset IIIa and BCI Competition IV Dataset 2a present these pooled means as if they were dataset-specific results; per-dataset means are not reported in the paper and should not be invented.
Results
Two-class EEG accuracy (%):
| Method | Mean ± SD | B1 | B2 | B3 | C1 | C2 | C3 | C4 | C5 | C6 | C7 | C8 | C9 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MDRM | 78.5 ± 16.1 | 97.8 | 63.3 | 88.3 | 88.2 | 52.8 | 92.4 | 71.5 | 58.3 | 64.6 | 75.0 | 95.8 | 94.4 |
| CSP+LDA | 79.4 ± 16.8 | 95.6 | 61.7 | 93.3 | 88.9 | 51.4 | 96.5 | 70.1 | 54.9 | 71.5 | 81.3 | 93.8 | 93.8 |
| Ga-DR | 78.2 ± 14.9 | 96.7 | 68.3 | 85.0 | 87.5 | 53.5 | 92.4 | 73.6 | 57.6 | 68.0 | 70.8 | 94.4 | 91.6 |
| Ga-PCA | 68.5 ± 13.0 | 80.0 | 63.3 | 68.3 | 77.8 | 50.0 | 84.7 | 64.5 | 53.4 | 56.9 | 56.2 | 84.0 | 84.0 |
| DPLM | 75.6 ± 15.3 | 85.6 | 63.3 | 75.0 | 89.6 | 56.9 | 93.1 | 70.8 | 56.9 | 58.3 | 68.0 | 95.1 | 94.4 |
| SPD-Net | 76.9 ± 17.1 | 97.7 | 66.7 | 88.3 | 84.7 | 56.3 | 93.8 | 68.1 | 56.9 | 62.5 | 56.3 | 95.8 | 95.1 |
| SPD-Mani-Net | 83.1 ± 14.9 | 100.0 | 66.7 | 98.3 | 94.4 | 57.6 | 93.1 | 75.0 | 71.5 | 66.7 | 83.3 | 96.5 | 94.4 |
Four-class multi-subject accuracy (%):
| Method | Mean ± SD | B1 | B2 | B3 | C1 | C2 | C3 | C4 | C5 | C6 | C7 | C8 | C9 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MDRM | 43.61 ± 16.71 | 67.78 | 39.17 | 27.50 | 61.46 | 27.08 | 64.93 | 39.93 | 25.00 | 22.22 | 61.46 | 46.88 | 39.93 |
| SPD-Net | 45.75 ± 17.56 | 70.56 | 36.67 | 38.33 | 66.67 | 26.74 | 69.80 | 37.50 | 25.35 | 26.04 | 55.21 | 60.07 | 36.11 |
| SPD-Mani-Net | 48.21 ± 15.73 | 65.00 | 33.33 | 35.00 | 65.28 | 28.13 | 68.75 | 45.49 | 29.51 | 34.03 | 51.74 | 64.58 | 57.64 |
| SPD-Mani-Net+Reg | 53.28 ± 17.78 | 83.30 | 43.30 | 32.50 | 61.11 | 33.33 | 69.44 | 42.71 | 39.24 | 32.99 | 62.85 | 69.44 | 69.10 |
Four-class multi-subject Kappa: MDRM 0.25 ± 0.22, SPD-Net 0.28 ± 0.24, SPD-Mani-Net 0.31 ± 0.21, SPD-Mani-Net+Reg 0.37 ± 0.24.
The synthetic-data ablation finds that shrinkage alone does not significantly improve accuracy over ReEig but shortens training; the Siamese architecture is identified as the main source of error reduction (exact values graphical only). FgMDM and SPDNetBN are not evaluated.
Notes
- The reported architecture is BiMap + Shrinkage with a Siamese contrastive objective and downstream MDRM; it does not contain a LogEig output layer.
- The paper attributes poor SPD-Net EEG results to LogEig vectorization followed by a Euclidean network.
- EEG preprocessing is incompletely specified: only the 8–30 Hz filter and 250 Hz sampling are explicit. Epoch timing, contrastive margin, epochs, optimizer details, and parameter counts are not reported.