1 SPDNet

SPDNet is the foundational end-to-end Riemannian neural network that maps SPD matrices through dimensionality-reducing bilinear maps and eigenvalue nonlinearities before Euclidean classification.

Ref: @huangRiemannianNetworkSPD2016

Overview

SPDNet receives SPD matrices directly and preserves positive definiteness through its manifold-valued layers. Its principal components are bilinear mapping (BiMap), eigenvalue rectification (ReEig), and eigenvalue logarithm (LogEig), followed by ordinary Euclidean fully connected and softmax layers. BiMap performs manifold-to-manifold dimensionality reduction, ReEig introduces a ReLU-like nonlinearity, and LogEig maps the result into the flat Log-Euclidean domain. The semi-orthogonal BiMap weights live on compact Stiefel manifolds and are optimized with Riemannian matrix backpropagation. The paper evaluates emotion recognition (AFEW), skeleton-based action recognition (HDM05), and video face verification (PaSC).

Architecture

For an input , an SPDNet with BiMap/ReEig blocks has the pipeline

The reported three-block model is

BiMap

The BiMap layer applies the congruence transformation

where and with . If has full column rank, then for every ,

so BiMap maps to . It is both a bilinear dimensionality-reduction layer and a learned manifold-to-manifold metric embedding. The semi-orthogonality constraint places the weights on the compact Stiefel manifold .

ReEig

The ReEig layer introduces nonlinearity by flooring small eigenvalues. For ,

with element-wise flooring to . It moves matrices away from the positive-semidefinite boundary and acts as an SPD analogue of ReLU.

LogEig and Classifier

The LogEig layer maps the final SPD representation into the Log-Euclidean tangent space:

The symmetric logarithmic matrix is vectorized and passed to a fully connected layer followed by softmax log loss.

Riemannian Matrix Backpropagation

For a composition , matrix backpropagation propagates both weight and data derivatives. Under the paper’s semi-orthogonal convention, the Euclidean BiMap gradient is

Its tangential component and retracted SGD update on the Stiefel manifold are

For eigendecomposition-based layers, with ,

ReEig differentiates with a masking matrix when else , and LogEig differentiates through with .

Model Parameters

DatasetInput SPD sizeSPDNet-3BiRe progressionOutput
AFEWFC to 7 expressions, softmax
HDM05FC to 130 actions, softmax
PaSCFC + softmax for verification

Shared settings: random semi-orthogonal BiMap initialization; ReEig threshold ; BiRe block counts evaluated: 0, 1, 2, 3.

Training Parameters

SettingAFEWHDM05PaSC
OptimizerSGD on Stiefel manifold (gradient + retraction)SameSame
Learning rate, fixed, fixed, fixed
Batch size303030
Epochs500500100
LossSoftmax log lossSoftmax log lossSoftmax log loss
Time/epoch≈2 min≈4 min≈15 min
Hardwarei7-2600K 3.40 GHz CPU, no GPUSameSame

Data protocols:

  • AFEW: frames normalized to ; each video is a covariance; training videos segmented into 1,747 clips for augmentation; results on the validation set (test labels unavailable).
  • HDM05: covariance of 31-joint 3-D coordinates (); ten random evaluations, half sequences train / half test; ≈18,000 training clips per evaluation.
  • PaSC: deep face features reduced by PCA to 400 dims, fused with the mean into SPD matrices; 280 training + 900 COX videos augmented to 12,529 clips; control and handheld protocols evaluated separately.

Results

Accuracy (%). Main comparison:

MethodAFEWHDM05PaSC controlPaSC handheld
STM-ExpLet31.73---
RSR-SPDML30.12--
CDL31.8178.2970.41
SPDML-AIM26.7265.4759.03
SPDNet-0BiRe26.3268.5263.92
SPDNet-1BiRe29.1271.7565.81
SPDNet-2BiRe31.5476.2369.64
SPDNet-3BiRe34.2380.1272.83

ReEig-threshold ablation on AFEW: → 34.23; → 33.15; → 32.35. Removing LogEig drops AFEW to 21.49 and HDM05 to 4.89.

EEG Evaluations in Later Papers

The original paper did not evaluate EEG; the following are third-party runs, not Huang and Van Gool’s experiments.

SourceProtocolSPDNet result
Ju et al. (2022), MI-KUWithin-session 10-fold CV (S1 / S2); cross-session holdout S1→S257.88 (8.68) / 58.88 (8.68) / 60.41 (12.13)
Ju et al. (2022), BCI Competition IV Dataset 2aWithin-session CV (T / E); holdout T→E65.91 (10.31) / 61.16 (10.50) / 55.67 (9.54)
Peng et al. (2023), IIIa + 2a2-class, original train/test split; 4-class multi-subject76.9 ± 17.1 (2-class, pooled); 45.75 ± 17.56 (4-class, pooled)
Carrara and Papadopoulo (2024), Phase-SPDNet, Zhou2016Binary L/R; within-session 5-fold CV; C3/Cz/C4; 8-32 Hz; AUCSPDNet 88.85 ± 8.05; coherence variants 62.89 / 75.07

Details per Source

Ju et al. (2022) - SPDNet Baseline on MI-KU and 2a

  • Dataset parameters: MI-KU - 54 subjects, 20 selected motor-cortex channels of 62, binary MI, imagery 1–3.5 s, 1,000 Hz, 2 sessions. 2a - 9 subjects, 22 EEG channels, four classes, imagery 0–4 s, 250 Hz, 2 sessions. Trials assumed band-pass filtered (nine causal Chebyshev Type-II filters in 4-Hz bins over 4-40 Hz), centered, and scaled.
  • Model parameters: classic SPDNet pipeline - spatial covariance per trial → BiMap → ReEig → LogEig → Euclidean classifier. Exact BiMap dimensions and ReEig threshold for the baseline are not tabulated in the paper (the paper’s own Tensor-CSPNet uses depthwise BiMaps with output dimension o and RBN; the SPDNet baseline uses the standard summing BiMap).
  • Training: the paper-wide protocol is cross-entropy, initial lr 0.01 with decay, batch 28, maximum 60 epochs, early-stopping patience 15, Riemannian optimization on the Stiefel manifold (see Tensor-CSPNet note); baseline-specific values not separately reported.
  • Results: see the summary table above; the paper notes SPDNet only exploits spatial patterns (no temporal dynamics), which the authors give as the reason it trails the tensor models.

Peng et al. (2023) - SPD-Net Baseline on IIIa and 2a

  • Dataset parameters: 250 Hz, 8-30 Hz band-pass, binary left/right-hand MI (2-class) or all four classes (multi-subject). Original train/test splits: IIIa 45 trials/class per set for B1 and 30 for B2/B3; 2a 72 trials/class per set (C1-C9). The 4-class experiment merges all subjects’ signals into one training/testing set (merged-subject, not held-out-subject).
  • Model parameters (2-class): five BiMap layers + five ReEig layers, reducing to a 6×6 SPD target: IIIa dims 60×50 → 50×40 → 40×30 → 30×20 → 20×6; 2a dims 22×19 → 19×16 → 16×13 → 13×10 → 10×6. ReEig threshold ε = 10⁻⁴. (The same 5-block structure is used by SPD-Mani-Net, with shrinkage γ = 0.1 replacing ReEig.)
  • Model parameters (4-class): two BiMap + two ReEig layers: IIIa 60×40 → 40×10; 2a 22×15 → 15×10; ε = 10⁻⁴.
  • Model parameters (toy experiment): three BiMap layers 30×25, 25×20, 20×10 + two ReEig; ε = 10⁻⁴ (shrinkage γ = 0.01 for SPD-Mani-Net).
  • Training: SGD with Riemannian gradient + retraction on the Stiefel manifold; lr 10⁻²; batch 50; random semi-orthogonal BiMap initialization. Epoch count not reported. Classification after LogEig vectorization into a Euclidean network (this vectorize-then-Euclidean tail is the paper’s stated reason SPD-Net underperforms on EEG).
  • Results: 2-class pooled mean 76.9 ± 17.1; 4-class pooled mean 45.75 ± 17.56 (kappa 0.28 ± 0.24). Both pooled over all 12 IIIa+2a subjects - not dataset-specific.

Carrara and Papadopoulo (2024), Phase-SPDNet - Standard SPDNet Baseline

  • Dataset parameters: six MOABB MI datasets (BNCI2014001, BNCI2014004, Cho2017, Schirrmeister2017, Weibo2014, Zhou2016); binary left vs right hand; only 3 electrodes C3/Cz/C4; band-pass 8–32 Hz (overlap-add), artifact rejection, per-channel zero-mean/unit-std standardization, full epoch; within-session 5-fold CV per session; metric ROC-AUC. See Phase-SPDNet note for per-dataset epochs/rates.
  • Model parameters: standard SPDNet but with the subspace dimension kept equal to the input dimension (rotation-only BiMap) so that GradCam++ explainability analyses are possible; halving the subspace consistently reduced standard-SPDNet performance. ReEig ε = 10⁻⁴; the used ReEig implementation has no trainable parameters.
  • Training: MOABB 1.1.0 common config - 300 epochs, batch 64, validation split 0.1, sparse categorical cross-entropy, Adam lr 10⁻³ (+ RiemannAdam on the Stiefel manifold for BiMap), early stopping patience 75, ReduceLROnPlateau factor 0.5.
  • Results: Zhou2016 88.85 ± 8.05 AUC (imag-coherence variant 62.89 ± 8.97; inst-coherence 75.07 ± 7.63). Under the same protocol Phase-SPDNet-OPT reaches 95.62 ± 2.36.

Notes

This gives an additional justification for ReEig beyond a mere activation heuristic.