1 EE(G)-SPDNet

EE(G)-SPDNet is an end-to-end Deep Riemannian Network for EEG that learns a filterbank with temporal convolutions (channel-specific or channel-independent), computes spatial covariance matrices, and classifies through an SPDNet, with a Bayesian-optimized filterbank variant (BO-SPDNet) as companion.

Ref: @wilsonDeepRiemannianNetworks2022

Code available at https://github.com/dcwil/eegspdnet.

Overview

Problem: DRNs for EEG so far used handcrafted filterbanks (Tensor-CSPNet 9-band, MRC-MLP 43-band, RFNet 50-band, TSMNet 4-band conv), and it was unclear how network size, filterbank learning, and end-to-end ability affect performance. The study proposes two DRNs with learnable filterbanks, systematically compared over filterbank size (1–8 bands), channel specificity, and interband covariance, against state-of-the-art ConvNets.

  • EE(G)-SPDNet: a second-order convolutional network — temporal convolution (mimicking band-pass filtering, learned end-to-end via backprop) → sample covariance pooling → SPDNet. Optimized jointly with classification; flexible across sampling rates; can learn band-/notch-/multiband filters; channel-specific mode gives one filter per electrode per band.
  • BO-SPDNet: identical SPDNet, but the filterbank (bandpass cut-off pairs) is searched by Bayesian optimization using rSVM cross-validation accuracy (Log-Euclidean kernel) as the objective; restricted to Sinc-like bandpass responses.
  • Datasets/tasks (exactly two): High Gamma Dataset (HGD) — motor movement, 14 participants, ~1,000 trials/subject (≈880 validation, ~160 final evaluation), 44 selected electrodes of 128, four classes (rest; movement of right hand, left hand, tongue/feet), 250 Hz after resampling. BCI Competition IV Dataset 2a — motor imagery, 9 participants, 22 channels, 2 sessions × 288 trials, four classes; session 1 for validation splits, session 2 as final evaluation.
  • Preprocessing (matching Schirrmeister et al. 2017): channel selection (44-channel HGD montage; 22 channels for 2a), clipping at 800 µV, resampling to 250 Hz; HGD band-pass 4–125 Hz, BCIC-IV-2a band-pass 4–38 Hz. Validation phase uses an 80:20 train/test split of the training data as a proxy for the holdout.
  • Contributions: first systematic comparison of DRN filterbank design; end-to-end DRNs outperform Deep4/ShallowFBCSP ConvNets using physiologically plausible frequencies (10–20, 20–35, 65–90 Hz); layer-by-layer (LBL) analysis revealing loss of Riemannian information through the SPDNet; multi-band filter learning analysis.

Architecture

Pipeline: EEG trial → Conv (learnable filterbank) → SCM pooling → SPDNet [3 × (BiMap → ReEig)] → LogEig → Vect → Linear.

  1. Convolution. Temporal kernels (length 25 samples) applied per band, either channel-independent (one filter per band for all electrodes) or channel-specific (one filter per electrode per band). For channel-independent multi-band models the per-band covariances are concatenated into a block-diagonal matrix, optionally including the interband covariance (off-diagonal blocks computed by stacking band-filtered signals along the electrode axis before covariance estimation) — three model sub-types: Spec, Ind, Ind [RMINT].
  2. Covariance pooling. Sample Covariance Matrix over each (filtered) trial (the CovLayer primitive):

Multi-band concatenation is provably SPD (block-diagonal of SPD blocks). Regularization applied only when matrices are non-PD (needed for < 0.01% of trials); BiMap/ReEig also force PD.
3. SPDNet (Huang & Van Gool 2017): three BiMap–ReEig pairs, each BiMap roughly halving the SPD dimension:

ReEig:

with rectification threshold (value not reported); then LogEig (Log-Euclidean metric, identity reference):

then vectorization (upper-triangle, redundant off-diagonals kept), then a linear layer with cross-entropy.
4. Meta-optimization (from the adavoudi/spdnet repo): the SPDNet SGD BiMap update is replaced by , then (any optimizer, here Adam) and retraction back to the Stiefel manifold; matrix backpropagation uses the Ionescu method (identical results to Daleckiĭ–Kreĭn in initial tests).
5. BO-SPDNet: Bayesian optimizer over bandpass cut-off pairs; objective = negative stratified 5-fold CV accuracy of an rSVM (linear kernel, LEM, C = 10) on the SPD matrices; the winning filterbank is then applied and the (identical) SPDNet trained on top.

Model Parameters

  • Convolutional kernel length: 25.
  • Number of bands ; per band: filters total (Ind) or per electrode (Spec); three sub-types (Spec / Ind / Ind[RMINT]).
  • SPDNet: 3 BiMap–ReEig pairs, each BiMap halves the dimension (e.g., 44→22→11→≈6 for HGD input 44×44; 22→11→6→3 for 2a 22×22; exact final sizes not tabulated).
  • Depth analysis suggests a minimum final matrix size of ~10–20 is required; “all models consistently favoured ”.
  • Random seeds: 4. SPD estimator: SCM. Total parameter counts: not reported.

Training Parameters

  • Optimizer: Adam (found best in initial tests; no thorough optimizer/LR search performed).
  • Learning rate: 0.001; batch size: 216; LR scheduler: cosine annealing; epochs: 1,000; loss: cross-entropy.
  • BO-SPDNet: BO stopping criteria 2,500 iterations or 12 hours; rSVM C = 10, linear kernel; stratified 5-fold CV for the objective.
  • Splits: validation 80:20; final evaluation on the holdout (HGD final set ~160 trials/subject; 2a session 2). 2a models re-trained/re-tested without hyperparameter optimization (transferred hyperparameters from HGD validation).
  • Software/hardware: Python 3.8, PyTorch, MNE, Braindecode, HyperOpt; bwForCluster NEMO.

Results

Final evaluation-set accuracy (%):

ModelType (# bands)HGDBCIC-IV-2a
EE(G)-SPDNetSpec (8)97.0 (p<0.001 vs both ConvNets)75.2
BO-SPDNetInd (8)94.574.7
ConvNet Deep4-93.173.0
ConvNet ShallowFBCSP-94.472.9
  • EE(G)-SPDNet beats Deep4 and ShallowFBCSP on HGD with statistical significance (Wilcoxon signed-rank); effects reproduced on 2a with decreased significance; BO-SPDNet improvements not significant. Best EE vs best BO on the evaluation set only marginally significant (p = 0.0495).
  • Channel specificity: EE(G)-SPDNet strongly favors channel-specific filtering (+2.2% average, p = 0.00781); BO-SPDNet favors channel-independent (+1.7%, p = 0.00781); EE’s preference shrinks as grows.
  • Interband covariance: including it gives +2.1% (EE, p = 0.0156) and +0.42% (BO, p = 0.0156) on validation.
  • Filterbank learning: BO slightly better in validation (except channel-specific); EE highest on evaluation, more robust from validation to evaluation.
  • Learned frequencies: gain spectra peak at 10–20 Hz and 20–35 Hz, plus a high-gamma peak at 65–90 Hz (channel-independent only when interband covariance kept); peaks flatten as increases.
  • Depth: number of BiMap–ReEig pairs has little effect as long as the final matrix size ≥ ~10–20.
  • Layer-by-layer analysis: an rSVM on intermediate representations often exceeds the final network accuracy (especially early/middle layers for channel-specific models); Euclidean SVMs are much worse than rSVM but improve through the network — evidence that the network optimizes features that lose Riemannian-specific benefit.