1 EE(G)-SPDNet
EE(G)-SPDNet is an end-to-end Deep Riemannian Network for EEG that learns a filterbank with temporal convolutions (channel-specific or channel-independent), computes spatial covariance matrices, and classifies through an SPDNet, with a Bayesian-optimized filterbank variant (BO-SPDNet) as companion.
Ref: @wilsonDeepRiemannianNetworks2022
Code available at https://github.com/dcwil/eegspdnet.
Overview
Problem: DRNs for EEG so far used handcrafted filterbanks (Tensor-CSPNet 9-band, MRC-MLP 43-band, RFNet 50-band, TSMNet 4-band conv), and it was unclear how network size, filterbank learning, and end-to-end ability affect performance. The study proposes two DRNs with learnable filterbanks, systematically compared over filterbank size (1–8 bands), channel specificity, and interband covariance, against state-of-the-art ConvNets.
- EE(G)-SPDNet: a second-order convolutional network — temporal convolution (mimicking band-pass filtering, learned end-to-end via backprop) → sample covariance pooling → SPDNet. Optimized jointly with classification; flexible across sampling rates; can learn band-/notch-/multiband filters; channel-specific mode gives one filter per electrode per band.
- BO-SPDNet: identical SPDNet, but the filterbank (bandpass cut-off pairs) is searched by Bayesian optimization using rSVM cross-validation accuracy (Log-Euclidean kernel) as the objective; restricted to Sinc-like bandpass responses.
- Datasets/tasks (exactly two): High Gamma Dataset (HGD) — motor movement, 14 participants, ~1,000 trials/subject (≈880 validation, ~160 final evaluation), 44 selected electrodes of 128, four classes (rest; movement of right hand, left hand, tongue/feet), 250 Hz after resampling. BCI Competition IV Dataset 2a — motor imagery, 9 participants, 22 channels, 2 sessions × 288 trials, four classes; session 1 for validation splits, session 2 as final evaluation.
- Preprocessing (matching Schirrmeister et al. 2017): channel selection (44-channel HGD montage; 22 channels for 2a), clipping at 800 µV, resampling to 250 Hz; HGD band-pass 4–125 Hz, BCIC-IV-2a band-pass 4–38 Hz. Validation phase uses an 80:20 train/test split of the training data as a proxy for the holdout.
- Contributions: first systematic comparison of DRN filterbank design; end-to-end DRNs outperform Deep4/ShallowFBCSP ConvNets using physiologically plausible frequencies (10–20, 20–35, 65–90 Hz); layer-by-layer (LBL) analysis revealing loss of Riemannian information through the SPDNet; multi-band filter learning analysis.
Architecture
Pipeline: EEG trial → Conv (learnable filterbank) → SCM pooling → SPDNet [3 × (BiMap → ReEig)] → LogEig → Vect → Linear.
- Convolution. Temporal kernels (length 25 samples) applied per band, either channel-independent (one filter per band for all electrodes) or channel-specific (one filter per electrode per band). For channel-independent multi-band models the per-band covariances are concatenated into a block-diagonal matrix, optionally including the interband covariance (off-diagonal blocks computed by stacking band-filtered signals along the electrode axis before covariance estimation) — three model sub-types: Spec, Ind, Ind [RMINT].
- Covariance pooling. Sample Covariance Matrix over each (filtered) trial (the CovLayer primitive):
Multi-band concatenation is provably SPD (block-diagonal of SPD blocks). Regularization applied only when matrices are non-PD (needed for < 0.01% of trials); BiMap/ReEig also force PD.
3. SPDNet (Huang & Van Gool 2017): three BiMap–ReEig pairs, each BiMap roughly halving the SPD dimension:
with rectification threshold (value not reported); then LogEig (Log-Euclidean metric, identity reference):
then vectorization (upper-triangle, redundant off-diagonals kept), then a linear layer with cross-entropy.
4. Meta-optimization (from the adavoudi/spdnet repo): the SPDNet SGD BiMap update is replaced by , then (any optimizer, here Adam) and retraction back to the Stiefel manifold; matrix backpropagation uses the Ionescu method (identical results to Daleckiĭ–Kreĭn in initial tests).
5. BO-SPDNet: Bayesian optimizer over bandpass cut-off pairs; objective = negative stratified 5-fold CV accuracy of an rSVM (linear kernel, LEM, C = 10) on the SPD matrices; the winning filterbank is then applied and the (identical) SPDNet trained on top.
Model Parameters
- Convolutional kernel length: 25.
- Number of bands ; per band: filters total (Ind) or per electrode (Spec); three sub-types (Spec / Ind / Ind[RMINT]).
- SPDNet: 3 BiMap–ReEig pairs, each BiMap halves the dimension (e.g., 44→22→11→≈6 for HGD input 44×44; 22→11→6→3 for 2a 22×22; exact final sizes not tabulated).
- Depth analysis suggests a minimum final matrix size of ~10–20 is required; “all models consistently favoured ”.
- Random seeds: 4. SPD estimator: SCM. Total parameter counts: not reported.
Training Parameters
- Optimizer: Adam (found best in initial tests; no thorough optimizer/LR search performed).
- Learning rate: 0.001; batch size: 216; LR scheduler: cosine annealing; epochs: 1,000; loss: cross-entropy.
- BO-SPDNet: BO stopping criteria 2,500 iterations or 12 hours; rSVM C = 10, linear kernel; stratified 5-fold CV for the objective.
- Splits: validation 80:20; final evaluation on the holdout (HGD final set ~160 trials/subject; 2a session 2). 2a models re-trained/re-tested without hyperparameter optimization (transferred hyperparameters from HGD validation).
- Software/hardware: Python 3.8, PyTorch, MNE, Braindecode, HyperOpt; bwForCluster NEMO.
Results
Final evaluation-set accuracy (%):
| Model | Type (# bands) | HGD | BCIC-IV-2a |
|---|---|---|---|
| EE(G)-SPDNet | Spec (8) | 97.0 (p<0.001 vs both ConvNets) | 75.2 |
| BO-SPDNet | Ind (8) | 94.5 | 74.7 |
| ConvNet Deep4 | - | 93.1 | 73.0 |
| ConvNet ShallowFBCSP | - | 94.4 | 72.9 |
- EE(G)-SPDNet beats Deep4 and ShallowFBCSP on HGD with statistical significance (Wilcoxon signed-rank); effects reproduced on 2a with decreased significance; BO-SPDNet improvements not significant. Best EE vs best BO on the evaluation set only marginally significant (p = 0.0495).
- Channel specificity: EE(G)-SPDNet strongly favors channel-specific filtering (+2.2% average, p = 0.00781); BO-SPDNet favors channel-independent (+1.7%, p = 0.00781); EE’s preference shrinks as grows.
- Interband covariance: including it gives +2.1% (EE, p = 0.0156) and +0.42% (BO, p = 0.0156) on validation.
- Filterbank learning: BO slightly better in validation (except channel-specific); EE highest on evaluation, more robust from validation to evaluation.
- Learned frequencies: gain spectra peak at 10–20 Hz and 20–35 Hz, plus a high-gamma peak at 65–90 Hz (channel-independent only when interband covariance kept); peaks flatten as increases.
- Depth: number of BiMap–ReEig pairs has little effect as long as the final matrix size ≥ ~10–20.
- Layer-by-layer analysis: an rSVM on intermediate representations often exceeds the final network accuracy (especially early/middle layers for channel-specific models); Euclidean SVMs are much worse than rSVM but improve through the network — evidence that the network optimizes features that lose Riemannian-specific benefit.