1 U-SPDNet
U-SPDNet extends SPDNet with a symmetric manifold-valued decoder, reconstruction supervision, and Log-Euclidean skip connections to mitigate structural-information degradation.
Date of publication: 14/12/2022
Ref: @wangUSPDNetSPDManifold2023
Source code: https://github.com/GitWR/U-SPDNet
Overview
U-SPDNet treats successive low-dimensional BiMap operations as a source of structural-information loss. Its contracting path is an SPDNet encoder, while a symmetric expanding path reconstructs the original SPD representation from the encoder output. A reconstruction error term encourages the complete manifold-to-manifold embedding to approximate an identity map, while cross-entropy supervises the encoder representation used for classification. Skip connections fuse corresponding encoder and decoder features through a Log-Fusion-Exp (LFE) implementation of the Log-Euclidean Riemannian barycenter. Evaluated on MDSD, Virus, FPHA, and UAV-Human.
Quote
The fundamental reasons for the effectiveness of Riemannian neural networking technique lie in two aspects: (1) the Riemannian geometry of the input data manifold can be preserved through all the layers during training; (2) the deep and nonlinear feature embedding mechanism.

Architecture
The full pipeline is
Encoder
The BiMap layer is with , ; full column rank preserves positive definiteness and semi-orthogonality places on a compact Stiefel manifold. The ReEig layer computes . The classification branch applies LogEig (), symmetric vectorization, a fully connected projection , softmax, and cross-entropy.
Decoder
The decoder mirrors the encoder and upsamples SPD matrices with where and . ReEig is applied in the decoder so that the upsampled outputs remain SPD. The symmetric structure avoids extra dimensional-alignment parameters at skip connections.
LFE Skip Connections
For corresponding encoded and decoded SPD features , the Log-Euclidean Fréchet mean has closed form
implemented as the LFE module:
\{M_h\} \xrightarrow{\ [[Deep Riemannian Network (DRN)#LogEig|LogEig]]\ } \{\log M_h\} \xrightarrow{\ [[Deep Riemannian Network (DRN)#AriM|AriM]]\ } \widetilde A = \frac{1}{H}\sum_h\log M_h \xrightarrow{\ [[Deep Riemannian Network (DRN)#ExpEig|ExpEig]]\ } P^*.The AriM (arithmetic mean) layer computes the Euclidean barycenter in the log domain. The paper calls this Riemannian optimization, but the skip connection is realized through the closed-form Log-Euclidean barycenter rather than explicit parallel transport.

Reconstruction and Total Losses
The reconstruction error term uses the Euclidean/Frobenius metric:
where is the reconstructed SPD matrix. The classification term is cross-entropy, and the complete objective is
where and balances reconstruction and classification.
The paper also derives the matrix backpropagation for ExpEig,
Model Parameters
| Dataset | Encoder dimensions | Decoder dimensions | Parameters |
|---|---|---|---|
| MDSD | 0.24 M | ||
| Virus | 0.23 M | ||
| FPHA | 0.06 M | ||
| UAV-Human | not reported |
| Dataset | Batch size | ReEig | Loss weight |
|---|---|---|---|
| MDSD | 20 | ||
| Virus | 10 | ||
| FPHA | 30 | ||
| UAV-Human | 30 |
Input covariance matrices are regularized as with .
Training Parameters
| Setting | MDSD | Virus | FPHA | UAV-Human |
|---|---|---|---|---|
| Optimizer | SGD on Stiefel manifolds (Riemannian matrix backprop) | Same | Same | Same |
| Initial learning rate | 0.01 | 0.01 | 0.01 | 0.01 |
| Schedule | ×0.8 every 50 ep | ×0.8 every 50 ep | not reported | ×0.9 every 50 ep |
| Maximum epochs | 500 | 300 | 1,400 | 1,400 |
| Time/epoch | 2.84 s | 2.35 s | 3.62 s | not reported |
Hardware: i7-9700 3.4 GHz CPU, 8 cores, 16 GB RAM. A GTX 2080 Ti did not accelerate training (the sequence of eigenvalue operations is the bottleneck). BiMap initialization: random semi-orthogonal.
Data protocols: MDSD 7 train / 3 test videos per class, frames → covariances; Virus 3 train / 2 test sets of ; FPHA 63-D joints → , 600 train / 575 test; UAV-Human 51-D (PCA 99% energy) → , 16,723 sequences split 70:30.
Results
Reconstruction metric comparison on MDSD:
| Reconstruction metric | Accuracy (%) | Time (s/epoch) |
|---|---|---|
| RET-EuM | 38.97 | 2.84 |
| RET-LEM | 39.49 | 6.97 |
LEM improves accuracy by 0.52 points but more than doubles epoch time, confirming the choice of EuM.
MDSD and Virus (accuracy %):
| Method | MDSD | Virus |
|---|---|---|
| SPDNet | ||
| SymNet | ||
| U-SPDNet |
FPHA (accuracy %): SPDNet 86.26, SymNet 82.96, U-SPDNet 87.83. UAV-Human: SPDNet 42.31, SymNet 35.89, U-SPDNet 43.39.
Ablations:
| Architecture | MDSD | Virus | FPHA |
|---|---|---|---|
| SPDNet baseline | 86.26 | ||
| U-SPDNet without skip connections | 87.13 | ||
| U-SPDNet | 87.83 |
| Fusion method | MDSD | Virus | FPHA |
|---|---|---|---|
| Direct AriM (arithmetic) | 87.42 | ||
| LFE Riemannian barycenter | 87.83 |
Trade-off parameter searched in ; best MDSD/Virus at . removes reconstruction supervision and reduces to the SPDNet baseline.
Notes and Caveats
- The decoder reconstructs SPD representations, the embedding is not bijective, so U-SPDNet cannot segment like a Euclidean U-Net.
- The skip connection uses the Log-Euclidean barycenter (LogEig → AriM → ExpEig); no explicit parallel transport is given.
- No momentum or weight decay is reported. U-SPDNet was not evaluated on EEG.