1 Deep Riemannian Network (DRN)
Implemented at: means there is a ready-made implementation in spdlearn. Ref: alone means the layer is described in a paper and the code shown here is only a sketch or a minimal custom implementation.
Core SPD Layers
CovLayer
Implemented at: https://spdlearn.org/generated/covariance/spd_learn.modules.CovLayer.html
This layer computes spatial covariance matrices from multivariate time series, transforming input data of shape (…, n_channels, n_times) into SPD matrices of shape (…, n_channels, n_channels).
BiMap
Implemented at: https://spdlearn.org/generated/bilinear/spd_learn.modules.BiMap.html
A bilinear mapping layer for Symmetric Positive Definite (SPD) matrices. It applies the congruence transform
where is a learnable matrix. If is SPD and has full column rank, then is also SPD, so the layer reduces dimensionality while staying on the manifold.
>>> import torch
>>> from spd_learn.modules import BiMap
>>> bimap = BiMap(in_features=8, out_features=4)
>>> X = torch.randn(2, 8, 8)
>>> X = X @ X.mT + 0.1 * torch.eye(8) # Make SPD
>>> Y = bimap(X)
>>> Y.shape
torch.Size([2, 4, 4])ReEig
Implemented at: https://spdlearn.org/generated/modeig/spd_learn.modules.ReEig.html
This layer rectifies eigenvalues to ensure numerical stability while preserving positive definiteness. It acts as a ReLU-like non linearity on SPD matrices:
where is the eigendecomposition and is a floor that prevents eigenvalues from collapsing toward zero.
LogEig
Implemented at: https://spdlearn.org/generated/modeig/spd_learn.modules.LogEig.html
This layer maps SPD matrices to the tangent space at the identity through the matrix logarithm:
where is the eigendecomposition and applies the logarithm element-wise to the eigenvalues. If upper=True, the symmetric matrix is then vectorized by taking its upper triangular part.
ExpEig
Implemented at: https://spdlearn.org/generated/modeig/spd_learn.modules.ExpEig.html
This layer maps symmetric matrices from the tangent space back to the SPD manifold through the matrix exponential. It is the inverse of Deep Riemannian Network (DRN) > LogEig:
where is the eigendecomposition and applies the exponential element-wise to the eigenvalues.
>>> import torch
>>> from spd_learn.modules import LogEig, ExpEig
>>> X = torch.randn(2, 4, 4)
>>> X = X @ X.mT + 0.1 * torch.eye(4) # Make SPD
>>> Y = LogEig(upper=True)(X)
>>> Y.shape
torch.Size([2, 10])
>>> Z = ExpEig(upper=True)(Y)
>>> Z.shape
torch.Size([2,4,4])Model-Specific Layers
GaussAgg
Ref: @nguyenNeuralNetworkBased2019
GaussAgg aggregates a set of feature vectors into a single SPD matrix by fitting a Gaussian and using its Gaussian embedding. Given vectors ,
and the layer outputs
In ST-TS-HGR-NET this layer is used to summarize joint features over fingers, time windows, or sub-sequences while preserving first- and second-order statistics in SPD form. It takes a set of vectors as input and returns one SPD matrix per aggregation window.
SPDAgg
Ref: @nguyenNeuralNetworkBased2019
SPDAgg fuses several SPD matrices into a new SPD matrix through a learned weighted sum of bilinear transforms:
where each is an input SPD matrix and each is a learnable transformation. It is the final manifold-valued aggregation step in ST-TS-HGR-NET before LogEig and classification. In tensor form, it typically maps (batch, n_inputs, d_in, d_in) to (batch, d_out, d_out).
import torch
import torch.nn as nn
class SPDAgg(nn.Module):
def __init__(self, n_inputs, d_in, d_out):
super().__init__()
self.weights = nn.Parameter(torch.randn(n_inputs, d_out, d_in))
def forward(self, xs): # xs: (batch, n_inputs, d_in, d_in)
ys = [W @ X @ W.mT for X, W in zip(xs.unbind(dim=1), self.weights)]
return torch.stack(ys, dim=0).sum(dim=0)SPDConv
Ref: @zhangDeepManifoldtoManifoldTransforming2020
This layer is the DMT-Net analogue of convolution for multi-channel SPD maps. It is also natural to implement it as SPDConv2d. Instead of unconstrained kernels, each kernel is parameterized to be SPD:
with . Convolving an SPD map with SPD kernels yields another SPD map, so the layer performs local filtering while preserving manifold structure. In tensor form, it acts like a convolution from (batch, in_channels, D, D) to (batch, out_channels, D', D').
import torch
import torch.nn as nn
import torch.nn.functional as F
class SPDConv2d(nn.Module):
def __init__(self, in_channels, out_channels, kernel_size, eps=1e-4):
super().__init__()
k = kernel_size
self.V = nn.Parameter(torch.randn(out_channels, in_channels, k, k))
self.eps = eps
def forward(self, x): # x: (batch, in_channels, D, D)
k = self.V.size(-1)
eye = torch.eye(k, device=x.device, dtype=x.dtype)
W = self.V.transpose(-1, -2) @ self.V + self.eps * eye
return F.conv2d(x, W)SPD Activation
Ref: @zhangDeepManifoldtoManifoldTransforming2020
DMT-Net replaces eigendecomposition-based rectification with an SVD-free element-wise nonlinearity. The paper shows that applying functions such as , , or element-wise to an SPD matrix still yields an SPD matrix. In practice this gives a cheap nonlinear layer that stays on the manifold and is followed by Frobenius normalization to keep eigenvalues bounded.
SPD Recursive Layer
Ref: @zhangDeepManifoldtoManifoldTransforming2020
This layer is DMT-Net’s manifold-valued analogue of a GRU cell. Instead of scalar or vector gates, it updates a hidden SPD state with SPD-valued reset and update gates built from bilinear maps:
The main idea is that every operation is chosen to remain compatible with SPD geometry: the bilinear maps preserve SPD structure, the gate is SPD-valued, and the update is performed with Hadamard products rather than ordinary vector gating. Conceptually, this layer lets the network remember or forget temporal information without leaving the SPD manifold. In code, it usually maps (batch, time, channels, d, d) to (batch, time, channels, d, d).
The downside is that it is much more intricate than the other DRN layers, both mathematically and numerically. A full working Torch implementation is in SPD Recursive Layer.
Diagonalizing Layer
Ref: @zhangDeepManifoldtoManifoldTransforming2020
The diagonalizing layer maps a full SPD matrix to a diagonal SPD matrix
where is a positive element-wise activation. Because is diagonal, the subsequent matrix logarithm becomes an element-wise logarithm on the diagonal entries, avoiding the expensive eigendecomposition otherwise needed for LogEig-style computation.
SPD Matrix Pooling
Ref: @wangSymNetSimpleSymmetric2022
This layer generalizes pooling to SPD-valued features. SymNet studies three variants:
- Tangent-space pooling: apply
LogEig, perform standard max or mean pooling in the flat tangent space, then return to the manifold withExpEig. - SPD max pooling: apply region-wise max pooling directly on SPD matrices. It is simple, but can distort geometry because the pooling window does not move along geodesics.
- SPD mean pooling: replace arithmetic averaging by the Fréchet mean over the pooling region so the pooled output stays on the manifold.
AriM
Ref: @wangUSPDNetSPDManifold2023
AriM is the arithmetic mean layer used in U-SPDNet’s Log-Fusion-Exp skip block. After several SPD matrices are mapped to the logarithmic domain with LogEig, AriM computes their Euclidean barycenter
The fused matrix is then mapped back to the SPD manifold with ExpEig. It is cheaper than computing a true Riemannian barycenter, but it is only geometry-faithful after the log-domain relaxation. So the AriM layer itself is just the mean in the logarithmic domain, typically mapping (batch, n_mats, d, d) to (batch, d, d):
import torch
import torch.nn as nn
class AriM(nn.Module):
def forward(self, xs): # xs: (batch, n_mats, d, d) in the log domain
return xs.mean(dim=1)
class LogEuclideanAriM(nn.Module):
def __init__(self, log_layer, exp_layer):
super().__init__()
self.log_layer = log_layer
self.arim = AriM()
self.exp_layer = exp_layer
def forward(self, xs): # xs: (batch, n_mats, d, d) on the SPD manifold
xs_log = torch.stack([self.log_layer(x) for x in xs.unbind(dim=1)], dim=1)
mean_log = self.arim(xs_log)
return self.exp_layer(mean_log)Normalization and Regularization Layers
SPDBatchNormMeanVar
Implemented at: https://spdlearn.org/generated/batchnorm/spd_learn.modules.SPDBatchNormMeanVar.html
This layer is a batch-normalization analogue for SPD matrices. It first recenters a batch by parallel-transporting each sample from the batch Fréchet mean to the identity:
then applies dispersion normalization and an optional learnable re-biasing through a congruence transform with a learnable SPD matrix :
Compared with Euclidean batch normalization, the point is the same, but all operations are carried out in a manifold-consistent way.
SPDDSMBN
Ref: @koblerSPDDomainspecificBatch2022
Source code: https://github.com/rkobler/TSMNet
SPDDSMBN stands for SPD domain-specific momentum batch normalization. It combines domain-specific batch normalization with SPD momentum batch normalization by maintaining separate running Fréchet statistics for each domain:
In TSMNet this layer transforms domain-specific SPD inputs into domain-aligned outputs before tangent-space mapping, which is especially useful for cross-session and cross-subject EEG transfer.
Shrinkage
Available at SPDLearn: https://spdlearn.org/generated/regularization/spd_learn.modules.Shrinkage.html#spd_learn.modules.Shrinkage
Ref: @pengReducingDimensionalitySPD2023
Shrinkage is a common approach to compensate for the estimated bias, and covariance matrices are usually replaced by:
where is the shrinkage parameter and is the average eigenvalue of . After eigendecomposition , the shrunk matrix can be written as
Unlike ReEig, which only floors small eigenvalues, shrinkage moves all eigenvalues toward the spherical covariance . In practice this acts like a softer regularizer and is often more robust when covariance estimates come from small sample sets.
TraceNorm
Implemented at: https://spdlearn.org/generated/regularization/spd_learn.modules.TraceNorm.html
TraceNorm removes global scale from a covariance matrix by dividing it by its trace and optionally adding a small diagonal regularizer:
This produces unit-trace SPD matrices, preserving correlation structure while making representations less sensitive to overall signal amplitude differences across subjects, sessions, or trials.
SPDDropout
Implemented at: https://spdlearn.org/generated/dropout/spd_learn.modules.SPDDropout.html
SPDDropout is a structured dropout layer for SPD matrices. During training it randomly drops entire channels instead of individual entries, and fills the diagonal of dropped channels with a small so the result remains SPD. This makes it a geometry-aware regularizer rather than a naive element-wise dropout.
LogEuclideanResidual
Implemented at: https://spdlearn.org/generated/residual/spd_learn.modules.LogEuclideanResidual.html
This layer implements a residual or skip connection for SPD matrices in the Log-Euclidean geometry:
Instead of adding SPD matrices directly in the ambient Euclidean space, it performs the addition in the tangent space and maps back to the manifold. This gives an SPD-preserving residual block analogous to the skip connections used in standard ResNets.
PatchEmbeddingLayer
Implemented at: https://spdlearn.org/generated/utils/spd_learn.modules.PatchEmbeddingLayer.html
PatchEmbeddingLayer extracts temporal patches from an input tensor of shape (batch, channels, time) by unfolding the signal into a sequence of local windows. The output has shape (batch, n_patches, channels, patch_size), so it is mainly a preprocessing layer used before covariance estimation or patchwise SPD modeling rather than a manifold operation by itself.