1 Deep Riemannian Network (DRN)

Implemented at: means there is a ready-made implementation in spdlearn. Ref: alone means the layer is described in a paper and the code shown here is only a sketch or a minimal custom implementation.

Core SPD Layers

CovLayer

Implemented at: https://spdlearn.org/generated/covariance/spd_learn.modules.CovLayer.html
This layer computes spatial covariance matrices from multivariate time series, transforming input data of shape (…, n_channels, n_times) into SPD matrices of shape (…, n_channels, n_channels).

BiMap

Implemented at: https://spdlearn.org/generated/bilinear/spd_learn.modules.BiMap.html
A bilinear mapping layer for Symmetric Positive Definite (SPD) matrices. It applies the congruence transform

where is a learnable matrix. If is SPD and has full column rank, then is also SPD, so the layer reduces dimensionality while staying on the manifold.

>>> import torch
>>> from spd_learn.modules import BiMap
>>> bimap = BiMap(in_features=8, out_features=4)
>>> X = torch.randn(2, 8, 8)
>>> X = X @ X.mT + 0.1 * torch.eye(8)  # Make SPD
>>> Y = bimap(X)
>>> Y.shape
torch.Size([2, 4, 4])

ReEig

Implemented at: https://spdlearn.org/generated/modeig/spd_learn.modules.ReEig.html
This layer rectifies eigenvalues to ensure numerical stability while preserving positive definiteness. It acts as a ReLU-like non linearity on SPD matrices:

where is the eigendecomposition and is a floor that prevents eigenvalues from collapsing toward zero.

LogEig

Implemented at: https://spdlearn.org/generated/modeig/spd_learn.modules.LogEig.html
This layer maps SPD matrices to the tangent space at the identity through the matrix logarithm:

where is the eigendecomposition and applies the logarithm element-wise to the eigenvalues. If upper=True, the symmetric matrix is then vectorized by taking its upper triangular part.

ExpEig

Implemented at: https://spdlearn.org/generated/modeig/spd_learn.modules.ExpEig.html
This layer maps symmetric matrices from the tangent space back to the SPD manifold through the matrix exponential. It is the inverse of Deep Riemannian Network (DRN) > LogEig:

where is the eigendecomposition and applies the exponential element-wise to the eigenvalues.

>>> import torch
>>> from spd_learn.modules import LogEig, ExpEig
>>> X = torch.randn(2, 4, 4)
>>> X = X @ X.mT + 0.1 * torch.eye(4)  # Make SPD
>>> Y = LogEig(upper=True)(X)
>>> Y.shape
torch.Size([2, 10])
>>> Z = ExpEig(upper=True)(Y)
>>> Z.shape
torch.Size([2,4,4])

Model-Specific Layers

GaussAgg

Ref: @nguyenNeuralNetworkBased2019
GaussAgg aggregates a set of feature vectors into a single SPD matrix by fitting a Gaussian and using its Gaussian embedding. Given vectors ,

and the layer outputs

In ST-TS-HGR-NET this layer is used to summarize joint features over fingers, time windows, or sub-sequences while preserving first- and second-order statistics in SPD form. It takes a set of vectors as input and returns one SPD matrix per aggregation window.

SPDAgg

Ref: @nguyenNeuralNetworkBased2019
SPDAgg fuses several SPD matrices into a new SPD matrix through a learned weighted sum of bilinear transforms:

where each is an input SPD matrix and each is a learnable transformation. It is the final manifold-valued aggregation step in ST-TS-HGR-NET before LogEig and classification. In tensor form, it typically maps (batch, n_inputs, d_in, d_in) to (batch, d_out, d_out).

import torch
import torch.nn as nn
 
class SPDAgg(nn.Module):
    def __init__(self, n_inputs, d_in, d_out):
        super().__init__()
        self.weights = nn.Parameter(torch.randn(n_inputs, d_out, d_in))
 
    def forward(self, xs):  # xs: (batch, n_inputs, d_in, d_in)
        ys = [W @ X @ W.mT for X, W in zip(xs.unbind(dim=1), self.weights)]
        return torch.stack(ys, dim=0).sum(dim=0)

SPDConv

Ref: @zhangDeepManifoldtoManifoldTransforming2020
This layer is the DMT-Net analogue of convolution for multi-channel SPD maps. It is also natural to implement it as SPDConv2d. Instead of unconstrained kernels, each kernel is parameterized to be SPD:

with . Convolving an SPD map with SPD kernels yields another SPD map, so the layer performs local filtering while preserving manifold structure. In tensor form, it acts like a convolution from (batch, in_channels, D, D) to (batch, out_channels, D', D').

import torch
import torch.nn as nn
import torch.nn.functional as F
 
class SPDConv2d(nn.Module):
    def __init__(self, in_channels, out_channels, kernel_size, eps=1e-4):
        super().__init__()
        k = kernel_size
        self.V = nn.Parameter(torch.randn(out_channels, in_channels, k, k))
        self.eps = eps
 
    def forward(self, x):  # x: (batch, in_channels, D, D)
        k = self.V.size(-1)
        eye = torch.eye(k, device=x.device, dtype=x.dtype)
        W = self.V.transpose(-1, -2) @ self.V + self.eps * eye
        return F.conv2d(x, W)

SPD Activation

Ref: @zhangDeepManifoldtoManifoldTransforming2020
DMT-Net replaces eigendecomposition-based rectification with an SVD-free element-wise nonlinearity. The paper shows that applying functions such as , , or element-wise to an SPD matrix still yields an SPD matrix. In practice this gives a cheap nonlinear layer that stays on the manifold and is followed by Frobenius normalization to keep eigenvalues bounded.

SPD Recursive Layer

Ref: @zhangDeepManifoldtoManifoldTransforming2020
This layer is DMT-Net’s manifold-valued analogue of a GRU cell. Instead of scalar or vector gates, it updates a hidden SPD state with SPD-valued reset and update gates built from bilinear maps:

The main idea is that every operation is chosen to remain compatible with SPD geometry: the bilinear maps preserve SPD structure, the gate is SPD-valued, and the update is performed with Hadamard products rather than ordinary vector gating. Conceptually, this layer lets the network remember or forget temporal information without leaving the SPD manifold. In code, it usually maps (batch, time, channels, d, d) to (batch, time, channels, d, d).

The downside is that it is much more intricate than the other DRN layers, both mathematically and numerically. A full working Torch implementation is in SPD Recursive Layer.

Diagonalizing Layer

Ref: @zhangDeepManifoldtoManifoldTransforming2020
The diagonalizing layer maps a full SPD matrix to a diagonal SPD matrix

where is a positive element-wise activation. Because is diagonal, the subsequent matrix logarithm becomes an element-wise logarithm on the diagonal entries, avoiding the expensive eigendecomposition otherwise needed for LogEig-style computation.

SPD Matrix Pooling

Ref: @wangSymNetSimpleSymmetric2022
This layer generalizes pooling to SPD-valued features. SymNet studies three variants:

  • Tangent-space pooling: apply LogEig, perform standard max or mean pooling in the flat tangent space, then return to the manifold with ExpEig.
  • SPD max pooling: apply region-wise max pooling directly on SPD matrices. It is simple, but can distort geometry because the pooling window does not move along geodesics.
  • SPD mean pooling: replace arithmetic averaging by the Fréchet mean over the pooling region so the pooled output stays on the manifold.

AriM

Ref: @wangUSPDNetSPDManifold2023
AriM is the arithmetic mean layer used in U-SPDNet’s Log-Fusion-Exp skip block. After several SPD matrices are mapped to the logarithmic domain with LogEig, AriM computes their Euclidean barycenter

The fused matrix is then mapped back to the SPD manifold with ExpEig. It is cheaper than computing a true Riemannian barycenter, but it is only geometry-faithful after the log-domain relaxation. So the AriM layer itself is just the mean in the logarithmic domain, typically mapping (batch, n_mats, d, d) to (batch, d, d):

import torch
import torch.nn as nn
 
class AriM(nn.Module):
    def forward(self, xs):  # xs: (batch, n_mats, d, d) in the log domain
        return xs.mean(dim=1)
 
class LogEuclideanAriM(nn.Module):
    def __init__(self, log_layer, exp_layer):
        super().__init__()
        self.log_layer = log_layer
        self.arim = AriM()
        self.exp_layer = exp_layer
 
    def forward(self, xs):  # xs: (batch, n_mats, d, d) on the SPD manifold
        xs_log = torch.stack([self.log_layer(x) for x in xs.unbind(dim=1)], dim=1)
        mean_log = self.arim(xs_log)
        return self.exp_layer(mean_log)

Normalization and Regularization Layers

SPDBatchNormMeanVar

Implemented at: https://spdlearn.org/generated/batchnorm/spd_learn.modules.SPDBatchNormMeanVar.html
This layer is a batch-normalization analogue for SPD matrices. It first recenters a batch by parallel-transporting each sample from the batch Fréchet mean to the identity:

then applies dispersion normalization and an optional learnable re-biasing through a congruence transform with a learnable SPD matrix :

Compared with Euclidean batch normalization, the point is the same, but all operations are carried out in a manifold-consistent way.

SPDDSMBN

Ref: @koblerSPDDomainspecificBatch2022
Source code: https://github.com/rkobler/TSMNet
SPDDSMBN stands for SPD domain-specific momentum batch normalization. It combines domain-specific batch normalization with SPD momentum batch normalization by maintaining separate running Fréchet statistics for each domain:

In TSMNet this layer transforms domain-specific SPD inputs into domain-aligned outputs before tangent-space mapping, which is especially useful for cross-session and cross-subject EEG transfer.

Shrinkage

Available at SPDLearn: https://spdlearn.org/generated/regularization/spd_learn.modules.Shrinkage.html#spd_learn.modules.Shrinkage
Ref: @pengReducingDimensionalitySPD2023
Shrinkage is a common approach to compensate for the estimated bias, and covariance matrices are usually replaced by:

where is the shrinkage parameter and is the average eigenvalue of . After eigendecomposition , the shrunk matrix can be written as

Unlike ReEig, which only floors small eigenvalues, shrinkage moves all eigenvalues toward the spherical covariance . In practice this acts like a softer regularizer and is often more robust when covariance estimates come from small sample sets.

TraceNorm

Implemented at: https://spdlearn.org/generated/regularization/spd_learn.modules.TraceNorm.html
TraceNorm removes global scale from a covariance matrix by dividing it by its trace and optionally adding a small diagonal regularizer:

This produces unit-trace SPD matrices, preserving correlation structure while making representations less sensitive to overall signal amplitude differences across subjects, sessions, or trials.

SPDDropout

Implemented at: https://spdlearn.org/generated/dropout/spd_learn.modules.SPDDropout.html
SPDDropout is a structured dropout layer for SPD matrices. During training it randomly drops entire channels instead of individual entries, and fills the diagonal of dropped channels with a small so the result remains SPD. This makes it a geometry-aware regularizer rather than a naive element-wise dropout.

LogEuclideanResidual

Implemented at: https://spdlearn.org/generated/residual/spd_learn.modules.LogEuclideanResidual.html
This layer implements a residual or skip connection for SPD matrices in the Log-Euclidean geometry:

Instead of adding SPD matrices directly in the ambient Euclidean space, it performs the addition in the tangent space and maps back to the manifold. This gives an SPD-preserving residual block analogous to the skip connections used in standard ResNets.

PatchEmbeddingLayer

Implemented at: https://spdlearn.org/generated/utils/spd_learn.modules.PatchEmbeddingLayer.html
PatchEmbeddingLayer extracts temporal patches from an input tensor of shape (batch, channels, time) by unfolding the signal into a sequence of local windows. The output has shape (batch, n_patches, channels, patch_size), so it is mainly a preprocessing layer used before covariance estimation or patchwise SPD modeling rather than a manifold operation by itself.