1 DMT-Net
DMT-Net: Deep Manifold-to-Manifold Transforming Network
DMT-Net is an end-to-end manifold-to-manifold network that learns spatial and temporal features from SPD representations of skeleton sequences without requiring SVD in its manifold-valued layers.
Ref: @zhangDeepManifoldtoManifoldTransforming2020

Overview
DMT-Net introduces three manifold-preserving operations: a local SPDConv, an SVD-free SPD Activation, and a GRU-inspired SPD Recursive Layer. A Diagonalizing Layer then makes the final Log-Euclidean mapping inexpensive.
The major contributions of the proposed DMT-Net model include:
- A true local SPD convolutional layer, with the advantages of sparsity and efficiency, while SPDNet employed a bilinear operation;
- A non-linear activation layer to avoid SVD, and, again, SPDNet still need SVD for the ReEig layer;
- A recursive layer;
- An diagonalizing trick to bypass the high computation of SVD when applying log-Euclidean mapping, while SPDNet employed the standard log-Euclidean metric which needs SVD.
Architecture
For each frame, the joint coordinates are arranged as . With sequence-level mean joint position and the corresponding repeated matrix , the frame descriptor is
Sequences are divided into 12 subsequences, and one frame is randomly selected from each subsequence during downsampling. F-DMT-Net supplements these descriptors with a first-order branch that applies an ordinary GRU directly to joint trajectories; the DMT-Net and GRU features are concatenated before the shared fully connected and softmax layers.
The DMT-Net pipeline:
- Frame or subclip SPD descriptors.
- Two local SPD convolutional layers, each comprising SPD convolution, element-wise SPD activation, and Frobenius normalization.
- One SPD recursive layer over the temporal sequence of SPD maps.
- One diagonalizing layer.
- Vectorization of the nonzero diagonal elements and element-wise logarithm.
- One fully connected layer and softmax classification with cross-entropy loss.
SPDConv
For a single-channel SPD matrix and an SPD kernel , local convolution is
For a multi-channel SPD map, the -th output channel is
with the kernel constrained to be SPD through
Convolution with these kernels is proven to preserve positive definiteness.
SPD Activation
The SPD Activation applies , , or element-wise (not as matrix functions):
where denotes the Hadamard product or element-wise power. By the Schur product theorem the series terms are PSD/PD, so their sums remain SPD. This avoids the eigendecomposition used by SPDNet’s ReEig. Frobenius normalization is applied afterward to keep eigenvalues bounded.
SPD Recursive Layer
For channel , the SPD Recursive Layer is
with
Bilinear projections with positive diagonal bias, element-wise , and Hadamard products preserve the SPD property. The initial state is the zero matrix.
Diagonalizing Layer
The Diagonalizing Layer computes
where is a positive element-wise activation such as or . Because is diagonal and SPD, its logarithm is obtained by taking the scalar logarithm of each diagonal element, avoiding the SVD normally required by LogEig.
Model Parameters
- DMT-Net: two SPD convolutional layers, one SPD recursive layer, one diagonalizing layer, one fully connected layer, one softmax layer.
- NTU RGB+D, LSC, and HDM05 SPD kernels: followed by .
- Florence SPD kernels: followed by .
- The second convolutional layer has twice as many output channels as the first.
- F-DMT-Net first-order GRU: 128 hidden nodes.
- Number of temporal subsequences: 12.
- Spatial kernel size ablation on HDM05: , , ; best.
- Exact FC dimensions, SPD-recursive channel dimensions, parameter count, and : not reported.
Training Parameters
- Epochs: 500.
- Learning rate: 0.003.
- Optimizer: RMSPropOptimizer.
- Momentum: 0.9.
- Batch size: not reported.
- Objective: cross-entropy loss.
- Data augmentation: skeleton scaling factors in and rotations around the , , and axes in .
- F-DMT-Net: two-stage training - train DMT-Net and the first-order branch separately while fixing the other branch, then jointly optimize both with a relatively lower learning rate (value not reported).
Results
- NTU RGB+D: F-DMT-Net is 0.6 pp below GCA-LSTM under cross-subject evaluation and 1.6 pp above it under cross-view evaluation.
- LSC: under RCSub, F-DMT-Net’s precision/recall are ≈8/6 pp higher than P-LSTM.
- HDM05, protocol 2: DMT-Net 81.52%, almost 20 pp above SPDNet; F-DMT-Net 85.30%.
- Florence, leave-one-subject-out: F-DMT-Net 99.55%, ≈4 pp above the cited best; 8 of 9 actions recognized at 100%.
- Ablation on HDM05: DMT-Net ≈2.3 pp better than the version without the SPD recursive layer and ≈1.4 pp better than the version without nonlinear activation.