1 DMT-Net

DMT-Net: Deep Manifold-to-Manifold Transforming Network

DMT-Net is an end-to-end manifold-to-manifold network that learns spatial and temporal features from SPD representations of skeleton sequences without requiring SVD in its manifold-valued layers.

Ref: @zhangDeepManifoldtoManifoldTransforming2020

center

Overview

DMT-Net introduces three manifold-preserving operations: a local SPDConv, an SVD-free SPD Activation, and a GRU-inspired SPD Recursive Layer. A Diagonalizing Layer then makes the final Log-Euclidean mapping inexpensive.

The major contributions of the proposed DMT-Net model include:

  1. A true local SPD convolutional layer, with the advantages of sparsity and efficiency, while SPDNet employed a bilinear operation;
  2. A non-linear activation layer to avoid SVD, and, again, SPDNet still need SVD for the ReEig layer;
  3. A recursive layer;
  4. An diagonalizing trick to bypass the high computation of SVD when applying log-Euclidean mapping, while SPDNet employed the standard log-Euclidean metric which needs SVD.

Architecture

For each frame, the joint coordinates are arranged as . With sequence-level mean joint position and the corresponding repeated matrix , the frame descriptor is

Sequences are divided into 12 subsequences, and one frame is randomly selected from each subsequence during downsampling. F-DMT-Net supplements these descriptors with a first-order branch that applies an ordinary GRU directly to joint trajectories; the DMT-Net and GRU features are concatenated before the shared fully connected and softmax layers.

The DMT-Net pipeline:

  1. Frame or subclip SPD descriptors.
  2. Two local SPD convolutional layers, each comprising SPD convolution, element-wise SPD activation, and Frobenius normalization.
  3. One SPD recursive layer over the temporal sequence of SPD maps.
  4. One diagonalizing layer.
  5. Vectorization of the nonzero diagonal elements and element-wise logarithm.
  6. One fully connected layer and softmax classification with cross-entropy loss.

SPDConv

For a single-channel SPD matrix and an SPD kernel , local convolution is

For a multi-channel SPD map, the -th output channel is

with the kernel constrained to be SPD through

Convolution with these kernels is proven to preserve positive definiteness.

SPD Activation

The SPD Activation applies , , or element-wise (not as matrix functions):

where denotes the Hadamard product or element-wise power. By the Schur product theorem the series terms are PSD/PD, so their sums remain SPD. This avoids the eigendecomposition used by SPDNet’s ReEig. Frobenius normalization is applied afterward to keep eigenvalues bounded.

SPD Recursive Layer

For channel , the SPD Recursive Layer is

with

Bilinear projections with positive diagonal bias, element-wise , and Hadamard products preserve the SPD property. The initial state is the zero matrix.

Diagonalizing Layer

The Diagonalizing Layer computes

where is a positive element-wise activation such as or . Because is diagonal and SPD, its logarithm is obtained by taking the scalar logarithm of each diagonal element, avoiding the SVD normally required by LogEig.

Model Parameters

  • DMT-Net: two SPD convolutional layers, one SPD recursive layer, one diagonalizing layer, one fully connected layer, one softmax layer.
  • NTU RGB+D, LSC, and HDM05 SPD kernels: followed by .
  • Florence SPD kernels: followed by .
  • The second convolutional layer has twice as many output channels as the first.
  • F-DMT-Net first-order GRU: 128 hidden nodes.
  • Number of temporal subsequences: 12.
  • Spatial kernel size ablation on HDM05: , , ; best.
  • Exact FC dimensions, SPD-recursive channel dimensions, parameter count, and : not reported.

Training Parameters

  • Epochs: 500.
  • Learning rate: 0.003.
  • Optimizer: RMSPropOptimizer.
  • Momentum: 0.9.
  • Batch size: not reported.
  • Objective: cross-entropy loss.
  • Data augmentation: skeleton scaling factors in and rotations around the , , and axes in .
  • F-DMT-Net: two-stage training - train DMT-Net and the first-order branch separately while fixing the other branch, then jointly optimize both with a relatively lower learning rate (value not reported).

Results

  • NTU RGB+D: F-DMT-Net is 0.6 pp below GCA-LSTM under cross-subject evaluation and 1.6 pp above it under cross-view evaluation.
  • LSC: under RCSub, F-DMT-Net’s precision/recall are ≈8/6 pp higher than P-LSTM.
  • HDM05, protocol 2: DMT-Net 81.52%, almost 20 pp above SPDNet; F-DMT-Net 85.30%.
  • Florence, leave-one-subject-out: F-DMT-Net 99.55%, ≈4 pp above the cited best; 8 of 9 actions recognized at 100%.
  • Ablation on HDM05: DMT-Net ≈2.3 pp better than the version without the SPD recursive layer and ≈1.4 pp better than the version without nonlinear activation.