Deep Manifold-to-Manifold Transforming Network for Skeleton-Based Action Recognition (2020)

Open in webOpen in zoteroOpen pdf

1 Abstract

In this paper, we will investigate skeleton-based action recognition by employing high-order statistics feature and first-order statistics feature, where the high-order statistics feature is characterized by symmetric positive definite (SPD) matrices. Noting that SPD matrices are theoretically embedded on Riemannian manifolds, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which can make SPD matrices flow from one Riemannian manifold to another one for facilitating the action recognition task. To learn discriminative SPD features from both spatial and temporal dependencies, we propose a neural network model with three novel layers on manifolds: i.e., (1) the local SPD convolutional layer, (2) the non-linear SPD activation layer, and (3) the Riemannian-preserved recursive layer. The SPD property is preserved through all layers without the singular value decomposition (SVD) operation, which has to be conducted in the existing methods with expensive computation cost. Furthermore, a diagonalizing SPD layer is designed to efficiently calculate the final metric for the classification task. Finally, DMT-Net is further fused with a first order layer to capture temporal evolution information. To evaluate our proposed method, we conduct extensive experiments on the task of action recognition, where the input signals are represented as SPD matrices. The experimental results demonstrate that the proposed method is competitive over state-of-the-art methods.

2 NOTES

This paper is mainly composed of three new layers: Local SPD Convolutional Layer, SPD Recursive Layer and a Diagonalizing layer. With the exception of the last one, the other two are generalization from their respective Euclidean layers in traditional Neural Networks. The Convolutional layer in here is exactly the same as usual, having only the need to restrain the kernel as SPD matrices. They also propose defining a matrix from which they obtain the kernels of the convolution as , so does not need the SPD, but using something like Geoopt one could define simply as SPD and let the Riemannian ADAM deal with it. Still, it is an interesting way to not rely on specific constraints. For the Recursive Layer something similar is done, but they put the constraint on the trainable parameters, so they are SPD, but do not change the logic from the usual recursive layer. The diagonalizing layer is meant to replace the LogEig layer, in here the authors propose a transformation to reduce from the Riemannian manifold onto a that is on the Euclidean manifold, where is some non-linear activation function (they mention that or can be used instead of ReEig without changing anything). However, in doing so the dimension grows such that becomes , so with 22 channels a covariance matrices would have dimensions and after the diagonalizaton it would become . Therefore is increases the dimension a lot, however, it avoids the use of SVD operation, and they mention using only the non-zero elements of the diagonal matrix, that is, the matrices become a -dimensional vector. Very interesting, I had not paid attention to this before…will run some tests and come here later to say if it works. Overall, I like the paper, mainly their proof on the Convolutional and Recursive layers.