Riemannian Multinomial Logistics Regression for SPD Neural Networks (2024)

Open in webOpen in zoteroOpen pdf

1 Abstract

Deep neural networks for learning Symmetric Positive Definite (SPD) matrices are gaining increasing attention in machine learning. Despite the significant progress, most existing SPD networks use traditional Euclidean classifiers on an approximated space rather than intrinsic classifiers that accurately capture the geometry of SPD manifolds. Inspired by Hyperbolic Neural Networks (HNNs), we propose Riemannian Multinomial Logistics Regression (RMLR) for the classification layers in SPD networks. We introduce a unified framework for building Riemannian classifiers under the metrics pulled back from the Euclidean space, and showcase our framework under the parameterized LogEuclidean Metric (LEM) and Log-Cholesky Metric (LCM). Besides, our framework offers a novel intrinsic explanation for the most popular LogEig classifier in existing SPD networks. The effectiveness of our method is demonstrated in three applications: radar recognition, human action recognition, and electroencephalography (EEG) classification. The code is available at https://github.com/GitZH-Chen/SPDMLR.git.

2 NOTES

The main idea of this paper is that instead of using, for instance, LogEig and then a fully-connected layer, they use a Riemannian Multinomial Logistics Regression (MLR) (RMLR) for classification which is used to build the SPD MLR model. The metric used is the Pullback Euclidean Metric (PEM), and since both the Log Euclidean Metric (LEM) and Log Cholesky Metric (LCM) are PEMs these are the evaluated ones. In order to do so they modify the Euclidean MLR by rewriting it with an SPD hyperplane which under any geometrically complete Riemannian metric is a regular submanifold of SPD manifolds. Therefore they obtain an general form for SPD MLR under a given PEM, then apply both and in it. It is interesting to note that when for , the SPD MLR is very similar to the LogEig MLR. The best result, however, for the Hinss2021 dataset, as seen in the following table, was for the case when , that is, for the deformed LCM.