A Lie Group Approach to Riemannian Batch Normalization (2024)

Open in web
Open in zotero
Open pdf

1 Abstract

Manifold-valued measurements exist in numerous applications within computer vision and machine learning. Recent studies have extended Deep Neural Networks (DNNs) to manifolds, and concomitantly, normalization techniques have also been adapted to several manifolds, referred to as Riemannian normalization. Nonetheless, most of the existing Riemannian normalization methods have been derived in an ad hoc manner and only apply to specific manifolds. This paper establishes a unified framework for Riemannian Batch Normalization (RBN) techniques on Lie groups. Our framework offers the theoretical guarantee of controlling both the Riemannian mean and variance. Empirically, we focus on Symmetric Positive Definite (SPD) manifolds, which possess three distinct types of Lie group structures. Using the deformation concept, we generalize the existing Lie groups on SPD manifolds into three families of parameterized Lie groups. Specific normalization layers induced by these Lie groups are then proposed for SPD neural networks. We demonstrate the effectiveness of our approach through three sets of experiments: radar recognition, human action recognition, and electroencephalography (EEG) classification. The code is available at https://github.com/GitZH-Chen/LieBN.git.

2 NOTES

The main proposal of this paper is a general framework for Riemannian Batch Normalization (RBN) over Lie Groups (LieBN), which they refer as LieBN. Their approach allows the control of first- and second-order statistics, a property that none of the other proposal could do, all while working with three deformed Lie Groups (AIM, LEM, LCM). To do so they had to redefine Gaussian distribution, centering, biasing, and scaling to Lie groups. Once this was done, they basically changed the simple RBN by Brooks with these new definitions, which leads to a three (instead of two) key operations: centering to the neutral element; scaling the dispersion; and, biasing towards a learned parameter. The principal result is that this is a generic approach, unlike the RBN by Brooks et al. which worked only with the AIM metric. Another interesting result is that due to the difference in the metrics, each produces a certain result, and different metrics works with different datasets. Further investigation in this direction might be interesting, since there does not seem to be a clear motive for one working better than another one. Also, SPDDSMBN and AIM are much slower since they require a bunch of eigendecompositions, which LEM and LCM do not (they only require a Cholesky decomposition if I’m not wrong).