ARMAGNAC: A New Parametric Batch Normalization Layer for SPDNet Architecture ()
Open in webOpen in zoteroOpen pdf
1 Abstract
This paper introduces ARMAGNAC (bAtch noRMalisation pArametric Geometric harmoNic ArithmetiC), a new batch normalization layer for deep learning architectures on Symmetric Positive Definite (SPD) matrices. Unlike standard existing approaches that rely on computationally expensive Fre´chet means with iterative eigenvalue decompositions, ARMAGNAC leverages a parametric mean that automatically adapts between simple arithmetic and harmonic means. This learnable parametrization enables the layer to adjust to the specific characteristics of the problem at hand while maintaining computational efficiency through explicit formulas of the backpropagation gradients. Experimental results demonstrate that ARMAGNAC correctly learn the parameter to use the most appropriate mean between arithmetic and harmonic on simulated data, and achieves superior performance results on real-data classification tasks.
2 NOTES
In this paper the authors propose a Batch Normalization method for SPD matrices that don’t rely on the Karcher Mean (Flow Algorithm). Instead they propose to use the Arithmetic and Geometric means to estimate a global mean that can normalize the data. That is, during inference phase, each SPD matrix is normalized using as where is the (learned) bias and (albeit not in the above formula but on the tables below) is the momentum.

Where and , the variables and are initialized as the identity
However, for their paper, the authors let the be so that the model can learn it during backpropagation. Thus, the resulting matrix is not exactly half arithmetic and half harmonic: and then to update the during learning it follows the GAH equation above but also replacing with :
They even describe how the backward pass is calculated:

The fact that it doesn’t require the Karcher Mean (Flow Algorithm), reduces the number of eigenvalue decompositions required to run the batch normalization. In their results, they mention that training on the Rices90 dataset, training with the geometric mean (from @brooksRiemannianBatchNormalization) took 9 times longer than training with ARMAGNAC. The accuracy results however are weird, it is 2% better on one dataset but almost 2% worse on another (considering for both ARMAGNAC vs Geometric Mean).
Overall, very interesting and I will try applying it to EEG data soon.