Structured Regularization for SPD Optimization with Side Information (2024)
Open in webOpen in zoteroOpen pdf
1 Abstract
Matrix-valued optimization tasks, including those involving symmetric positive definite (SPD) matrices, arise in a wide range of applications in machine learning, data science and statistics. Classically, such problems are solved via constrained Euclidean optimization, where the domain is viewed as a Eu-clidean space and the structure of the matrices (e.g., positive definiteness) enters as constraints. More recently, geometric approaches that leverage parametrizations of the problem as unconstrained tasks on the corresponding matrix manifold have been proposed. While they exhibit algorithmic benefits in many settings, they cannot directly handle additional constraints, such as side information on the solution. A remedy comes in the form of constrained Riemannian optimization methods, notably, Rie-mannian Frank-Wolfe and Projected Gradient Descent. However, both algorithms require potentially expensive subroutines that can introduce computational bottlenecks in practise. To mitigate these shortcomings, we propose a structured regularization approach based on symmetric gauge functions. On the example of computing optimistic likelihoods, we show that the regularizer preserves crucial structure in the objective, including geodesic convexity. This allows for solving the regularized problem with a fast unconstrained method with global optimality certificate. We demonstrate the effectiveness of our approach in numerical experiments and through theoretical analysis.
2 NOTES
This a complex paper, to me currently, that involves the optimization in the Riemannian manifold using some side information (some constraint on the manifold). The idea that I got it that this side information is similar to clamping the inputs from a GAN generator, which means you are not passing directly information but is limiting it enough so that it is information after all. This is particularly interesting if you think about using a mean covariance matrix from a class, for instance, as a restriction on the optimization, so that you know this is not the one you want to be at but it should go far from it.