Classification of covariance matrices for EEG: How to handle the low-rank case? (2022)
Open in webOpen in zoteroOpen pdf
1 Abstract
Brain-Computer Interfaces (BCIs) allow humans to communicate with a device solely via brain activity. Most BCIs rely on the classification of electroencephalography (EEG) signals to translate brain activity into commands. Existing state-of-the-art classification methods, based on Riemannian geometry, show their limits when the covariance matrices used to represent the signals become ill-conditioned. In this work, we consider a new Riemannian metric, the fixed-rank Wasserstein metric, which allows to take this low-rank structure into account. We compare this new metric with two other metrics, the Euclidean metric and the Affine-Invariant Riemannian metric, via two distance-based classification methods, Minimum Distance to the Mean and k-Nearest Neighbors. We also evaluate the impact of shrunk covariance estimators which are common estimators for alleviating the effect of ill-conditioning. Our results show that the Wasserstein metric achieves similar classification performance to the affine-invariant metric and uses less computation time and memory resources if the rank is wisely chosen. We also show that the new metric is practically independent of the matrix shrinkage and hence does not suffer from ill-conditioning. The Wasserstein metric is therefore of great interest for high-dimensional EEG signals and could outperform state-of-the-art Riemannian approaches in BCI.
2 NOTES
In this Gailly explores the use of the fixed-rank Wasserstein metric, a new Riemannian metric, which allows to take low-rank structure into account. She compares it against other metrics and also evaluates the impact of shrunk covariance estimators. The most interesting point of this is that the new metric is practically independent of the matrix shrinkage, which means it does not suffer from ill-conditioning. The text is brilliant, I love the section on covariance estimation and Riemannian geometry. Unfortunately she did not make a single image showing that by decreasing the rank the Wasserstein is better then the other, but for ranks < 10 (RMDM) and ranks < 15 (k-NN), it is always faster, which is an important point, given that in a figure bellow we see that by using all channels, a full rank covariance is much slower. On the other other hand, clearly show that it is indeed practically independent of shrinkage, unlike AIRM which varies a lot depending on the value.


