Improving Mean Covariance Matrix Estimation by Minimizing Within-class Dissimilarities Using Asymmetry of Kullback-Leibler Divergence in MI-Based BCI (2024)

Open in webOpen in zoteroOpen pdf

1 Abstract

Brain-computer interface (BCI) systems create a direct communication channel from the brain to an output device by translating brain activity into a sequence of control commands. Feature extraction algorithms play a key role in the performance of the BCI systems. Some feature extraction algorithms use the mean (class) covariance matrices containing information about the dispersion and variability of the data and are estimated from empirical data. Common spatial pattern is one of the most widespread feature extraction algorithms in this field that needs a robust and well-estimated mean covariance matrix for each class. To this end, we have proposed a novel method to estimate a mean covariance matrix for each class of data which has the minimum within-class dissimilarity based on the asymmetry of Kullback–Leibler divergence using a linear combination of its two asymmetric directions. We applied the proposed algorithm to three public electroencephalographic datasets (BCI competition IV dataset I, BCI competition IV dataset IIa, and BCI competition III dataset IVa) to validate its effectiveness, compared with several methods in the same category of covariance matrix improvement as the benchmark. The proposed method showed at least 8% superiority in grand mean classification accuracy of the three datasets over the benchmark methods, while the accuracies of most individuals are improved. The experimental results reveal that the proposed algorithm provides statistically significant improvements in classification results by reducing within-class variations. In addition, it led to more compact and separable features and may cause more robust BCI applications.

2 NOTES

In this paper the authors combine the asymmetry of Kullback-Leibler divergence to estimate a mean covariance matrix for each class, and apply it to CSP which led to their method called asymmetric KL-based covariance matrix in CSP (AKLCSP). The main reason for using CSP is that it is highly dependent on the wellness of the estimation of the Mean Covariance Matrix (MCM). Since there are two parts that need to be solved (both sides of the KL divergence) for all trials (and all classes), it involves an optimization procedure. The proof on how to find the solution of this is long, but the idea is that there emerges an arithmetic and a harmonic mean, and then the proposed method is a trade-off between each but that benefits from both. Albeit not a very easy proof, their algorithm is straightforward, as shown bellow. They only compared its performance against other CSP algorithms (which is fair), and in all cases (two classes only) they achieved a higher mean average then all other. Overall, it is a good paper, however, would like to see if it could be extended to other methods other than CSP.

\begin{algorithm}
\caption{The pseudocode of the AKLCSP mean covariance matrix computation}
 
\begin{algorithmic}
\Input Filtered (8-30 Hz) train data of one class with $n$ trials: $X = \{x_1, \ldots, x_n\}, \ x \in \mathbb{R}^{L \times M}$; Hyper-parameter $\alpha$, $0 \leq \alpha \leq 1$\\
\Output $\bar{C}$
\State Extract the shrinkage hyper-parameter $\gamma = \text{DLCSPauto}(X)$;
\State $A = \text{zeros}(M)$; $B = \text{zeros}(M)$; $I = \text{eye}(M)$;
\For{$i = 1 : n$}
    \State Make the EEG signal zero mean: $x_i \leftarrow x_i - \text{mean}(x_i)$;
    \State Compute covariance matrix: $C_i = \frac{x_i^T x_i}{\text{tr}(x_i^T x_i)}$;
    \State $P = \text{mean}(\text{diag}(C_i))$;
    \State Shrink the covariance matrix of trial $i$: $C_i \leftarrow (1 - \gamma)C_i + \gamma P I$;
    \State $A \leftarrow A + C_i$;
    \State $B \leftarrow B + \text{pinv}(C_i)$;
\EndFor
\State Apply the share of the sum of covariance matrices: $A \leftarrow (1 - \alpha)A$;
\State Apply the share of the sum of the inversed covariance matrices: $B \leftarrow \alpha B$;
\State Compute the needed variables: $\rho = n(1 - 2\alpha)$; $\rho_p = -\rho$;
\If{$0 \leq \alpha < 0.5$}
    \State Compute the GEVD of $A^{0.5}BA^{0.5}$: $[U, \Lambda] = \text{eig}(A^{0.5}BA^{0.5})$;
    \State Compute the mean covariance matrix: $\bar{C} = \text{inv}(A^{-0.5} U ((\Lambda + \frac{\rho^2}{4} I)^{0.5} + \frac{\rho}{2}) U' A^{-0.5})$;
\ElsIf{$0.5 \leq \alpha \leq 1$}
    \State Compute the GEVD of $B^{0.5}AB^{0.5}$: $[U, \Lambda] = \text{eig}(B^{0.5}AB^{0.5})$;
    \State Compute the mean covariance matrix: $\bar{C} = \text{inv}(B^{-0.5} U ((\Lambda + \frac{\rho_p^2}{4} I)^{0.5} + \frac{\rho_p}{2}) U' B^{-0.5})$;
\EndIf
\end{algorithmic}
\end{algorithm}