Regularized estimation of large covariance matrices (2008)

Open in webOpen in zoteroOpen pdf

1 Abstract

This paper considers estimating a covariance matrix of p variables from n observations by either banding or tapering the sample covariance matrix, or estimating a banded version of the inverse of the covariance. We show that these estimates are consistent in the operator norm as long as (log p)/n→0, and obtain explicit rates. The results are uniform over some fairly natural well-conditioned families of covariance matrices. We also introduce an analogue of the Gaussian white noise model and show that if the population covariance is embeddable in that model and well-conditioned, then the banded approximations produce consistent estimates of the eigenvalues and associated eigenvectors of the covariance matrix. The results can be extended to smooth versions of banding and to non-Gaussian distributions with sufficiently short tails. A resampling approach is proposed for choosing the banding parameter in practice. This approach is illustrated numerically on both simulated and real data.

2 NOTES

This paper is very mathematical, but the main idea is to show that by banding a covariance matrix they can achieve similar or better efficiency. To band it is necessary to define by how much, which is determined by the parameter . The authors propose a method to find it using cross-validation. They use it in a synthetic and a real dataset. In both experiments they got a lower loss compared to the SCM estimator, performing (almost) exactly as well as the oracle estimator. Even when they achieved a better result, and much better when , which is the relevant case for EEG.