Ranking and combining multiple predictors without labeled data (2014)

Open in webOpen in zoteroOpen pdf

1 Abstract

In a broad range of classification and decision-making problems, one is given the advice or predictions of several classifiers, of unknown reliability, over multiple questions or queries. This scenario is different from the standard supervised setting, where each classifier’s accuracy can be assessed using available labeled data, and raises two questions: Given only the predictions of several classifiers over a large set of unlabeled test data, is it possible to (i) reliably rank them and (ii) construct a metaclassifier more accurate than most classifiers in the ensemble? Here we present a spectral approach to address these questions. First, assuming conditional independence between classifiers, we show that the off-diagonal entries of their covariance matrix correspond to a rank-one matrix. Moreover, the classifiers can be ranked using the leading eigenvector of this covariance matrix, because its entries are proportional to their balanced accuracies. Second, via a linear approximation to the maximum likelihood estimator, we derive the Spectral Meta-Learner (SML), an unsupervised ensemble classifier whose weights are equal to these eigenvector entries. On both simulated and real data, SML typically achieves a higher accuracy than most classifiers in the ensemble and can provide a better starting point than majority voting for estimating the maximum likelihood solution. Furthermore, SML is robust to the presence of small malicious groups of classifiers designed to veer the ensemble prediction away from the (unknown) ground truth.

2 NOTES

This is the paper that proposed Spectral Meta-Learner. The idea is that instead of using majority voting the authors use a meta-learner, where they construct a covariance matrix with the outputs of the models (for each input). The algorithm seems to be pretty simple, of implements can be seen in https://github.com/learn-ensemble/PY-SUMMA/blob/master/pySUMMA/sml.py and is indeed pretty straightforward. Also, the method works very well as can be seen by their results and in @liTTIMETesttimeInformation2024. However, even if I don’t get it all, the paper most interesting part is that the authors derive all the equations and prove that the method is good and works. There is nothing much to be said thought, the details of the proof are in the paper and details of the method are in the Spectral Meta-Learner note.