Statistical comparisons of classifiers over multiple data sets (2006)
Open in webOpen in zoteroOpen pdf
1 Abstract
While methods for comparing two learning algorithms on a single data set have been scrutinized for quite some time already, the issue of statistical tests for comparisons of more algorithms on multiple data sets, which is even more essential to typical machine learning studies, has been all but ignored. This article reviews the current practice and then theoretically and empirically examines several suitable tests. Based on that, we recommend a set of simple, yet safe and robust non-parametric tests for statistical comparisons of classifiers: the Wilcoxon signed ranks test for comparison of two classifiers and the Friedman test with the corresponding post-hoc tests for comparison of more classifiers over multiple data sets. Results of the latter can also be neatly presented with the newly introduced CD (critical difference) diagrams.
2 NOTES
In this paper the authors discuss the use of statistical tests for comparison of classifiers over multiple datasets in two scenarios: two and more than two classifiers.
The main take away is:
- if the comparison is of two classifiers then use the Wilcoxon Signed-Ranks Tests.
- if the comparison is of more than two classifiers then the Friedman Test is to be used. Then, if the null-hypothesis is rejected, use:
- the Nemenyi Test if comparing all classifiers to each other, and
- the Bonferroni-Dunn Test or the Holm’s procedure, which is more powerful but don’t have the advantage of the Bonferroni-Dunn of being easily visualisable with Critical Difference (CD) Diagrams.
Then, the authors propose (at least it seems like they are proposing it), the Critical Difference (CD) Diagrams, which allows the visualization of the post-hoc tests when using multiple classifiers. And, more important, it doesn’t take much space on the page, being easily understandable.
Quotes:
Failure on Post-Hoc Test
“Sometimes the Friedman test reports a significant difference but the post-hoc test fails to detect it. This is due to the lower power of the latter. No other conclusions than that some algorithms do differ can be drawn in this case. In our experiments this has, however, occurred only in a few cases out of one thousand.” (Demsˇar and Demsar, 2006, p. 13)