Pseudo-label : The simple and efficient semi-supervised learning method for deep neural networks ()
Open in webOpen in zoteroOpen pdf
1 Abstract
We propose the simple and efficient method of semi-supervised learning for deep neural networks. Basically, the proposed network is trained in a supervised fashion with labeled and unlabeled data simultaneously. For unlabeled data, Pseudo-Label s, just picking up the class which has the maximum predicted probability, are used as if they were true labels. This is in effect equivalent to Entropy Regularization. It favors a low-density separation between classes, a commonly assumed prior for semi-supervised learning. With Denoising Auto-Encoder and Dropout, this simple method outperforms conventional methods for semi-supervised learning with very small labeled data on the MNIST handwritten digit dataset.
2 NOTES
This is probably one of the first papers that came with this idea of using the own model to generate the labels for unlabeled data. If not the first then certainly one of the first to do so in such a simple way. The authors simple propose that the model is fine-timed using simultaneously training and testing (pseudo-labeled) data, having a loss function which is a combination of the two, which a factor determining their proportion. That is, at first should be small so that the model relies mostly on the training data, as the test pseudo-labels are gonna be wrong, then, as time goes by this moves to a more balanced case so that unlabeled data can be beneficial. The bottom diagram demonstrates how this works. It is a very simple method, but is still being applied to new papers. The reason why I decided to read it is that I popped up on a couple of recent papers (2024~2025). Pretty good and quite straightforward, as is expected since it was published for a workshop.
