Contrastive Representation Learning for Electroencephalogram Classification (2020)
Open in webOpen in zoteroOpen pdf
1 Abstract
Interpreting and labeling human electroencephalogram (EEG) is a challenging task requiring years of medical training. We present a framework for learning representations from EEG signals via contrastive learning. By recombining channels from multi-channel recordings, we increase the number of samples quadratically per recording. We train a channel-wise feature extractor by extending the SimCLR framework to time-series data. We introduce a set of augmentations for EEG and study their efficacy on different classification tasks. We demonstrate that the learned features improve EEG classification and significantly reduce the amount of labeled data needed on three separate tasks: (1) Emotion Recognition (SEED), (2) Normal/Abnormal EEG classification (TUH), and (3) Sleep-stage scoring (SleepEDF). Our models show improved performance over previously reported supervised models on SEED and SleepEDF and self-supervised models on all three tasks.
2 NOTES
In this paper the authors present a self-supervised framework for learning representations of EEG signals. Their method is based on a channel encoder which used learning. That is, it takes a single channels as inputs, applies two augmentation methods (time shift, amplitude scale, dc shift, masking, band-stop filter, or additive noise), pass each through a channel encoder then a project, and at the end compute the contrastive loss. Therefore, the model should learn representations that are invariant under a set of augmentations. As the authors say, a contrastive learning algorithm learns representation that are maximally similar for augmented instances of the same data-point and minimally similar for different data-points. They tested two architectures for the encoder using RNNs and CNNs. For two datasets the CNN-based one achieved the best results and one for the RNN-based. However, the most interesting aspect of the whole work is that since it is all channel based, they can easily combine multiple datasets. They do so, by pre-training the model on those and achieve even better results. The tested augmentation, on the other hand, did not help the accuracy by themselves, only by the constrative loss.