A practical guide to applying machine learning to infant EEG data (2022)
Open in webOpen in zoteroOpen pdf
1 Abstract
Electroencephalography (EEG) has been widely adopted by the developmental cognitive neuroscience community, but the application of machine learning (ML) in this domain lags behind adult EEG studies. Applying ML to infant data is particularly challenging due to the low number of trials, low signal-to-noise ratio, high inter-subject variability, and high inter-trial variability. Here, we provide a step-by-step tutorial on how to apply ML to classify cognitive states in infants. We describe the type of brain attributes that are widely used for EEG classification and also introduce a Riemannian geometry based approach for deriving connectivity estimates that account for intertrial and inter-subject variability. We present pipelines for learning classifiers using trials from a single infant and from multiple infants, and demonstrate the application of these pipelines on a standard infant EEG dataset of forty 12-month-old infants collected under an auditory oddball paradigm. While we classify perceptual states induced by frequent versus rare stimuli, the presented pipelines can be easily adapted for other experimental designs and stimuli using the associated code that we have made publicly available.
2 NOTES
In this paper the authors present a introduction and guideline on infant EEG data classification. There are two mainly relevant things: the infant EEG data discussion and the statistical tests. First things first, I wasn’t aware but this is a real problem, and a very much complex one. In this case the data came from 12-month-old infants, which means that their brain is still very much in development, therefore the variability in their responses is much greater then in adults. This variation is due to myelination and neuronal responses selectivity, and they continue even in adulthood, which might explain the variability. Another issue is that an infants have a limited attention span, so they can’t keep focused on a task (especially an tiresome EEG task). About their experiment, it was about rare vs frequent stimuli, whether infants could distinguish between two phonemes: /ra/ vs /la/. Their results were bad, barely over chance. The second interesting things was the statistical test, where they applied McNemar’s test of each model against each other, looking at test folds and seeing which model was better in that particular case. This is a great way to compare if there was a real variation and maybe even to see if there are some easy/hard samples to classify and what can be extracted from them. Overall a great paper, many open questions, possible model classification on the dataset and possible experiments construction on this idea.