Neural decoding from stereotactic EEG: Accounting for electrode variability across subjects (2024)
Open in webOpen in zoteroOpen pdf
1 Abstract
Deep learning based neural decoding from stereotactic electroencephalography (sEEG) would likely benefit from scaling up both dataset and model size. To achieve this, combining data across multiple subjects is crucial. However, in sEEG cohorts, each subject has a variable number of electrodes placed at distinct locations in their brain, solely based on clinical needs. Such heterogeneity in electrode number/placement poses a significant challenge for data integration, since there is no clear correspondence of the neural activity recorded at distinct sites between individuals. Here we introduce seegnificant: a training framework and architecture that can be used to decode behavior across subjects using sEEG data. We tokenize the neural activity within electrodes using convolutions and extract long-term temporal dependencies between tokens using self-attention in the time dimension. The 3D location of each electrode is then mixed with the tokens, followed by another self-attention in the electrode dimension to extract effective spatiotemporal neural representations. Subject-specific heads are then used for downstream decoding tasks. Using this approach, we construct a multi-subject model trained on the combined data from 21 subjects performing a behavioral task. We demonstrate that our model is able to decode the trial-wise response time of the subjects during the behavioral task solely from neural data. We also show that the neural representations learned by pretraining our model across individuals can be transferred in a few-shot manner to new subjects. This work introduces a scalable approach towards sEEG data integration for multi-subject model training, paving the way for cross-subject generalization for sEEG decoding.
2 NOTES
In this paper the authors present an neural network for behavioral prediction using self-attention. The main difference against many other paper is the separation into two attention blocks: one for time and one for electrodes. Essentially, the authors construct the vector embedding by first applying a temporal CNN over the signals. This CNN takes takes a trial , apply a temporal convolution electrode by electrode with outputs for each one, but the main thing is that they also apply (after a batch normalization layer) an average pooling layer so that the and the number of electrodes , which might vary among trials, is reduced to an equal value. Therefore, the input signal becomes the vector . I understood the electrodes part wrongly, I thought every input would have the same but that is not true, an input with 5 electrodes is gonna have different vector then one with 8 electrodes, however, is gonna be the same for them. What happens is that this is simply done to enable flexibility, not uniformity, in the electrode dimension. The attention mechanism is the one that is gonna deal with and this is the one that doesn’t care about its number. Both attentions are gonna be quite similar, however, one is gonna arrange the electrode-latents and the other time-latents . Just as the transformer add positional encoding, in here the same is done, however, instead of sine and cosine the authors use the coordinates of the position from each electrode in the brain to construct the positional encoding. The procedure is not very simple, at least I did not understood it completely, but seems quite interesting and smart. These are the main ideas from the network. Actually, something else interesting is the way they fine-tune the model for new subjects. They first train the regression head of the new subject for 400 training, then for the remaining 600 all the model parameters were trained. They did not compare to many models, but that may be for lack of literature at the moment, so it is hard to evaluate their models, but their results are promising.
