PSDNorm: Test-time temporal normalization for deep learning on EEG signals (2025)

Open in webOpen in zoteroOpen pdf

1 Abstract

Distribution shift poses a significant challenge in machine learning, particularly in biomedical applications such as EEG signals collected across different subjects, institutions, and recording devices. While existing normalization layers, Batch-Norm, LayerNorm and InstanceNorm, help address distribution shifts, they fail to capture the temporal dependencies inherent in temporal signals. In this paper, we propose PSDNorm, a layer that leverages Monge mapping and temporal context to normalize feature maps in deep learning models. Notably, the proposed method operates as a test-time domain adaptation technique, addressing distribution shifts without additional training. Evaluations on 10 sleep staging datasets using the U-Time model demonstrate that PSDNorm achieves state-of-the-art performance at test time on datasets not seen during training while being 4x more data-efficient than the best baseline. Additionally, PSDNorm provides a significant improvement in robustness, achieving markedly higher F1 scores for the 20% hardest subjects.

2 NOTES

In here the use of normalization is on the power spectral density (PSD) of each feature map, from which they are normalized as is done in the SPD Batch Norm, using the barycenter. Essentially, there are three main operations: 1) PSD estimation, 2) running Riemannian barycenter update, and 3) Monge mapping application. The main difference here comes from another work from the main author (Théo Gnassounou), which is the -Monge Mapping, based on the classical Monge mapping between Gaussian distributions. Its okay.

The main interesting thing here is how they did their training/evaluations, where they chose 10 datasets, setting 7 for training and 3 for testing. During training they also divide the datasets into training, validation, and testing sets with a 64%/16%/20% subject split. And, to evaluate model robustness, they also use the F1@20% score, which measures performance on the 20% of subjects with the lowest F1 scores under the BatchNorm baseline.