Eigen-value: Efficient domain-robust data valuation via eigenvalue-based approach (2025)
Open in webOpen in zoteroOpen pdf
1 Abstract
Data valuation has become central in the era of data-centric AI. It drives efficient training pipelines and enables objective pricing in data markets by assigning a numeric value to each data point. Most existing data valuation methods estimate the effect of removing individual data points by evaluating changes in model validation performance under in-distribution (ID) settings, as opposed to out-of-distribution (OOD) scenarios where data follow different patterns. Since ID and OOD data behave differently, data valuation methods based on ID loss often fail to generalize to OOD settings, particularly when the validation set contains no OOD data. Furthermore, although OOD-aware methods exist, they involve heavy computational costs, which hinder practical deployment. To address these challenges, we introduce \emph{Eigen-Value} (EV), a plug-and-play data valuation framework for OOD robustness that uses only an ID data subset, including during validation. EV provides a new spectral approximation of domain discrepancy, which is the gap of loss between ID and OOD using ratios of eigenvalues of ID data’s covariance matrix. EV then estimates the marginal contribution of each data point to this discrepancy via perturbation theory, alleviating the computational burden. Subsequently, EV plugs into ID loss-based methods by adding an EV term without any additional training loop. We demonstrate that EV achieves improved OOD robustness and stable value rankings across real-world datasets, while remaining computationally lightweight. These results indicate that EV is practical for large-scale settings with domain shift, offering an efficient path to OOD-robust data valuation.
2 NOTES
The authors mentions some interesting ideas about using the eigenvalues from the distribution to compare against a single sample and identity if it can be considered an in-distribution (ID) or out-of-distribution (OOD) sample. The main thing they mention is that removing a single data point from a dataset can be viewed as applying a perturbation to the covariance matrix. Also, that if a matrix is modified by a small term such that , one can express the eigenvalues and eigenvectors of as expansions of . However, their results are not good, so it is hard to care about the paper.