Combining datasets to increase the number of samples and improve model fitting (2023)
Open in webOpen in zoteroOpen pdf
1 Abstract
For many use cases, combining information from different datasets can be of interest to improve a machine learning model’s performance, especially when the number of samples from at least one of the datasets is small. An additional challenge in such cases is that the features from these datasets are not identical, even though there are some commonly shared features among the datasets. To tackle this, we propose a novel framework called Combine datasets based on Imputation (ComImp). In addition, we propose PCA-ComImp, a variant of ComImp that utilizes Principle Component Analysis (PCA), where dimension reduction is conducted before combining datasets. This is useful when the datasets have a large number of features that are not shared across them. Furthermore, our framework can also be utilized for data preprocessing by imputing missing data, i.e., filling in the missing entries while combining different datasets. To illustrate the performance and practicability of the proposed methods and their potential usages, we conduct experiments for various tasks (regression, classification) and for different data types (tabular data, time series data) when the datasets to be combined have missing data. We also investigate how the devised methods can be used with transfer learning to provide even further model training improvement. Our results indicate that can provide extra improvement when being used in combination with transfer learning.
2 NOTES
In this paper the authors propose a method to easily combine datasets, in particular ones related to the same context but with distinct features, called ComImp. The working of their algorithm is rather simple, it stacks the datasets inputs and outputs, for the missing values they use some imputation method. These method are abundant, the authors mentions kNN, MICE, missForest, Bayesian methods and DIMV. For their experiments, however, they used soft-impute, but don’t explain it at all, just say they used it. Their results are also not outstanding, but it is interesting to see that such a simple algorithms works, even in their EEG example, which is a quite complex one. Overall, its an interesting technique, not spectacular, but might serve as a foundation for better ones.
