VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain

VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain
复制标题

VIME:将自监督和半监督学习的成功扩展到表格领域

DOI:
--
复制
发表时间:
2020
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
M. Schaar
M. Schaar
中科院分区:
--
文献类型:
--
作者:
Jinsung Yoon;Yao Zhang;James Jordon;M. Schaar

文献摘要

被引文献

相似文献

自监督和半监督学习框架在训练图像和语言领域有限的标记数据的机器学习模型方面取得了显着的进展。fi。这些方法严重依赖于领域数据集中的独特结构(如图像中的空间关系或语言中的语义关系)。它们不适用于与图像和语言数据不具有相同显式结构的一般表格数据。在本文中,我们通过提出新的针对表格数据的自监督和半监督学习框架来弥补这一差距,我们统称为VIME(Value InputationfiMASK Estiments)。除了用于自监督学习的重建借口任务外,我们还创建了一种新的借口任务,即从受破坏的表格数据估计掩码向量。我们还介绍了一种用于自监督和半监督学习框架的新的表格数据扩充方法。在实验中,我们在来自不同应用领域的多个表格数据集中对所提出的框架进行了评估,例如基因组学和临床数据。与现有的基准方法相比,VIME的性能超过了最先进的性能。
Self-and semi-supervised learning frameworks have made significant progress in training machine learning models with limited labeled data in image and language domains. These methods heavily rely on the unique structure in the domain datasets (such as spatial relationships in images or semantic relationships in language). They are not adaptable to general tabular data which does not have the same explicit structure as image and language data. In this paper, we fill this gap by proposing novel self-and semi-supervised learning frameworks for tabular data, which we refer to collectively as VIME (Value Imputation and Mask Estimation). We create a novel pretext task of estimating mask vectors from corrupted tabular data in addition to the reconstruction pretext task for self-supervised learning. We also introduce a novel tabular data augmentation method for self-and semi-supervised learning frameworks. In experiments, we evaluate the proposed framework in multiple tabular datasets from various application domains, such as genomics and clinical data. VIME exceeds state-of-the-art performance in comparison to the existing baseline methods.