TabTransformer: Tabular Data Modeling Using Contextual Embeddings

TabTransformer: Tabular Data Modeling Using Contextual Embeddings
复制标题

DOI:
--
复制
发表时间:
2020-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Xin Huang;A. Khetan;Milan Cvitkovic;Zohar S. Karnin
Xin Huang;A. Khetan;Milan Cvitkovic;Zohar S. Karnin
中科院分区:
其他
文献类型:
--
作者:
Xin Huang;A. Khetan;Milan Cvitkovic;Zohar S. Karnin

文献摘要

被引文献

相似文献

我们提出了 TabTransformer,一种用于监督和半监督学习的新型深度表格数据建模架构。 TabTransformer 是基于基于自注意力的 Transformer 构建的。 Transformer 层将分类特征的嵌入转换为鲁棒的上下文嵌入,以实现更高的预测精度。通过对 15 个公开数据集进行的广泛实验,我们表明 TabTransformer 在平均 AUC 上优于最先进的表格数据深度学习方法至少 1.0%,并且与基于树的集成模型的性能相匹配。此外,我们证明从 TabTransformer 学习的上下文嵌入对于缺失和噪声数据特征都具有高度鲁棒性,并提供更好的可解释性。最后,对于半监督设置,我们开发了一种无监督预训练程序来学习数据驱动的上下文嵌入,与最先进的方法相比,AUC 平均提升了 2.1%。
We propose TabTransformer, a novel deep tabular data modeling architecture for supervised and semi-supervised learning. The TabTransformer is built upon self-attention based Transformers. The Transformer layers transform the embeddings of categorical features into robust contextual embeddings to achieve higher prediction accuracy. Through extensive experiments on fifteen publicly available datasets, we show that the TabTransformer outperforms the state-of-the-art deep learning methods for tabular data by at least 1.0% on mean AUC, and matches the performance of tree-based ensemble models. Furthermore, we demonstrate that the contextual embeddings learned from TabTransformer are highly robust against both missing and noisy data features, and provide better interpretability. Lastly, for the semi-supervised setting we develop an unsupervised pre-training procedure to learn data-driven contextual embeddings, resulting in an average 2.1% AUC lift over the state-of-the-art methods.