PTab: Using the Pre-trained Language Model for Modeling Tabular Data

PTab: Using the Pre-trained Language Model for Modeling Tabular Data
复制标题

PTab:使用预先训练的语言模型对表格数据进行建模

DOI:
10.48550/arxiv.2209.08060
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Ledell Yu Wu
Ledell Yu Wu
中科院分区:
--
文献类型:
--
作者:
Guangyi Liu;Jie Yang;Ledell Yu Wu

文献摘要

参考文献

被引文献

相似文献

表格数据是信息时代的基础,已经得到了广泛的研究。最近的研究表明,基于神经的模型是有效的学习上下文表示的表格数据。有效的上下文表征的学习需要有意义的特征和大量的数据。然而,目前的方法往往无法正确地学习上下文表示的功能没有语义信息。此外,由于数据集之间的差异,很难通过混合表格数据集来扩大训练集。为了解决这些问题,我们提出了一个新的框架PTab,使用预训练的语言模型来建模表格数据。PTab通过三个阶段的处理来学习表格数据的上下文表示:模态转换(MT),掩蔽语言微调(MF)和分类微调(CF)。我们使用预训练模型(PTM)初始化我们的模型,该模型包含从大规模语言数据中学习的语义信息。因此,上下文表征可以有效地学习在微调阶段。此外,我们可以自然地混合文本化的表格数据来扩大训练集,以进一步提高表示学习。我们评估PTAB八个流行的表格分类数据集。实验结果表明,与最先进的基线(例如XGBoost)相比,我们的方法在监督设置中实现了更好的平均AUC得分,并且在半监督设置下优于对应的方法。我们目前的可视化结果表明,PTab具有良好的基于实例的可解释性。
Tabular data is the foundation of the information age and has been extensively studied. Recent studies show that neural-based models are effective in learning contextual representation for tabular data. The learning of an effective contextual representation requires meaningful features and a large amount of data. However, current methods often fail to properly learn a contextual representation from the features without semantic information. In addition, it's intractable to enlarge the training set through mixed tabular datasets due to the difference between datasets. To address these problems, we propose a novel framework PTab, using the Pre-trained language model to model Tabular data. PTab learns a contextual representation of tabular data through a three-stage processing: Modality Transformation(MT), Masked-Language Fine-tuning(MF), and Classification Fine-tuning(CF). We initialize our model with a pre-trained Model (PTM) which contains semantic information learned from the large-scale language data. Consequently, contextual representation can be learned effectively during the fine-tuning stages. In addition, we can naturally mix the textualized tabular data to enlarge the training set to further improve representation learning. We evaluate PTab on eight popular tabular classification datasets. Experimental results show that our method has achieved a better average AUC score in supervised settings compared to the state-of-the-art baselines(e.g. XGBoost), and outperforms counterpart methods under semi-supervised settings. We present visualization results that show PTab has well instance-based interpretability.
DOI: 10.18653/v1/2021.naacl-main.270
发表时间: 2021-05
期刊: --
影响因子: --
作者:
H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer
通讯作者: H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer