TABBIE: Pretrained Representations of Tabular Data

TABBIE: Pretrained Representations of Tabular Data
复制标题

DOI:
10.18653/v1/2021.naacl-main.270
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer
H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer
中科院分区:
其他
文献类型:
--
作者:
H. Iida;Dung Ngoc Thai;Varun Manjunatha;Mohit Iyyer

文献摘要

被引文献

相似文献

现有的表格表示学习工作使用来自预训练语言模型(如BERT)的自监督目标函数对表格和相关文本进行联合建模。虽然这种联合预训练改进了涉及配对表格和文本的任务(例如,回答关于表格的问题),我们表明它在没有任何关联文本的表格上操作的任务上表现不佳(例如,填充缺失的单元格)。我们设计了一个简单的预训练目标(损坏的细胞检测),它只从表格数据中学习,并在一套基于表格的预测任务中达到最先进的水平。与竞争方法不同,我们的模型(TABBIE)提供了所有表子结构(单元格,行和列)的嵌入,并且它也需要更少的计算来训练。我们的模型的学习单元格,列和行表示的定性分析表明,它理解复杂的表语义和数值趋势。
Existing work on tabular representation-learning jointly models tables and associated text using self-supervised objective functions derived from pretrained language models such as BERT. While this joint pretraining improves tasks involving paired tables and text (e.g., answering questions about tables), we show that it underperforms on tasks that operate over tables without any associated text (e.g., populating missing cells). We devise a simple pretraining objective (corrupt cell detection) that learns exclusively from tabular data and reaches the state-of-the-art on a suite of table-based prediction tasks. Unlike competing approaches, our model (TABBIE) provides embeddings of all table substructures (cells, rows, and columns), and it also requires far less compute to train. A qualitative analysis of our model’s learned cell, column, and row representations shows that it understands complex table semantics and numerical trends.