AnyPredict: Foundation Model for Tabular Prediction

AnyPredict: Foundation Model for Tabular Prediction
复制标题

DOI:
10.48550/arxiv.2305.12081
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Zifeng Wang;Chufan Gao;Cao Xiao;Jimeng Sun
Zifeng Wang;Chufan Gao;Cao Xiao;Jimeng Sun
中科院分区:
其他
文献类型:
--
作者:
Zifeng Wang;Chufan Gao;Cao Xiao;Jimeng Sun

文献摘要

被引文献

相似文献

基础模型在大量数据上进行了预训练,以便在许多下游任务中表现良好。他们在自然语言处理和计算机视觉方面取得了巨大的成功。尽管如此,这些模型在表格预测任务中的使用受到限制,主要障碍包括:(1)缺乏具有标准化标签的大规模和多样化的表格数据集;(2)跨领域的模式不匹配和预测目标异质性。本文提出了一种使用域内和广泛的域外数据集为表格预测基础模型(AnyPredict)构建大规模训练数据的方法。该方法使用数据引擎,该数据引擎利用大型语言模型(LLM)来合并表格样本,以克服具有不同模式的表之间的障碍,并使用“学习,注释和审计”管道将域外数据与目标任务对齐。扩展的训练数据使预训练的AnyPredict能够支持域中的每个表格数据集,而无需微调,从而比监督基线有了显着改进:它在7个患者结局预测数据集和3个试验结局预测数据集上的平均排名分别达到1.57和1.00。此外,AnyPredict表现出令人印象深刻的零射击性能:它比有监督的XGBoost模型高出8。9%和17。在两个预测任务中,平均分别为2%。
Foundation models are pre-trained on massive data to perform well across many downstream tasks. They have demonstrated significant success in natural language processing and computer vision. Nonetheless, the use of such models in tabular prediction tasks has been limited, with the main hurdles consisting of (1) the lack of large-scale and diverse tabular datasets with standardized labels and (2) the schema mismatch and predictive target heterogeneity across domains. This paper proposes a method for building training data at scale for tabular prediction foundation models ( AnyPredict ) using both in-domain and a wide range of out-domain datasets. The method uses a data engine that leverages large language models (LLMs) to consolidate tabular samples to overcome the barrier across tables with varying schema and align out-domain data with the target task using a “learn, annotate, and audit” pipeline. The expanded training data enables the pre-trained AnyPredict to support every tabular dataset in the domain without fine-tuning, resulting in significant improvements over supervised baselines: it reaches an average ranking of 1.57 and 1.00 on 7 patient outcome prediction datasets and 3 trial outcome prediction datasets, respectively. In addition, AnyPredict exhibits impressive zero-shot performances: it outperforms supervised XGBoost models by 8 . 9% and 17 . 2% on average in two prediction tasks, respectively.