SuperTML: Two-Dimensional Word Embedding for the Precognition on Structured Tabular Data

SuperTML: Two-Dimensional Word Embedding for the Precognition on Structured Tabular Data
复制标题

DOI:
10.1109/cvprw.2019.00360
复制
发表时间:
2019-02
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Baohua Sun;Lin Yang;Wenhan Zhang;Michael Lin;Patrick Dong;Charles Young;Jason Dong
Baohua Sun;Lin Yang;Wenhan Zhang;Michael Lin;Patrick Dong;Charles Young;Jason Dong
中科院分区:
其他
文献类型:
--
作者:
Baohua Sun;Lin Yang;Wenhan Zhang;Michael Lin;Patrick Dong;Charles Young;Jason Dong

文献摘要

被引文献

相似文献

根据 Kaggle ML 和 DS 调查,表格数据是行业中最常用的数据形式。梯度提升树、支持向量机、随机森林和逻辑回归通常用于表格数据的分类任务。使用分类嵌入的 DNN 模型也应用于此任务,但迄今为止所有尝试都使用一维嵌入。最近使用二维词嵌入的超级字符方法在文本分类任务中取得了最先进的结果,展示了这种新方法的前景。在本文中,我们提出了SuperTML方法,它借用超级字符方法和二维嵌入的思想来解决表格数据的分类问题。对于表格数据的每个输入,特征首先像图像一样投影到二维嵌入中,然后将该图像输入到微调的二维 CNN 模型中进行分类。所提出的 SuperTML 方法自动处理表格数据中的分类数据和缺失值,无需预处理为数值。模型性能的比较是在 Kaggle 平台上最大、最活跃的竞赛之一以及 UCI 机器学习存储库中最受欢迎的三个数据集上进行的。实验结果表明,所提出的 SuperTML 方法在大型和小型数据集上都取得了最先进的结果。
Tabular data is the most commonly used form of data in industry according to a Kaggle ML and DS Survey. Gradient Boosting Trees, Support Vector Machine, Random Forest, and Logistic Regression are typically used for classification tasks on tabular data. DNN models using categorical embeddings are also applied in this task, but all attempts thus far have used one-dimensional embeddings. The recent work of Super Characters method using two-dimensional word embeddings achieved state-of-the-art results in text classification tasks, showcasing the promise of this new approach. In this paper, we propose the SuperTML method, which borrows the idea of Super Characters method and two-dimensional embeddings to address the problem of classification on tabular data. For each input of tabular data, the features are first projected into two-dimensional embeddings like an image, and then this image is fed into fine-tuned two-dimensional CNN models for classification. The proposed SuperTML method handles the categorical data and missing values in tabular data automatically, without any need to pre-process into numerical values. Comparisons of model performance are conducted on one of the largest and most active competitions on the Kaggle platform, as well as on the top three most popular data sets in the UCI Machine Learning Repository. Experimental results have shown that the proposed SuperTML method have achieved state-of-the-art results on both large and small datasets.