Text classification in shipping industry using unsupervised models and Transformer based supervised models

Text classification in shipping industry using unsupervised models and Transformer based supervised models
复制标题

DOI:
10.48550/arxiv.2212.12407
复制
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Yingyi Xie;Dongping Song
Yingyi Xie;Dongping Song
中科院分区:
其他
文献类型:
--
作者:
Yingyi Xie;Dongping Song

文献摘要

相似文献

在特定上下文中获取标记数据可能既昂贵又耗时。虽然采用了不同的算法,包括无监督学习、半监督学习、自学习,但文本分类的性能因上下文而异。在缺乏标记数据集的情况下,我们提出了一种新的简单的无监督文本分类模型,使用标准国际贸易分类(SITC)代码对国际航运业货物内容进行分类。我们的方法源于使用预训练的手套词嵌入来表示单词,并使用余弦相似度来找到最可能的标签。为了比较无监督文本分类模型和监督分类模型,我们还应用了几个Transformer模型对货物内容进行分类。由于缺乏训练数据,我们使用SITC数字代码和相应的文本描述作为训练数据。利用少量人工标注的货物内容数据,对无监督分类和基于Transformer的监督分类的分类性能进行了评价。对比表明,即使在训练数据集的大小增加30%之后,无监督分类也明显优于基于Transformer的监督分类。缺乏训练数据是阻碍深度学习模型(如Transformers)成功实际应用的关键瓶颈。在训练数据稀缺的情况下,无监督分类为文本分类提供了另一种高效的方法。
Obtaining labelled data in a particular context could be expensive and time consuming. Although different algorithms, including unsupervised learning, semi-supervised learning, self-learning have been adopted, the performance of text classification varies with context. Given the lack of labelled dataset, we proposed a novel and simple unsupervised text classification model to classify cargo content in international shipping industry using the Standard International Trade Classification (SITC) codes. Our method stems from representing words using pretrained Glove Word Embeddings and finding the most likely label using Cosine Similarity. To compare unsupervised text classification model with supervised classification, we also applied several Transformer models to classify cargo content. Due to lack of training data, the SITC numerical codes and the corresponding textual descriptions were used as training data. A small number of manually labelled cargo content data was used to evaluate the classification performances of the unsupervised classification and the Transformer based supervised classification. The comparison reveals that unsupervised classification significantly outperforms Transformer based supervised classification even after increasing the size of the training dataset by 30%. Lacking training data is a key bottleneck that prohibits deep learning models (such as Transformers) from successful practical applications. Unsupervised classification can provide an alternative efficient and effective method to classify text when there is scarce training data.