Rank over Class: The Untapped Potential of Ranking in Natural Language Processing

Rank over Class: The Untapped Potential of Ranking in Natural Language Processing
复制标题

DOI:
10.1109/bigdata52589.2021.9671386
复制
发表时间:
2020-09
期刊:
2021 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Amir Atapour-Abarghouei;Stephen Bonner;A. Mcgough
Amir Atapour-Abarghouei;Stephen Bonner;A. Mcgough
中科院分区:
其他
文献类型:
--
作者:
Amir Atapour-Abarghouei;Stephen Bonner;A. Mcgough

文献摘要

被引文献

相似文献

长期以来,文本分类一直是自然语言处理(NLP)中的主要内容,其应用程序跨越了不同的领域,如情感分析、推荐系统和垃圾邮件检测。有了这样一个强大的解决方案,人们往往会忍不住把它作为解决所有NLP问题的首选工具,因为当你拿着锤子时,一切看起来都像钉子。然而,我们在这里认为,目前使用分类解决的许多任务实际上被硬塞进了分类模型中,如果我们将它们作为一个排名问题来处理,我们不仅可以改进模型,而且我们可以获得更好的性能。我们提出了一种新的端到端排序方法,该方法由一个转换器网络组成,该网络负责产生一对文本序列的表示,这些文本序列又被传递到上下文聚合网络,该网络输出排序分数,用于根据一些相关性概念来确定序列的排序。我们在公开可用的数据集上进行了大量的实验,并研究了排序在经常使用分类解决的问题中的应用。在一个严重倾斜的情感分析数据集上的实验中,将排序结果转换为分类标签比最新的文本分类产生了大约22%的改进,展示了在某些场景下文本排序相对于文本分类的有效性。
Text classification has long been a staple within Natural Language Processing (NLP) with applications spanning across diverse areas such as sentiment analysis, recommender systems and spam detection. With such a powerful solution, it is often tempting to use it as the go-to tool for all NLP problems since when you are holding a hammer, everything looks like a nail. However, we argue here that many tasks which are currently addressed using classification are in fact being shoehorned into a classification mould and that if we instead address them as a ranking problem, we not only improve the model, but we achieve better performance. We propose a novel end-to-end ranking approach consisting of a Transformer network responsible for producing representations for a pair of text sequences, which are in turn passed into a context aggregating network outputting ranking scores used to determine an ordering to the sequences based on some notion of relevance. We perform numerous experiments on publicly-available datasets and investigate the applications of ranking in problems often solved using classification. In an experiment on a heavily- skewed sentiment analysis dataset, converting ranking results to classification labels yields an approximately 22% improvement over state-of-the-art text classification, demonstrating the efficacy of text ranking over text classification in certain scenarios.