On the Power of Pre-Trained Text Representations: Models and Applications in Text Mining

On the Power of Pre-Trained Text Representations: Models and Applications in Text Mining
复制标题

DOI:
10.1145/3447548.3470810
复制
发表时间:
2021-08
期刊:
Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Yu Meng;Jiaxin Huang;Yu Zhang;Jiawei Han
Yu Meng;Jiaxin Huang;Yu Zhang;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Yu Meng;Jiaxin Huang;Yu Zhang;Jiawei Han

文献摘要

相似文献

近年来,文本表示学习在广泛的文本挖掘任务中取得了巨大的成功。早期的词嵌入学习方法将词表示为固定的低维向量,以捕获其语义。这样学习到的词嵌入被用作任务特定模型的输入特征。最近,预训练语言模型(PLMs)通过在大规模文本语料库上预训练基于transformer的神经模型来学习通用语言表示,已经彻底改变了自然语言处理(NLP)领域。这种预训练的表示对通用语言特征进行编码,这些特征可以转移到几乎任何与文本相关的应用程序中。plm在许多应用程序中优于以前的特定于任务的模型,因为它们只需要在目标语料库上进行微调,而不是从头开始训练。在本教程中,我们介绍了预训练文本嵌入和语言模型的最新进展,以及它们在广泛的文本挖掘任务中的应用。具体来说,我们首先概述了一组最近开发的自监督和弱监督文本嵌入方法以及作为下游任务基础的预训练语言模型。然后,我们提出了几种基于预训练文本嵌入和语言模型的新方法,用于各种文本挖掘应用,如主题发现和文本分类。我们专注于弱监督、领域独立、语言不可知、有效和可扩展的方法,用于从大规模文本语料库中挖掘和发现结构化知识。最后,我们用真实世界的数据集演示了预训练的文本表示如何帮助减轻人工注释负担,并促进自动、准确和高效的文本分析。
Recent years have witnessed the enormous success of text representation learning in a wide range of text mining tasks. Earlier word embedding learning approaches represent words as fixed low-dimensional vectors to capture their semantics. The word embeddings so learned are used as the input features of task-specific models. Recently, pre-trained language models (PLMs), which learn universal language representations via pre-training Transformer-based neural models on large-scale text corpora, have revolutionized the natural language processing (NLP) field. Such pre-trained representations encode generic linguistic features that can be transferred to almost any text-related applications. PLMs outperform previous task-specific models in many applications as they only need to be fine-tuned on the target corpus instead of being trained from scratch. In this tutorial, we introduce recent advances in pre-trained text embeddings and language models, as well as their applications to a wide range of text mining tasks. Specifically, we first overview a set of recently developed self-supervised and weakly-supervised text embedding methods and pre-trained language models that serve as the fundamentals for downstream tasks. We then present several new methods based on pre-trained text embeddings and language models for various text mining applications such as topic discovery and text classification. We focus on methods that are weakly-supervised, domain-independent, language-agnostic, effective and scalable for mining and discovering structured knowledge from large-scale text corpora. Finally, we demonstrate with real-world datasets how pre-trained text representations help mitigate the human annotation burden and facilitate automatic, accurate and efficient text analyses.