Probabilistic Relational Supervised Topic Modelling using Word Embeddings

Probabilistic Relational Supervised Topic Modelling using Word Embeddings
复制标题

DOI:
10.1109/bigdata.2018.8622326
复制
发表时间:
2018-12
期刊:
2018 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Jabir Alshehabi Al-Ani;Maria Fasli
Jabir Alshehabi Al-Ani;Maria Fasli
中科院分区:
其他
文献类型:
--
作者:
Jabir Alshehabi Al-Ani;Maria Fasli

文献摘要

被引文献

相似文献

语言变化的速度越来越快,影响了文本处理的许多应用程序和算法。自然语言处理(NLP)的研究人员一直在努力寻找更通用的解决方案,以科普不断变化的情况。当应用于来自社交媒体的短文本时,这更具挑战性。此外,越来越多的社交媒体对语言的发展和使用产生了重大影响。我们的工作是出于开发NLP技术的需要,这些技术可以科普社交媒体中使用的简短非正式文本,以及每天在社交媒体上上传的大量文本数据。在本文中,我们描述了一种新的短文本主题建模方法,使用词嵌入,并考虑到任何非正式的社会媒体文本中的单词,目的是解决在混乱的文本中减少噪音的挑战。本文提出了一种基于词频-逆文档词频(TF-IDF)的新算法,命名为词频-逆上下文词频(TF-ICTF)。TF-ICTF依赖于单词和上下文之间相对于时间的概率关系。我们的实验工作显示出对其他国家的最先进的方法有希望的结果。
The increasing pace of change in languages affects many applications and algorithms for text processing. Researchers in Natural Language Processing (NLP) have been striving for more generalized solutions that can cope with continuous change. This is even more challenging when applied on short text emanating from social media. Furthermore, increasingly social media have been casting a major influence on both the development and the use of language. Our work is motivated by the need to develop NLP techniques that can cope with short informal text as used in social media alongside the massive proliferation of textual data uploaded daily on social media. In this paper, we describe a novel approach for Short Text Topic Modelling using word embeddings and taking into account any informality of words in the social media text with the aim of addressing the challenge of reducing noise in messy text. We present a new algorithm derived from the Term Frequency -Inverse Document Frequency (TF-IDF), named Term Frequency - Inverse Context Term Frequency (TF-ICTF). TF-ICTF relies on a probabilistic relation between words and context with respect to time. Our experimental work shows promising results against other state-of-the-art methods.