Beyond the topics: how deep learning can improve the discriminability of probabilistic topic modelling.

Beyond the topics: how deep learning can improve the discriminability of probabilistic topic modelling.
复制标题

DOI:
10.7717/peerj-cs.252
复制
发表时间:
2020
期刊:
PeerJ. Computer science
影响因子:
--
通讯作者:
Awwad Shiekh Hasan B
Awwad Shiekh Hasan B
中科院分区:
其他
文献类型:
--
作者:
Al Moubayed N;McGough S;Awwad Shiekh Hasan B

文献摘要

相似文献

本文提出了一种判别方法来补充主题建模的无监督概率性质。该框架将每个文档主题的概率转换为类相关的深度学习模型,该模型提取适合分类的高度歧视性特征。然后将该框架用于最小特征工程的情感分析。该方法将情感分析问题从单词/文档领域转换为主题领域,使其对噪声更健壮,并结合了其他方式无法表示的复杂上下文信息。然后使用堆叠去噪自编码器(SDA)在最小假设下对每个情感主题之间的复杂关系进行建模。为了实现这一点,构建了不同的主题模型和每个情感极性的SDA,并添加了用于分类的决策层。该框架在样本量、类偏差和分类任务不同的基准数据集的综合收集上进行了测试。在不需要情感词典或过度设计的功能的情况下,实现了对艺术状态的重大改进。进行了进一步的分析来解释所观察到的精度的提高。
The article presents a discriminative approach to complement the unsupervised probabilistic nature of topic modelling. The framework transforms the probabilities of the topics per document into class-dependent deep learning models that extract highly discriminatory features suitable for classification. The framework is then used for sentiment analysis with minimum feature engineering. The approach transforms the sentiment analysis problem from the word/document domain to the topics domain making it more robust to noise and incorporating complex contextual information that are not represented otherwise. A stacked denoising autoencoder (SDA) is then used to model the complex relationship among the topics per sentiment with minimum assumptions. To achieve this, a distinct topic model and SDA per sentiment polarity is built with an additional decision layer for classification. The framework is tested on a comprehensive collection of benchmark datasets that vary in sample size, class bias and classification task. A significant improvement to the state of the art is achieved without the need for a sentiment lexica or over-engineered features. A further analysis is carried out to explain the observed improvement in accuracy.