Sentiment analysis with knowledge resource and NLP tools

Sentiment analysis with knowledge resource and NLP tools
复制标题

利用知识资源和NLP工具进行情感分析

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
S. Ananiadou
S. Ananiadou
中科院分区:
--
文献类型:
--
作者:
S. Piao;Yoshimasa Tsuruoka;S. Ananiadou

文献摘要

被引文献

相似文献

自动情感分析是人类语言技术(HLT)和文本挖掘中的一个重要且具有挑战性的课题,在社会科学中有许多应用。近年来,人们在这个问题上付出了很大的努力。许多关于这个主题的出版作品都使用了各种机器学习技术。在我们的工作中,我们研究了通过结合知识资源和自然语言处理技术自动识别文本情感倾向的可行性,作为基于机器学习的方法的补充。我们的方法涉及资源和工具,包括主观性词典(Wilson et al., 2005)、一套NLP工具和加权算法。我们开发了一个情感分析工具与网络演示(http://text0.mib.man.ac.uk:8080/opminpackage/opinion_analysis)。该工具已与其他文本挖掘工具一起使用,以支持社会科学家从报纸文章中进行框架分析。在本文中,我们描述了我们的系统,并报告了我们在句子和文档级别上识别情感倾向的功能的评估。我们使用多视角问答(MPQA)语料库(Wiebe et al., 2005)和2000篇人工分类的影评(Pang and Lee, 2004)作为测试数据。作为评估指标,我们使用了f-score,这是一种结合了准确率和召回率的性能指标,范围在0(最差表现)和1(最佳表现)之间。在评估中,我们的情感分析工具获得了令人鼓舞的结果,句子级情感分析的f值为0.6012至0.8333,文档级情感分析的平均f值为0.7196。
Automatic sentiment analysis is an important and challenging topic in Human Language Technology (HLT) and text mining, with several applications for social sciences. Over recent years, much effort has been devoted to this subject. Many published works on this subject employ various machine learning techniques. In our work, we investigate the feasibility of automatically identifying text sentiment orientation by combining knowledge resources and NLP techniques, as a complementary method to those based on machine learning. Our approach involves resources and tools including a subjectivity lexicon (Wilson et al., 2005), a set of NLP tools and weighting algorithms. We developed a sentiment analysis tool with a web demonstrator (http://text0.mib.man.ac.uk:8080/opminpackage/opinion_analysis). This tool has been used jointly with other text mining tools to support social scientists in frame analysis from newspaper articles. In this paper, we describe our system and report on our evaluation of the functionality of identifying sentiment orientation at the sentence and document levels. We used the Multi-Perspective Question Answering (MPQA) Corpus (Wiebe et al., 2005) and a collection of 2,000 manually classified film reviews (Pang and Lee, 2004) as the test data. As evaluation measure, we used f-score, a performance measure that combines precision and recall and ranges between 0 (worst performance) and 1 (best performance). In the evaluation, our sentiment analysis tool obtained encouraging results, producing f-scores ranging from 0.6012 to 0.8333 for sentence-level sentiment analysis and an average f-score of 0.7196 for document-level sentiment analysis.