Overview of Opinion Analysis Pilot Task at NTCIR-6

Overview of Opinion Analysis Pilot Task at NTCIR-6
复制标题

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
Yohei Seki;D. Evans;Lun-Wei Ku;Hsin-Hsi Chen;N. Kando;Chin-Yew Lin
Yohei Seki;D. Evans;Lun-Wei Ku;Hsin-Hsi Chen;N. Kando;Chin-Yew Lin
中科院分区:
其他
文献类型:
--
作者:
Yohei Seki;D. Evans;Lun-Wei Ku;Hsin-Hsi Chen;N. Kando;Chin-Yew Lin

文献摘要

被引文献

相似文献

摘要本文概述了2006年至2007年在第六届NT-CIR研讨会上的意见分析试点任务。我们用中文、日文和英文分别创建了32、30和28个主题(11,907、15,279和8,379句)的测试集。使用这个测试集,我们进行了意见提取子任务。该子任务从四个角度进行定义:(a)意见句判断,(B)意见保持器提取,(c)相关句判断,(d)极性判断。14名参赛者提交了21个运行结果,组织者提交了5个结果。并对参与意见抽取子任务的群体的评价结果进行了分析。关键词:意见抽取,意见保持器,相关性,极性,和NTCIR。1引言本文概述了2006年至2007年第六届NTCIR研讨会[4](NTCIR-6意见)上的意见分析试点任务[5]。这是第一次尝试产生一个多语言的测试集合,用于评估NTCIR的意见提取。意见和情感分析最近在自然语言处理研究界受到了很多关注[2,9,7]。随着网络上广泛的信息来源,以及促进用户生成内容的面向社会社区的网站的快速增长,商业和政府方面都对自动分析和监测网络上流行的态度趋势产生了进一步的兴趣。因此,自动检测表达意见的句子的兴趣([12]等),表达式的极性([13]等),目标和意见持有者([1]等)在研究界受到了越来越多的关注。应用包括跟踪对商业产品、政府政策的反应和意见,跟踪潜在政治丑闻的博客等等。在第六届NTCIR研讨会上,介绍了一个新的试点任务,用于意见分析。试点任务有三种语言的轨道:中文,英文和日文。在本文中,我们对测试集、任务设计和使用测试集的评估结果进行了概述,我们认为,由于语料库的可比性,这个试点任务为扩展跨语言的有观点文本分析研究提供了一个独特的机会。在跨语言信息检索任务中,基于人工相关性判断,对文档进行了仔细的选择,确保了三种语言的高质量语料库都是相关的。在第2节中,我们解释了
Abstract This paper describes an overview of the OpinionAnalysis Pilot Task from 2006 to 2007 at the Sixth NT-CIR Workshop. We created test collection for 32, 30,and 28 topics (11,907, 15,279, and 8,379 sentences)in Chinese, Japanese and English. Using this test col-lection, we conducted opinion extraction subtask. Thesubtask was defined from four perspectives: (a) opin-ionated sentence judgment, (b) opinion holder extrac-tion, (c) relevance sentence judgment, and (d) polarityjudgment. 21 run results were submitted by 14 partici-pants with five results submitted by the organizers. Weshow the evaluation results of the groups participatingin opinion extraction subtask. Keywords: Opinion Extraction, Opinion Holder, Rel-evance, Polarity, and NTCIR. 1 Introduction This paper describes an overview of the OpinionAnalysis Pilot Task [5] from 2006 to 2007 at the SixthNTCIR Workshop [4] (NTCIR-6 Opinion). This wasthe first effort to produce a multi-lingual test collectionfor evaluating opinion extraction at NTCIR.Opinion and sentiment analysis has been receivinga lot of attention in the natural language processing re-search community recently [2, 9, 7]. With the broadrange of information sources available on the web,and rapid increase in the uptake of social community-oriented websites that foster user-generated contentthere has been further interest by both commercial andgovernmental parties in trying to automatically ana-lyze and monitor the tide of prevalent attitudes on theweb. As a result, interest in automatically detectingsentences in which an opinion is expressed ([12] etc.),the polarity of the expression ([13] etc.), targets, andopinion holders ([1] etc.) has been receiving more at-tention in the research community. Applications in-clude tracking response to and opinions about com-mercial products, governmental policies, tracking blogentries for potential political scandals and so on.In the Sixth NTCIR Workshop, a new pilot task foropinion analysis has been introduced. The pilot taskhas tracks in three languages: Chinese, English, andJapanese. In this paper, we present an overview of thetest collection, task design, and evaluation results us-ing the test collection across the Chinese, Japanese,and English data.We believe that this pilot task presents a unique op-portunity to expand the study of opinionated text anal-ysis across languages due to the comparable nature ofthe corpus. The documents have been carefully se-lected based on the manual relevance judgments as-signed in a cross-lingual Information Retrieval task,ensuring a high quality corpus that is relevant in allthree languages.This paper is organized as follows. In Section 2,we explain the task design for the