Creating an Annotated Corpus for Sentiment Analysis of German Product Reviews

Creating an Annotated Corpus for Sentiment Analysis of German Product Reviews
复制标题

创建用于德国产品评论情感分析的注释语料库

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
R. Messerschmidt
R. Messerschmidt
中科院分区:
--
文献类型:
--
作者:
K. Boland;Andias Wira;R. Messerschmidt

文献摘要

被引文献

相似文献

带注释数据的可用性是开发用于情感分析的机器学习算法的重要前提。然而,由于手动标记大型数据集既耗时又昂贵,因此可用的数据集很少,并且大多数数据集代表了非常狭窄领域的小样本,例如电影评论或某种产品类型的评论。此外,许多带注释的数据集仅适用于英文文本。然而,如果只有来自一个特定领域的训练数据可用,或者测试语料库中混合了特定领域,则输入数据集的不同特征对情感分析算法性能的影响仍不清楚。因此,我们为各种产品类型的德国产品评论引入了一个新的数据集,并调查该特定领域(不同产品类型)中即使很小的差异是否已经表现出不同的特征,例如关于情感标注的难度。该语料库的注释为未来类似语料库的增强注释以及将我们的注释扩展到本质上不同领域的语料库奠定了基础。然后,这些将用于研究不同语料库特征对不同情感分析算法的影响,并作为应用机器学习方法对德语文本进行句子情感分析的基础。 6 GESIS-技术报告2013|15
The availability of annotated data is an important prerequisite for the development of machine learning algorithms for sentiment analysis. However, as manually labeling large datasets is time-consuming and expensive, few datasets are available and most of them represent a small sample of a very narrow domain, e.g. movie reviews or reviews of a certain product type. Additionally, many annotated datasets are available for English texts only. However, the influence of different characteristics of the input dataset on the performance of algorithms for sentiment analysis remains unclear if only training data from one specific domain is available or if specific domains are mixed in the test corpus. We therefore introduce a new dataset for German product reviews of various product types and investigate whether even small variances in this specific domain (different product types) already exhibit different characteristics, e.g. with regard to the difficulty of sentiment annotation. The annotation of this corpus lays the basis for future enhanced annotations of similar corpora and for the extension of our annotations to corpora of inherently different domains. These will then serve to investigate the influence of different corpus characteristics on different algorithms for sentiment analysis and as a basis to apply machine learning methods for sentence-wise sentiment analysis for German texts. 6 GESIS-Technical Report 2013|15