SQUINKY! A Corpus of Sentence-level Formality, Informativeness, and Implicature

SQUINKY! A Corpus of Sentence-level Formality, Informativeness, and Implicature
复制标题

斯奎克!

DOI:
--
复制
发表时间:
2015
期刊:
arXiv.org
影响因子:
--
通讯作者:
Shibamouli Lahiri
Shibamouli Lahiri
中科院分区:
--
文献类型:
--
作者:
Shibamouli Lahiri

文献摘要

被引文献

相似文献

我们介绍了一个7032个句子的语料库,由人类注释者对正式性、信息量和含义进行了1-7级的评分。使用Amazon Mechanical Turk对语料库进行注释。通过比较两个MTurk实验的平均评分,以及在更受控的环境中进行的与飞行员注释(关于句子形式)的相关性,来检验所获得判断的可靠性。尽管标注任务具有主观性和固有的难度,但平均评分之间的相关性非常令人鼓舞,特别是在正式性和信息性方面。我们进一步探讨了三个语言变量之间的相关性,体裁分级的体裁变化和体裁内部的相关性,与自动风格评分的兼容性,以及文档在风格方面的句子组成。到目前为止,我们的语料库是最大的句子级标注语料库发布的形式,信息和含义。
We introduce a corpus of 7,032 sentences rated by human annotators for formality, informativeness, and implicature on a 1-7 scale. The corpus was annotated using Amazon Mechanical Turk. Reliability in the obtained judgments was examined by comparing mean ratings across two MTurk experiments, and correlation with pilot annotations (on sentence formality) conducted in a more controlled setting. Despite the subjectivity and inherent difficulty of the annotation task, correlations between mean ratings were quite encouraging, especially on formality and informativeness. We further explored correlation between the three linguistic variables, genre-wise variation of ratings and correlations within genres, compatibility with automatic stylistic scoring, and sentential make-up of a document in terms of style. To date, our corpus is the largest sentence-level annotated corpus released for formality, informativeness, and implicature.