The N2 corpus: A semantically annotated collection of Islamist extremist stories

The N2 corpus: A semantically annotated collection of Islamist extremist stories
复制标题

N2 语料库:带有语义注释的伊斯兰极端主义故事集

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Steven R. Corman
Steven R. Corman
中科院分区:
--
文献类型:
--
作者:
Mark A. Finlayson;Jeffry R. Halverson;Steven R. Corman

文献摘要

被引文献

相似文献

我们描述了一种新的语言资源--N2(叙事网络)语料库。语料库在三个重要方面是独一无二的。首先,语料库中的每个文本都是一个故事,这与其他语言资源不同,其他语言资源可能包含故事或类似故事的文本,但不是专门策划为只包含故事。其次,语料库的统一主题是与伊斯兰极端分子相关的材料,由他们编制或经常被他们引用。第三,语料库中的每个文本都经过了14层句法和语义的注释,包括:指称表达和共指;事件、时间表达和时间关系;语义角色;以及词义。在无法使用分析器进行高质量自动注解的情况下,需要手动对层进行双重注解,并由训练有素的注释员进行评判。该语料库由100篇文本和42480个单词组成。大多数文本最初是阿拉伯语的,但都是英文翻译的。我们解释了构建语料库的动机,文本的选择过程,语料库本身的详细内容,选择注释层的理由,以及注解程序。
We describe the N2 (Narrative Networks) Corpus, a new language resource. The corpus is unique in three important ways. First, every text in the corpus is a story, which is in contrast to other language resources that may contain stories or story-like texts, but are not specifically curated to contain only stories. Second, the unifying theme of the corpus is material relevant to Islamist Extremists, having been produced by or often referenced by them. Third, every text in the corpus has been annotated for 14 layers of syntax and semantics, including: referring expressions and co-reference; events, time expressions, and temporal relationships; semantic roles; and word senses. In cases where analyzers were not available to do high-quality automatic annotations, layers were manually double-annotated and adjudicated by trained annotators. The corpus comprises 100 texts and 42,480 words. Most of the texts were originally in Arabic but all are provided in English translation. We explain the motivation for constructing the corpus, the process for selecting the texts, the detailed contents of the corpus itself, the rationale behind the choice of annotation layers, and the annotation procedure.