Building a Diverse Document Leads Corpus Annotated with Semantic Relations

Building a Diverse Document Leads Corpus Annotated with Semantic Relations
复制标题

构建多样化的文档导致语料库带有语义关系注释

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
S. Kurohashi
S. Kurohashi
中科院分区:
--
文献类型:
--
作者:
Masatsugu Hangyo;Daisuke Kawahara;S. Kurohashi

文献摘要

被引文献

相似文献

近年来,语义分析在自然语言处理中得到了积极的研究。对于语义分析的研究来说,具有语义标注的语料库是必不可少的。尽管有这样的语料库被标注在报纸文章上,但也有各种各样的体裁和风格,包括报纸文章中没有的语言表达。在本文中,我们建立了一个带有语义关系标注的多样化文档引导语语料库。为了减少标注人员的工作量,并尽可能多地标注各种文档,我们将每个文档的标注目标限制为只有前三句话。我们已经完成了1000个文档的语料库的建立,并报告了该语料库的统计数据。
In these days, semantic analysis has been actively studied in natural language processing. For the study of semantic analysis, corpora with semantic annotations are essential. Although there are such corpora annotated on newspaper articles, there are various genres and styles, including linguistic expressions that are not found in newspaper articles. In this paper, we build a diverse document leads corpus annotated with semantic relations. To reduce the workload of annotators and annotate as many various documents as possible, we restrict the annotation target of each document to only the first three sentences. We have completed building a corpus of 1,000 documents and report the statistics of this corpus.