NoSta-D: A Corpus of German Non-standard Varieties

NoSta-D: A Corpus of German Non-standard Varieties
复制标题

NoSta-D:德国非标准品种语料库

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Marc Reznicek
Marc Reznicek
中科院分区:
--
文献类型:
--
作者:
Stefanie Dipper;Anke Lüdeling;Marc Reznicek

文献摘要

被引文献

相似文献

直到最近,计算语言学的大部分研究都是在报纸文本上进行的。现在,重点已经扩展到其他类型的语言数据。这意味着许多语言描述和自动工具需要适应或扩展到非报纸语言。德语非标准语料库(NoSta-D)将为依赖分析、命名实体识别和域外文本类型的共指解析的评价和训练数据提供首个金标准。
Until recently, most research in computational linguistics has been done on newspaper texts. Nowadays, the focus has been extended to other types of language data. This means that many linguistic descriptions and automatic tools need to be adapted or extended to non-newspaper language. The non-standard varieties corpus of German (NoSta-D) will provide a first gold standard for evaluation and training data of dependency analysis, named entity recognition and coreference resolution for out-of-domain text types.