Morphosyntactic annotation of CHILDES transcripts.

Morphosyntactic annotation of CHILDES transcripts.
复制标题

DOI:
10.1017/s0305000909990407
复制
发表时间:
2010-06
影响因子:
2.2
通讯作者:
Wintner, Shuly
Wintner, Shuly
中科院分区:
人文科学3区
文献类型:
--
作者:
Sagae, Kenji;Davis, Eric;Lavie, Alon;MacWhinney, Brian;Wintner, Shuly

文献摘要

被引文献

相似文献

儿童语言语料库是儿童语言习得和心理语言学研究的基础。语料库的语言学注释为研究者探索语法结构的发展及其用法提供了更好的手段。我们描述了一个项目,其目标是用标记的依存结构的形式用语法关系注释Childes数据库的英语部分。我们已经产生了一个超过18,800个话语(大约65,000个单词)的语料库,其中包含人工精选的黄金标准语法关系注释。使用这个语料库,我们已经开发了一个高度准确的数据驱动的英语Childes数据解析器,我们使用它来自动标注Childes英语部分的其余部分。我们还将解析器扩展到西班牙语,目前正在努力支持更多语言。解析器以及手动和自动注释的数据可免费用于研究目的。
Corpora of child language are essential for research in child language acquisition and psycholinguistics. Linguistic annotation of the corpora provides researchers with better means for exploring the development of grammatical constructions and their usage. We describe a project whose goal is to annotate the English section of the CHILDES database with grammatical relations in the form of labeled dependency structures. We have produced a corpus of over 18,800 utterances (approximately 65,000 words) with manually curated gold-standard grammatical relation annotations. Using this corpus, we have developed a highly accurate data-driven parser for the English CHILDES data, which we used to automatically annotate the remainder of the English section of CHILDES. We have also extended the parser to Spanish, and are currently working on supporting more languages. The parser and the manually and automatically annotated data are freely available for research purposes.