Building a Treebank for Italian: a Data-driven Annotation Schema

Building a Treebank for Italian: a Data-driven Annotation Schema
复制标题

构建意大利语树库:数据驱动的注释模式

DOI:
--
复制
发表时间:
2000
期刊:
--
影响因子:
--
通讯作者:
L. Lesmo
L. Lesmo
中科院分区:
--
文献类型:
--
作者:
C. Bosco;V. Lombardo;D. Vassallo;L. Lesmo

文献摘要

被引文献

相似文献

目前,许多自然语言研究人员将注意力转向树库开发,并试图在其表示格式中实现准确性和语料库数据覆盖率。本文提出了一种数据驱动的注释模式,确保数据覆盖率和语言现象之间的注释一致性的意大利树库。该模式是一种基于依赖关系的格式,以谓词-论元结构的概念为中心,用痕迹来表示不连续的成分。树库的开发涉及到一个注释过程,该过程由一个交互式解析工具帮助人类注释者执行,该工具逐步构建句子的句法表示。为了增加这个解析器的语法知识,一个特定的数据驱动的策略已被应用。我们描述了注释模式的周期性发展,强调了格式的丰富性和灵活性,并提出了一些代表性问题。
Many natural language researchers are currently turning their attention to treebank development and trying to achieve accuracy and corpus data coverage in their representation formats. This paper presents a data-driven annotation schema developed for an Italian treebank ensuring data coverage and consistency between annotation of linguistic phenomena. The schema is a dependency-based format centered upon the notion of predicate-argument structure augmented with traces to represent discontinuous constituents. The treebank development involves an annotation process performed by a human annotator helped by an interactive parsing tool that builds incrementally syntactic representation of the sentence. To increase the syntactic knowledge of this parser, a specific data-driven strategy has been applied. We describe the cyclical development of the annotation schema highlighting the richness and flexibility of the format, and we present some representational issues.