Building a Treebank for Italian: a Data-driven Annotation Schema
Building a Treebank for Italian: a Data-driven Annotation Schema
复制标题
构建意大利语树库:数据驱动的注释模式
DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
L. Lesmo
中科院分区:
文献类型:
--
作者:
C. Bosco;V. Lombardo;D. Vassallo;L. Lesmo
Many natural language researchers are currently turning their attention to treebank development and trying to achieve accuracy and corpus data coverage in their representation formats. This paper presents a data-driven annotation schema developed for an Italian treebank ensuring data coverage and consistency between annotation of linguistic phenomena. The schema is a dependency-based format centered upon the notion of predicate-argument structure augmented with traces to represent discontinuous constituents. The treebank development involves an annotation process performed by a human annotator helped by an interactive parsing tool that builds incrementally syntactic representation of the sentence. To increase the syntactic knowledge of this parser, a specific data-driven strategy has been applied. We describe the cyclical development of the annotation schema highlighting the richness and flexibility of the format, and we present some representational issues.