A formal framework for linguistic annotation

A formal framework for linguistic annotation
复制标题

DOI:
10.1016/s0167-6393(00)00068-6
复制
发表时间:
2001-01-01
影响因子:
3.2
通讯作者:
Liberman, M
Liberman, M
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bird, S;Liberman, M

文献摘要

被引文献

相似文献

“语言注释”涵盖应用于原始语言数据的任何描述性或分析性符号。基本数据可以是时间函数的形式-音频、视频和/或生理记录-或者它可以是文本。增加的注释可能包括各种类型的音译(从语音特征到语篇结构)、词性和意义标注、句法分析、“命名实体”识别、共指注释等等。虽然有几个正在进行的努力为这样的注释提供格式和工具,并出版注释的语言数据库,但缺乏广泛接受的标准正成为一个关键问题。现有的拟议标准都侧重于文件格式。本文的重点是语言注释的逻辑结构。我们调查了各种现有的注释格式,并展示了一个共同的概念核心,注释图。这为构建、维护和搜索语言注释提供了一个正式的框架,同时与许多其他数据结构和文件格式保持一致。(C)2001爱思唯尔科技有限公司。保留所有权利。
'Linguistic annotation' covers any descriptive or analytic notations applied to raw language data. The basic data may be in the form of time functions - audio, video and/or physiological recordings - or it may be textual. The added notations may include transcriptions of all sorts (from phonetic features to discourse structures), part-of-speech and sense tagging, syntactic analysis,'named entity' identification, coreference annotation, and so on. While there are several ongoing efforts to provide formats and tools for such annotations and to publish annotated linguistic databases, the lack of widely accepted standards is becoming a critical problem. Proposed standards, to the extent they exist, have focused on file formats. This paper focuses instead on the logical structure of linguistic annotations. We survey a wide variety of existing annotation formats and demonstrate a common conceptual core, the annotation graph. This provides a formal framework for constructing, maintaining and searching linguistic annotations, while remaining consistent with many alternative data structures and file formats. (C) 2001 Elsevier Science B.V. All rights reserved.