The TEI and Current Standards for Structuring Linguistic Data. An Overview

The TEI and Current Standards for Structuring Linguistic Data. An Overview
复制标题

TEI 和当前构建语言数据的标准。

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Maik Stührenberg
Maik Stührenberg
中科院分区:
--
文献类型:
--
作者:
Maik Stührenberg

文献摘要

被引文献

相似文献

TEI作为一种成熟的标注格式已经使用多年,适用于不同类型的语料库,包括经过语言标注的数据。虽然它是基于一个大型社区的共识,但它不具有标准的法律地位。在过去十年中,已经努力为语言数据制定明确的法律标准,这些标准不仅作为语言语料库交换的规范基础,而且还涉及最近的技术进步,例如基于网络的标准,以及大型和多重注释语料库的使用。在本文中,我们将概述国际标准化的过程,并讨论目前在ISO/TC 37(一个名为“术语和其他语言和内容资源”的技术委员会)主持下制定的一些国际标准。之后,将根据其正式模型、符号格式和注释模型,讨论TEI指南和这些规范之间的关系。本文的结论对语料库的处理提出了建议。
The TEI has served for many years as a mature annotation format for corpora of different types, including linguistically annotated data. Although it is based on the consensus of a large community, it does not have the legal status of a standard. During the last decade, efforts have been undertaken to develop definitive de jure standards for linguistic data that not only act as a normative basis for the exchange of language corpora but also address recent advancements in technology, such as web-based standards, and the use of large and multiply annotated corpora. In this article we will provide an overview of the process of international standardization and discuss some of the international standards currently being developed under the auspices of ISO/TC 37, a technical committee called “Terminology and other Language and Content Resources”. After that the relationship between the TEI Guidelines and these specifications, according to their formal model, notation format, and annotation model, will be discussed. The conclusion of the paper provides recommendations for dealing with language corpora.