The DTA “Base Format”: A TEI Subset for the Compilation of a Large Reference Corpus of Printed Text from Multiple Sources

The DTA “Base Format”: A TEI Subset for the Compilation of a Large Reference Corpus of Printed Text from Multiple Sources
复制标题

DOI:
10.4000/jtei.1114
复制
发表时间:
2014-12
期刊:
--
影响因子:
--
通讯作者:
S. Haaf;Alexander Geyken;Frank Wiegand
S. Haaf;Alexander Geyken;Frank Wiegand
中科院分区:
其他
文献类型:
--
作者:
S. Haaf;Alexander Geyken;Frank Wiegand

文献摘要

被引文献

相似文献

在本文中,我们描述了DTA“基本格式”(DTABf),它是TEI P5标记集的一个严格子集。DTABf的目的是在表达性和准确性之间提供一种平衡,并为来自多个来源的印刷文本的历史语料库的各种文本类型提供一种可互操作的注释方案。DTABf是在Deutsches Textarchiv (DTA)项目核心语料库中的大量历史文本数据和15个合作项目的文本集合的基础上开发的,目前总共有2.1亿个令牌。DTABf是一种“活的”TEI格式,当遇到包含新结构现象的DTA新文本候选时,它会不断进行调整。我们还关注DTABf的其他方面,包括一致性,与其他TEI方言的互操作性,HTML和TEI文本的其他表示,以及转换为其他格式,以及语言分析。我们包括一些最佳实践示例,以说明如何将外部语料库无损地转换为DTABf,从而使第三方能够在其特定项目中使用DTABf。DTABf有全面的文档记录,并且有几个软件工具可用于处理它,使其成为广泛使用的编码历史印刷德文文本的格式。
In this article we describe the DTA “Base Format” (DTABf), a strict subset of the TEI P5 tag set. The purpose of the DTABf is to provide a balance between expressiveness and precision as well as an interoperable annotation scheme for a large variety of text types of historical corpora of printed text from multiple sources. The DTABf has been developed on the basis of a large amount of historical text data in the core corpus of the project Deutsches Textarchiv (DTA) and text collections from 15 cooperating projects with a current total of 210 million tokens. The DTABf is a “living” TEI format which is continuously adjusted when new text candidates for the DTA containing new structural phenomena are encountered. We also focus on other aspects of the DTABf including consistency, interoperability with other TEI dialects, HTML and other presentations of the TEI texts, and conversion into other formats, as well as linguistic analysis. We include some examples of best practices to illustrate how external corpora can be losslessly converted into the DTABf, thus enabling third parties to use the DTABf in their specific projects. The DTABf is comprehensively documented, and several software tools are available for working with it, making it a widely used format for the encoding of historical printed German text.