MULTIPLE TOKENIZATIONS IN A DIACHRONIC CORPUS

MULTIPLE TOKENIZATIONS IN A DIACHRONIC CORPUS
复制标题

历时语料库中的多种标记化

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Amir Zeldes
Amir Zeldes
中科院分区:
--
文献类型:
--
作者:
Thomas Krause;C. Odebrecht;Amir Zeldes

文献摘要

被引文献

相似文献

本文探讨了一个最大限度灵活的语料库体系结构,以建立和分析历时语料库。历史数据在表现和分析方面提出了许多挑战,而历时语料库更加多样化和不系统(Claridge,2008)。由于历史和历时语料库的构建是如此困难和昂贵,因此将它们存储在允许在任何时间点添加新文本和注释层的架构中至关重要。在本文中,我们集中在两个问题的语料库建设,多重规范化和多重标记在一个多层架构。我们使用德国科学文本的历时语料库(山脊草药学1)来解释我们的方法论问题。该语料库包含1543年至1870年间12个不同来源的草药文本摘录。
This paper deals with the construction of a maximally flexible corpus architecture for building and analyzing diachronic corpora. Historical data poses many challenges with regard to representation and analysis, and diachronic corpora are even more varied and unsystematic (Claridge, 2008). Since historical and diachronic corpora are so difficult and expensive to build, it is crucial that they be stored in an architecture that permits the addition of new texts and annotation layers at any point in time. In this paper we focus on two issues of corpus construction multiple normalizations and multiple tokenizations in a multi-layer architecture. We exemplify our methodological issues using a diachronic corpus of German scientific texts (Ridges Herbology1). The corpus contains excerpts from texts about herbs from 12 different sources written between 1543 and 1870.