MULTIPLE TOKENIZATIONS IN A DIACHRONIC CORPUS
MULTIPLE TOKENIZATIONS IN A DIACHRONIC CORPUS
复制标题
历时语料库中的多种标记化
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Amir Zeldes
中科院分区:
文献类型:
--
作者:
Thomas Krause;C. Odebrecht;Amir Zeldes
This paper deals with the construction of a maximally flexible corpus architecture for building and analyzing diachronic corpora. Historical data poses many challenges with regard to representation and analysis, and diachronic corpora are even more varied and unsystematic (Claridge, 2008). Since historical and diachronic corpora are so difficult and expensive to build, it is crucial that they be stored in an architecture that permits the addition of new texts and annotation layers at any point in time. In this paper we focus on two issues of corpus construction multiple normalizations and multiple tokenizations in a multi-layer architecture. We exemplify our methodological issues using a diachronic corpus of German scientific texts (Ridges Herbology1). The corpus contains excerpts from texts about herbs from 12 different sources written between 1543 and 1870.