The Czech National Corpus
The Czech National Corpus
复制标题
捷克国家语料库
DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
Věra Schmiedtová
中科院分区:
文献类型:
--
作者:
J. Kocek;Marie Koprivová;Věra Schmiedtová
The paper deals with the history of the Czech National Corpus (CNC) project. It reports on the present
stage of its development, describes what type of corpus it is, and the text processing methods and morphological
annotation used in its compilation. It also briefly discusses the software used in the CNC.
The Bank of Czech (BoC) has now 330 million word forms. It is the basis of a representative corpus
(SYN2000 - 100 million word forms) which was created in spring 2000, and is intended as a material
source for future dictionaries. At the moment the lexical saturation of the material is tested.