The Czech National Corpus

The Czech National Corpus
复制标题

捷克国家语料库

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
Věra Schmiedtová
Věra Schmiedtová
中科院分区:
--
文献类型:
--
作者:
J. Kocek;Marie Koprivová;Věra Schmiedtová

文献摘要

被引文献

相似文献

本文介绍了捷克国家语料库(CNC)项目的历史。它报告了目前 介绍了它是什么类型的语料库,以及文本的处理方法和词法 在其编译中使用的注释。并简要介绍了数控系统中所使用的软件。 捷克银行(Bank of Czech)目前拥有3.3亿个单词。它是代表性语料库的基础 (SYN 2000 - 1亿字的形式),这是在2000年春季创建,并打算作为一种材料, 未来字典的来源。目前,测试材料的词汇饱和度。
The paper deals with the history of the Czech National Corpus (CNC) project. It reports on the present stage of its development, describes what type of corpus it is, and the text processing methods and morphological annotation used in its compilation. It also briefly discusses the software used in the CNC. The Bank of Czech (BoC) has now 330 million word forms. It is the basis of a representative corpus (SYN2000 - 100 million word forms) which was created in spring 2000, and is intended as a material source for future dictionaries. At the moment the lexical saturation of the material is tested.