The Czech Academic Corpus 2.0 Guide

The Czech Academic Corpus 2.0 Guide
复制标题

捷克学术语料库 2.0 指南

DOI:
10.2478/v10108-009-0003-9
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
J. Raab
J. Raab
中科院分区:
--
文献类型:
--
作者:
Barbora Hladká;Jan Hajic;Jirka Hana;Jaroslava Hlavácová;Jirí Mírovský;J. Raab

文献摘要

被引文献

相似文献

捷克学术语料库2.0版是一个包含650,000个单词的词法和句法注释语料库。捷克学术语料库(CAC)是由捷克共和国科学院捷克语言研究所的一个团队于1971年至1985年创建的。当CAC项目开始时,自20世纪60年代以来只有两个计算机注释语料库可用-美国英语布朗语料库和英国英语LOB语料库。这两个语料库都为语料库语言学家所熟知,而CAC仍然被隐藏,主要是因为20世纪80年代捷克共和国的政治制度。将CAC的内部格式和注释方案转移到布拉格部门树库(PDT)概念中的想法出现在PDT第二版的工作中。主要目标是使CAC和PDT完全兼容,从而使CAC能够集成到PDT中。目前发布的CAC第二版提供了内部格式以及形态和句法注释方案的完整转换。捷克学术语料库v.2.0由语言数据联盟出版。
The Czech Academic Corpus 2.0 Guide The Czech Academic Corpus version 2.0 is a morphologically and syntactically annotated corpus of 650,000 words. The Czech Academic Corpus (CAC) was created by a team from the Institute of the Czech Language of the Academy of Sciences of the Czech Republic from 1971 to 1985. When the CAC project began there were only two computerized annotated corpora available since the 1960s - the Brown Corpus of American English and the LOB Corpus of British English. Both corpora became well known to corpus linguists, whereas the CAC remained hidden mainly because of the 1980s political regime in the Czech Republic. The idea of transferring the internal format and annotation scheme of the CAC into the Prague Dependency Treebank (PDT) concept emerged during the work on the PDT's second version. The main goal was to make the CAC and the PDT fully compatible and thus enable the integration of the CAC into the PDT. The currently released second version of the CAC presents the complete conversion of the internal format and morphological and syntactical annotation schemes. The Czech Academic Corpus v. 2.0 is being published by the Linguistic Data Consortium.