Corpus databases with feature pre-calculation

Corpus databases with feature pre-calculation
复制标题

具有特征预计算的语料数据库

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
E. Komen
E. Komen
中科院分区:
--
文献类型:
--
作者:
E. Komen

文献摘要

被引文献

相似文献

可靠编码的树库是语言学研究的金矿。查询一个典型的研究问题包括:(a)查询树库以提取包含要研究的特征的句子,(B)识别并跟踪确定语言特征编码方式的特征,以及(c)使用统计数据来找出这些特征中的哪一个(组合)确定语言特征的结果。虽然在这一进程中,步骤(a)和(c)有足够的工具,但步骤(B)尚未得到重视。本文描述了如何使用“Cesax”和“CorpusStudio”程序来共同构建一个“语料库研究数据库”,该数据库包含步骤(a)中选择的感兴趣的句子,以及步骤(B)的用户可定义的预先计算的特征。
Reliably coded treebanks are a goldmine for linguistics research. Answering a typical research question involves: (a) querying a treebank to extract sentences containing the feature to be investigated, (b) recognizing and keeping track of characteristics that determine the way in which the linguistic feature is encoded, and (c) using statistics to find out which (combination) of these characteristics determines the outcome of the linguistic feature. While sufficient tools are available for steps (a) and (c) in this process, step (b) has not received much attention yet. This paper describes how the programs “Cesax” and “CorpusStudio” can be used jointly to construct a “corpus research database”, a database that contains the sentences of interest selected in step (a), as well as user-definable pre-calculated characteristics for step (b).