Corpus databases with feature pre-calculation
Corpus databases with feature pre-calculation
复制标题
具有特征预计算的语料数据库
DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
E. Komen
中科院分区:
文献类型:
--
作者:
E. Komen
Reliably coded treebanks are a goldmine for linguistics research. Answering a typical research question involves: (a) querying a treebank to extract sentences containing the feature to be investigated, (b) recognizing and keeping track of characteristics that determine the way in which the linguistic feature is encoded, and (c) using statistics to find out which (combination) of these characteristics determines the outcome of the linguistic feature. While sufficient tools are available for steps (a) and (c) in this process, step (b) has not received much attention yet. This paper describes how the programs “Cesax” and “CorpusStudio” can be used jointly to construct a “corpus research database”, a database that contains the sentences of interest selected in step (a), as well as user-definable pre-calculated characteristics for step (b).