Estimating the latent number of types in growing corpora with reduced cost–accuracy trade-off.

Estimating the latent number of types in growing corpora with reduced cost–accuracy trade-off.
复制标题

在降低成本与准确性权衡的情况下估计不断增长的语料库中的潜在类型数量。

DOI:
10.1017/s0305000915000094
复制
发表时间:
2015
期刊:
Journal of Child Language.
影响因子:
--
通讯作者:
S.
S.
中科院分区:
--
文献类型:
--
作者:
Hidaka;S.

文献摘要

相似文献

儿童言语中独特词的数量是反映儿童语言发展的最基本的统计指标之一。然而,我们可能会面临困难时,试图准确地评估随着时间的推移与有限的样本大小的儿童不断增长的语料库中的独特的单词的数量。本研究提出了一种新的技术来估计潜在的话从一系列的字由儿童说出。该技术利用类型数量的统计特性作为采样标记数量的函数。通过对横截面和纵向样本的实证数据分析,验证了该方法的实用性。收敛的经验证据表明,建议的估计提高了一组现有的估计词汇量估计的准确性。利用这个有效的估计,我们提出了一个新的抽样方案的词汇评估,具有较低的成本和较高的准确性相比,现有的方法。
The number of unique words in children's speech is one of most basic statistics indicating their language development. We may, however, face difficulties when trying to accurately evaluate the number of unique words in a child's growing corpus over time with a limited sample size. This study proposes a novel technique to estimate the latent number of words from a series of words uttered by children. This technique utilizes statistical properties of the number of types as a function of the number of sampled tokens. We tested the practical effectiveness of the proposed method in the empirical data analysis of the cross-sectional and longitudinal samples. The converging empirical evidence indicates that the proposed estimator improves the accuracy of vocabulary size estimation over a set of existing estimators. Utilizing this efficient estimator, we propose a new sampling scheme for vocabulary assessment that has lower cost and higher accuracy compared to existing methods.