Optimal prediction of the number of unseen species

Optimal prediction of the number of unseen species
复制标题

DOI:
10.1073/pnas.1607774113
复制
发表时间:
2016-11-22
影响因子:
11.1
通讯作者:
Wu, Yihong
Wu, Yihong
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Orlitsky, Alon;Suresh, Ananda Theertha;Wu, Yihong

文献摘要

被引文献

相似文献

在许多科学努力中,估计看不见的物种的数量是一个重要的问题。 Fisher等人推出的最受欢迎的配方。 [Fisher RA,Corbet AS,Williams CB(1943)J动物ECOL 12(1):42-58],使用N样品预测迄今未见物种的u数量,如果收集了T.N新样本,则会观察到。引人注目的是可以准确预测的新样本和现有样本数量之间的最大比例。在开创性的作品中,Good and Toulmin [Good I,Toulmin G(1956)Biometrika 43(102):45-63]构建了一个有趣的估计量,该估计量可以预测所有T 1,但没有可证明的保证。我们得出了一类估计器,这些估计器可以预测到t的所有过程。 log n。我们还表明,此范围是最好的,并且估算器的均方误差对于任何t来说都是最佳的。我们的方法可为EFRON THESTIST估算器提供可证明的保证,此外,与在各种合成数据集上的现有方法相比,具有更强理论和实验性能的变体。估计器是简单的,线性的,计算上有效的,并且可扩展到大量数据集。它们的性能保证可用于所有分布,并适用于在各种科学学科中常用的所有四种标准抽样模型:多项式,泊松,超几何和伯努利产品。
Estimating the number of unseen species is an important problem in many scientific endeavors. Its most popular formulation, introduced by Fisher et al. [Fisher RA, Corbet AS, Williams CB (1943) J Animal Ecol 12(1): 42-58], uses n samples to predict the number U of hitherto unseen species that would be observed if t.n new samples were collected. Of considerable interest is the largest ratio t between the number of new and existing samples for which U can be accurately predicted. In seminal works, Good and Toulmin [Good I, Toulmin G (1956) Biometrika 43(102): 45-63] constructed an intriguing estimator that predicts U for all t 1, but without provable guarantees. We derive a class of estimators that provably predict U all of the way up to t. log n. We also show that this range is the best possible and that the estimator's mean-square error is near optimal for any t. Our approach yields a provable guarantee for the Efron-Thisted estimator and, in addition, a variant with stronger theoretical and experimental performance than existing methodologies on a variety of synthetic and real datasets. The estimators are simple, linear, computationally efficient, and scalable to massive datasets. Their performance guarantees hold uniformly for all distributions, and apply to all four standard sampling models commonly used across various scientific disciplines: multinomial, Poisson, hypergeometric, and Bernoulli product.