Estimating the size of the telephone universe: a Bayesian Mark-recapture approach

Estimating the size of the telephone universe: a Bayesian Mark-recapture approach
复制标题

估计电话宇宙的大小:贝叶斯标记重捕获方法

DOI:
10.1145/1014052.1014136
复制
发表时间:
2004
期刊:
Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining
影响因子:
--
通讯作者:
David Poole
David Poole
中科院分区:
--
文献类型:
--
作者:
David Poole

文献摘要

被引文献

相似文献

多年来,标记-再捕获模型一直被用来估计未知的动物和鸟类种群的规模。在本文中,我们采用有限混合标记-重捕获模型来估计美国的活跃电话线数量。这个想法是使用在长途网络上观察到的线路的呼叫模式来估计没有出现在网络上的线路的数量。我们提出了一种贝叶斯方法,并使用马尔科夫链蒙特卡罗方法从模型参数的后验分布中得到推论。在州一级,我们的结果与最近公布的在线计数报告相当一致。对于容易分类为商业或住宅的线,估计具有低方差。当分类未知时,变异性显著增加。结果对先验分布的变化不敏感。我们讨论了由数据规模引起的重大计算和数据挖掘挑战,在数周内每天观察到大约3.5亿个呼叫详细记录。
Mark-recapture models have for many years been used to estimate the unknown sizes of animal and bird populations. In this article we adapt a finite mixture mark-recapture model in order to estimate the number of active telephone lines in the USA. The idea is to use the calling patterns of lines that are observed on the long distance network to estimate the number of lines that do not appear on the network. We present a Bayesian approach and use Markov chain Monte Carlo methods to obtain inference from the posterior distributions of the model parameters. At the state level, our results are in fairly good agreement with recent published reports on line counts. For lines that are easily classified as business or residence, the estimates have low variance. When the classification is unknown, the variability increases considerably. Results are insensitive to changes in the prior distributions. We discuss the significant computational and data mining challenges caused by the scale of the data, approximately 350 million call-detail records per day observed over a number of weeks.