Estimation of Global Network Statistics from Incomplete Data

Estimation of Global Network Statistics from Incomplete Data
复制标题

DOI:
10.1371/journal.pone.0108471
复制
发表时间:
2014-10-22
期刊:
影响因子:
3.7
通讯作者:
Dodds, Peter Sheridan
Dodds, Peter Sheridan
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Bliss, Catherine A.;Danforth, Christopher M.;Dodds, Peter Sheridan

文献摘要

被引文献

相似文献

复杂网络是各种社会、生物、物理和虚拟系统的基础。复杂网络科学的一个深刻的复杂性是,在大多数情况下,观察所有节点和所有网络交互是不可能的。以前的工作解决部分网络数据的影响是令人惊讶的有限,主要集中在丢失的节点,并建议从二次抽样数据得出的网络统计数据是不合适的估计相同的网络统计描述的整体网络拓扑结构。我们生成缩放方法来预测真实的网络统计数据,包括度分布,仅从节点,链接或权重的部分知识。我们的方法是透明的,并且不假设网络的已知生成过程,从而能够预测各种应用的网络统计数据。我们验证了四个模拟网络类和经验数据集的各种规模的分析结果。我们通过不同比例的采样数据进行子采样实验,并证明我们的缩放方法可以提供非常好的真实网络统计估计,同时承认限制。最后,我们将我们的技术应用于一组丰富和不断发展的大规模社交网络,Twitter回复网络。基于1亿条推文,我们使用我们的缩放技术,提出了一个统计特性的Twitter互动组从2008年9月至2008年11月。我们的治疗使我们能够找到邓巴的假设在检测一个上限阈值的数量活跃的社会接触,个人保持在一周的过程中的支持。
Complex networks underlie an enormous variety of social, biological, physical, and virtual systems. A profound complication for the science of complex networks is that in most cases, observing all nodes and all network interactions is impossible. Previous work addressing the impacts of partial network data is surprisingly limited, focuses primarily on missing nodes, and suggests that network statistics derived from subsampled data are not suitable estimators for the same network statistics describing the overall network topology. We generate scaling methods to predict true network statistics, including the degree distribution, from only partial knowledge of nodes, links, or weights. Our methods are transparent and do not assume a known generating process for the network, thus enabling prediction of network statistics for a wide variety of applications. We validate analytical results on four simulated network classes and empirical data sets of various sizes. We perform subsampling experiments by varying proportions of sampled data and demonstrate that our scaling methods can provide very good estimates of true network statistics while acknowledging limits. Lastly, we apply our techniques to a set of rich and evolving large-scale social networks, Twitter reply networks. Based on 100 million tweets, we use our scaling techniques to propose a statistical characterization of the Twitter Interactome from September 2008 to November 2008. Our treatment allows us to find support for Dunbar's hypothesis in detecting an upper threshold for the number of active social contacts that individuals maintain over the course of one week.