课题基金 / 基金详情

Statistical Methods for Network Data

Statistical Methods for Network Data
网络数据的统计方法
批准号:
1106772
负责人:
Elizaveta Levina
金额:
$29.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-07-01 至 2015-06-30

项目摘要

项目成果

Elizaveta Levina的其他基金

相似基金

相关文献

中文摘要
翻译
网络数据已经在广泛的领域中变得普遍,并且大量不同的研究人员已经研究了网络的各个方面,但很少应用统计方法。该项目提出了新的理论、方法和算法,采用有原则的统计方法来解决这些问题,评估不确定性,并为一致性等理想属性建立条件。网络社区结构是网络实践中的一个普遍现象,也是网络分析中的一个基本问题。提出了新的拟合网络块模型的伪似然算法,以及几种允许块内非均匀度分布的推广,消除了经典块模型的主要局限性。基于聚合数据的伪似然大大加快了计算速度,允许将这些模型拟合到比以前更大更稀疏的网络中。研究了用于群体检测的准则的渐近分布,从而发展了群体结构、一致性条件和渐近正确划分阈值的显著性检验,具有重要的实际意义。还提出了新的、更稳健的标准,在较弱的条件下保持一致。该建议还开发了一种正式的非参数检验,用于比较两个网络,这是一个在实践中经常出现的问题,但目前只通过非正式的汇总统计比较来解决。最后,节点和边缘上的协变量被纳入模型,并用于预测网络中未观察到的链接。许多提出的方法为相应的网络问题提供了第一个统计解决方案。网络社区检测的统计方法的发展,在促进核心统计理论和方法发展的同时,对网络分析和复杂网络研究的跨学科领域产生了直接影响。这些应用非常广泛,涵盖了传染病建模、国家安全、通信、社会学和基因组学等不同领域。提出的新统计工具采用了一种更正式、更严格的方法,并且有可能改变许多科学家进行网络分析的方式。
英文摘要
Network data have become common in a wide range of fields, and a large and diverse community of researchers have studied various aspects of networks, yet statistical methods are rarely applied. This project proposes new theory, methodology, and algorithms that take a principled statistical approach to these problems, assess uncertainty, and establish conditions for desirable properties such as consistency. The focus is primarily on discovering community structure in networks, a common phenomenon in practice and a fundamental question in network analysis. New pseudo-likelihood algorithms are proposed for fitting the block model for networks, as well as several generalizations that allow for non-uniform degree distribution within blocks, removing the main limitation of the classic block model. The pseudo-likelihood based on aggregated data substantially speeds up computation, allowing fitting these models to larger and sparser networks than previously possible. The asymptotic distribution of criteria used for community detection is also studied, which leads to development of significance tests for community structure, consistency conditions, and asymptotically correct partition thresholds, which have important practical implications. New, more robust criteria are also proposed, consistent under weaker conditions. The proposal also develops a formal non-parametric test for comparing two networks, a problem that arises frequently in practice but is currently addressed only through informal comparisons of summary statistics. Finally, covariates on nodes and edges are incorporated into the models and used for predicting unobserved links in the networks. Many of the proposed methods provide the first statistical solutions to the corresponding network problems. Development of statistical methods for community detection in networks, while contributing to the development of core statistical theory and methodology, has direct impact on the interdisciplinary field of network analysis and the study of complex networks. The applications of these are wide-spread, covering such diverse areas as infectious disease modeling, national security, communications, sociology, and genomics. The new statistical tools proposed take a more formal, rigorous approach, and have the potential to change how many scientists approach network analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FRG: Collaborative Research: Flexible Network Inference
Multivariate Analysis for Samples of Networks
RTG: Understanding dynamic big data with complex structure
Conference proposal: From Industrial Statistics to Data Science
国内基金
海外基金
Computational Methods for Analyzing Toponome Data