Reliability of relational event model estimates under sampling: How to fit a relational event model to 360 million dyadic events

Reliability of relational event model estimates under sampling: How to fit a relational event model to 360 million dyadic events
复制标题

抽样下关系事件模型估计的可靠性:如何将关系事件模型拟合到 3.6 亿个二元事件

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
1.7
通讯作者:
A. Lomi
A. Lomi
中科院分区:
--
文献类型:
--
作者:
J. Lerner;A. Lomi

文献摘要

被引文献

相似文献

摘要我们评估了两种抽样方案下估计的关系事件模型(REM)参数的可靠性:(1)从观察到的事件中均匀抽样和(2)病例对照抽样,从适当定义的风险集中抽取非事件或空二对(“对照”)。我们实验确定的变异性估计参数作为一个函数的采样事件和控制每个事件的数量,分别。结果表明,REM可以可靠地适合网络连接超过1200万个节点,超过3.6亿个二元事件,通过分析一个样本的一些数万个事件和少量的控制每个事件。使用我们在维基百科编辑网络上收集的数据,我们说明了基于REM的实证研究中通常包含的网络效应如何需要广泛不同的样本量来可靠地估计。在我们的分析中,我们使用了一个开源软件,该软件实现了两种抽样方案,允许分析人员将REM拟合和分析到可能在不同的经验设置、不同的样本参数或模型规格中收集的相同或其他数据。
Abstract We assess the reliability of relational event model (REM) parameters estimated under two sampling schemes: (1) uniform sampling from the observed events and (2) case–control sampling which samples nonevents, or null dyads (“controls”), from a suitably defined risk set. We experimentally determine the variability of estimated parameters as a function of the number of sampled events and controls per event, respectively. Results suggest that REMs can be reliably fitted to networks with more than 12 million nodes connected by more than 360 million dyadic events by analyzing a sample of some tens of thousands of events and a small number of controls per event. Using the data that we collected on the Wikipedia editing network, we illustrate how network effects commonly included in empirical studies based on REMs need widely different sample sizes to be reliably estimated. For our analysis we use an open-source software which implements the two sampling schemes, allowing analysts to fit and analyze REMs to the same or other data that may be collected in different empirical settings, varying sample parameters or model specification.