Latent network models to account for noisy, multiply reported social network data

Latent network models to account for noisy, multiply reported social network data
复制标题

DOI:
10.1093/jrsssa/qnac004
复制
发表时间:
2023-02-08
影响因子:
2
通讯作者:
Power, Eleanor A.
Power, Eleanor A.
中科院分区:
数学4区
文献类型:
--
作者:
De Bacco, Caterina;Contisciani, Martina;Power, Eleanor A.

文献摘要

被引文献

相似文献

社交网络数据通常通过合并来自多个人的报告来构建。然而,如何调和来自个人的不一致反应并不明显。如果人们的反应反映了规范性的期望,比如对平衡、互惠关系的期望,那么多重报告的数据可能会有特殊的风险。在这里,我们提出了一个概率模型,它结合了多个人报告的关系,以估计未观察到的网络结构。除了估计每个报告者的参数,这是与他们的倾向,过度或报告不足的关系,该模型明确纳入了一个术语的“相互性”,倾向于报告关系在两个方向涉及相同的改变。我们的模型的算法实现是基于变分推理,这使得它高效和可扩展到大型系统。我们将我们的模型应用于从一个以名册为基础的设计和75个印度村庄收集的一个名称生成器的设计,从一个巴布亚社区收集的数据。我们在两个数据集中观察到“相互性”的强有力证据,并发现该值因关系类型而异。因此,我们的模型估计网络的互易值与标准确定性聚合方法产生的互易值有很大不同,这表明在收集、构建和分析基于调查的网络数据时需要考虑这些问题。
Social network data are often constructed by incorporating reports from multiple individuals. However, it is not obvious how to reconcile discordant responses from individuals. There may be particular risks with multiply reported data if people's responses reflect normative expectations-such as an expectation of balanced, reciprocal relationships. Here, we propose a probabilistic model that incorporates ties reported by multiple individuals to estimate the unobserved network structure. In addition to estimating a parameter for each reporter that is related to their tendency of over- or under-reporting relationships, the model explicitly incorporates a term for 'mutuality', the tendency to report ties in both directions involving the same alter. Our model's algorithmic implementation is based on variational inference, which makes it efficient and scalable to large systems. We apply our model to data from a Nicaraguan community collected with a roster-based design and 75 Indian villages collected with a name-generator design. We observe strong evidence of 'mutuality' in both datasets, and find that this value varies by relationship type. Consequently, our model estimates networks with reciprocity values that are substantially different than those resulting from standard deterministic aggregation approaches, demonstrating the need to consider such issues when gathering, constructing, and analysing survey-based network data.