Federated queries of clinical data repositories: the sum of the parts does not equal the whole

Federated queries of clinical data repositories: the sum of the parts does not equal the whole
复制标题

DOI:
10.1136/amiajnl-2012-001299
复制
发表时间:
2013-06-01
影响因子:
6.4
通讯作者:
Weber, Griffin M.
Weber, Griffin M.
中科院分区:
管理学2区
文献类型:
--
作者:
Weber, Griffin M.

文献摘要

被引文献

相似文献

背景和目的2008年,我们开发了一个共享健康研究信息网络(SHRINE),首次实现了对波士顿四家医院全部患者群体的研究查询。它使用联邦体系结构,其中每家医院只返回符合查询的患者总数。这允许医院保留对本地数据库的控制权,并遵守联邦和州隐私法。但是,由于患者可能从多家医院接受治疗,因此联邦查询的结果可能与针对单个中央存储库运行查询的结果不同。本文描述了发生这种情况的情况,并提出了纠正这些错误的技术。方法通过比较患者人口统计数据的单向散列值,我们使用一次性过程来识别哪些患者在多个存储库中拥有数据。这使我们能够对本地数据库进行分区,以便给定分区内的所有患者都拥有同一医院子集的数据。然后在每个分区上分别独立地运行联邦查询,并将合并后的结果呈现给用户。使用理论边界和模拟医院网络,我们证明了一旦进行了划分,SHRINE可以更精确地估计匹配查询的患者数量。结论:不同医院患者群体重叠的不确定性限制了SHRINE和其他联合查询工具的有效性。我们的技术在保留聚合联邦架构的同时减少了这种不确定性。
Background and objective In 2008 we developed a shared health research information network (SHRINE), which for the first time enabled research queries across the full patient populations of four Boston hospitals. It uses a federated architecture, where each hospital returns only the aggregate count of the number of patients who match a query. This allows hospitals to retain control over their local databases and comply with federal and state privacy laws. However, because patients may receive care from multiple hospitals, the result of a federated query might differ from what the result would be if the query were run against a single central repository. This paper describes the situations when this happens and presents a technique for correcting these errors.Methods We use a one-time process of identifying which patients have data in multiple repositories by comparing one-way hash values of patient demographics. This enables us to partition the local databases such that all patients within a given partition have data at the same subset of hospitals. Federated queries are then run separately on each partition independently, and the combined results are presented to the user.Results Using theoretical bounds and simulated hospital networks, we demonstrate that once the partitions are made, SHRINE can produce more precise estimates of the number of patients matching a query.Conclusions Uncertainty in the overlap of patient populations across hospitals limits the effectiveness of SHRINE and other federated query tools. Our technique reduces this uncertainty while retaining an aggregate federated architecture.