Privacy-protecting distributed analysis of biomedical big data
Privacy-protecting distributed analysis of biomedical big data
批准号:
9159815
负责人:
Darren Toh
金额:
$50.01万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-30 至 2019-06-30
关键词:
AgreementBig DataBioinformaticsBiomedical ResearchClinical ResearchCodeComplexComputer softwareConfidentiality of Patient InformationDataData AnalysesData ProtectionData ScienceData SetData SourcesDatabasesDevelopmentDistantElectronic Health RecordEnvironmentFundingHealthHealthcare SystemsHousingIndividualInsuranceLinear RegressionsLinkLogistic RegressionsMethodsMulticenter StudiesPatientsPerformancePrivacyProcessProgramming LanguagesPublic HealthRegistriesRegression AnalysisResearchResearch PersonnelSecuritySentinelSiteSoftware ToolsSourceStatistical ComputingStatistical Data InterpretationStatistical ModelsSystemTechnologyTestingUnited States National Institutes of HealthWorkbasebig biomedical datacollaboratorydata sharingdata structuredesigndistributed dataexperiencehandheld mobile deviceimprovedmultidisciplinaryopen dataopen sourcepatient orientedprecision medicineprogramsreal world applicationsocial mediastatisticssystems researchtool
中文摘要
摘要
技术、生物信息学和数据科学的进步使分析大型和复杂的数据成为可能
生成证据的数据库,以改善公共健康并加快精确度的发展
医药。然而,大数据的出现也引发了人们对隐私和保密的担忧。这
应用程序专注于垂直分区数据中的数据隐私,这是一个数据环境,在这种环境中,信息
有关个人的信息在两个或更多数据源中可用。这种类型的数据结构在
生物医学研究,预计将呈指数级增长,因为来自同一个人的信息
越来越多地从多个来源收集数据,如保险索赔数据库、电子健康记录、
注册表、社交媒体、可穿戴设备和移动设备。组合多个数据库可提供更多
完整的关于患者的健康档案,并产生更有力的证据。然而,对数据的担忧
隐私、机密性和安全性,以及治理和机构协议中的限制,使其高度
在物理上汇集不同的数据源具有挑战性,有时甚至是不可能的。我们建议开发一种
开源的免费软件工具,它将采用一种尖端的方法--分布式回归--来
分析垂直分区的数据集。该方法不需要在物理上合并数据,但
产生统计上相同的结果,就像数据集被链接并集中在一个站点上一样。
参与的网站将只传输不可识别的信息,而不是共享患者级别的信息
矩阵(用于拟合统计模型的设计矩阵)和需要的其他汇总级统计
统计建模过程。这种方法为数据隐私提供了更好的保护,同时允许
进行复杂的统计分析。软件工具将使用以下工具进行开发、测试和微调
既有模拟数据集,也有来自Optus Labs的真实数据,该实验室拥有最大的垂直
分割了美国的数据集,其中包含500多万名患者的索赔和电子健康记录数据。这个
工具将与PopMedNetTM兼容,PopMedNetTM是目前由
几个大型国家倡议,如NIH医疗保健系统研究合作实验室分发
研究网络,PCORI资助的以患者为中心的国家临床研究网络(PCORnet),以及
FDA资助的哨兵计划。因此,该工具具有高度可伸缩性,并可立即对
真实世界的大数据分析。多学科研究团队包括开创了一些新技术的研究人员
分布式回归方法和在多中心研究方面拥有丰富经验的专家。这个
分布式回归方法有很大的潜力改变多中心大生物医学的范式
研究,从传输潜在可识别的患者级别数据到共享不可识别的
汇总级信息。拟议的软件工具将是迈向现实世界应用的重要一步
这种最先进的隐私保护分析方法。
英文摘要
ABSTRACT
Advances in technology, bioinformatics, and data science have made it possible to analyze large and complex
databases to generate evidence that improves public health and accelerates the development of precision
medicine. However, the advent of big data has also raised concerns about privacy and confidentiality. This
application is focused on data privacy in vertically partitioned data, a data environment where information
about an individual is available in two or more data sources. This type of data structure is common in
biomedical research and is expected to grow exponentially as information from the same individual is
increasingly collected in multiple sources, such as insurance claims databases, electronic health records,
registries, social media, wearables, and mobile devices. Combining multiple databases provides a more
complete health profile about the patient and generates more robust evidence. However, concerns about data
privacy, confidentiality, and security, and constraints in governance and institutional agreements make it highly
challenging or sometimes impossible to physically pool different data sources. We propose to develop an
open-source, freely available software tool that will employ a cutting-edge method – distributed regression – to
analyze vertically partitioned datasets. The method does not require data to be combined physically, but
produces statistically equivalent results as if the datasets were linked and pooled centrally at one site.
Instead of sharing patient-level information, participating sites will only transfer non-identifiable information
matrix (a design matrix used in fitting of statistical models) and other summary-level statistics needed in the
statistical modeling process. This approach offers much greater protection for data privacy while allowing one
to perform sophisticated statistical analysis. The software tool will be developed, tested, and fine-tuned using
both simulated datasets and the real-world data from Optum Labs, which houses one of the largest vertically
partitioned datasets in the U.S. with claims and electronic health record data from over 5 million patients. The
tool will be made compatible with PopMedNetTM, an open-source data-sharing platform currently used by
several large national initiatives such as the NIH Health Care Systems Research Collaboratory Distributed
Research Network, the PCORI-funded National Patient-Centered Clinical Research Network (PCORnet), and
the FDA-funded Sentinel program. The tool is therefore highly scalable and can have immediate impacts on
real-world big data analysis. The multidisciplinary study team includes researchers who pioneered some of the
distributed regression approaches and experts who have extensive experience in multi-center studies. The
distributed regression method has great potential to shift the paradigm of multi-center big biomedical
research, from transferring of potentially identifiable patient-level data to the sharing of non-identifiable
summary-level information. The proposed software tool will be a major step towards real-world application
of this state-of-the-art privacy-protecting analytic approach.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Identifying treatment-resistant depression in automated databases
-
批准号:8110228
-
项目类别:
-
资助金额:$9.86万
-
财政年份:2011
-
负责人:Darren Toh
-
依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: