课题基金 / 基金详情

Statistical Infrastructure for Combining Information from Multiple Data Sources

Statistical Infrastructure for Combining Information from Multiple Data Sources
用于组合来自多个数据源的信息的统计基础设施
批准号:
7939906
负责人:
TRIVELLORE E RAGHUNATHAN
金额:
$41.12万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-24 至 2012-08-31

项目摘要

项目成果

TRIVELLORE E RAGHUNATHAN的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供): 该应用解决了“处理医疗数据的信息技术”的广泛挑战领域(第10区)和两个具体的挑战主题(10-RR-101):促进医疗数据二次研究的信息技术示范项目和(10-HL-101):开发数据共享和分析方法,以从大规模观察数据中获得数据,特别是那些来自电子健康记录的数据,对心血管、肺和血液疾病的比较治疗效果和结果进行可靠的估计。这项请求中的几个挑战领域表明,没有单一的数据集来处理健康的生物、社会、心理、经济和环境决定因素之间的复杂相互作用。尽管如此,要做出明智的政策决策,还需要对这些方面有更深的理解。这项提案的首席调查员一直在参与制定综合来自多个数据来源的信息的方法。如果向实质性研究人员提供实施这些方法的适当的计算基础设施,这些方法可以得到改进并变得更加有用。这项提议的目标是开发一个统计基础设施,以便利汇集来自多个来源的数据,创建用于分析目的的数据集,并提供适当分析这种组合数据集的软件。在这种方法下,两个或多个数据集将被连接,跨数据源的公共变量将被对齐,而未对齐的变量将被视为一个或多个数据集中的缺失数据。多重归罪方法将进行调整,以便使用相当普遍的和半参数模型将来自多个来源的信息结合起来。最终产品将是建立在首席调查员开发的现有基础设施上的两种软件模式。在第一种模式中,用户将输入两个或更多个数据集,列出这些数据集中要分析的所有变量,指定统计模型来检验研究假设,然后软件将适当地分析数据并传播结果。当数据因保密问题而不能自由分发给研究人员时,第二种模式将为保存机密数据集的计算机系统增加一个基于网络的远程数据分析工具。然后,合法用户将能够通过网络登录到系统,并像以前一样指定分析请求。该软件的两个版本将被开发,一个作为许多研究人员使用的流行软件SAS的附加组件。第二个版本将是一个独立版本,可供无法访问SAS的研究人员使用。该软件将开发成可以在Windows和Linux平台上运行。这项研究的目的是建立一个统计软件系统,使生物统计学家、临床医生、流行病学家和其他公共卫生研究人员能够适应需要组合来自多个数据源的信息的统计模型。其目的是汇集行政、流行病学和临床研究数据库中的信息,以解决公共卫生研究问题。该软件系统可供数据协调中心使用,在那里网络研究人员可以远程提交分析请求。该系统还将对数据储存库有用,因为数据集因保密问题而无法发布,但允许研究人员请求根据这些数据集进行分析。统计软件系统将通过网络免费向公众分发。
英文摘要
DESCRIPTION (provided by applicant): This application addresses broad Challenge Area of "Information Technology for Processing Health Care Data" (Area # 10) and two specific Challenge topics (10-RR-101): Information Technology Demonstration Projects Facilitating Secondary Use of Healthcare Data for Research and (10-HL-101): Develop data sharing and analytic approaches to obtain from large- scale observational data, especially those derived from electronic health records, reliable estimates of comparative treatment effects and outcomes of cardiovascular, lung, and blood diseases. Several challenge areas in this request suggest that there is no single data set for addressing the complex interplay of biological, social, psychological, economic and environmental determinants of health. Still, deeper understanding of these facets is needed to make informed policy decisions. The principal investigator of this proposal has been involved in developing methods for combining information from multiple data sources. These methods can be enhanced and become more useful if a proper computational infrastructure implementing them were made available to substantive researchers. The goal of this proposal is to develop a statistical infrastructure to facilitate pooling of data from multiple sources, create data sets for analytical purposes and provide software for properly analyzing such combined data sets. Under this approach the two or more data sets will be concatenated, common variables across the data sources will be aligned and the nonaligned variables will be treated as missing data in one or more data sets. The multiple imputation methodology will be tailored to combine information from multiple sources using fairly general and semi-parametric models. The end product will be two modes of software will be built upon the existing infrastructure developed by the principal investigator. In the first mode, the user will input two or more data sets, list all the variables to be analyzed from these data sets, specify the statistical model to test a research hypothesis and the software will then properly analyze the data and disseminate the results. When data cannot be freely distributed to researchers due to confidentiality concerns, the second mode will add a web-based remote data analysis tool for the computer system holding the confidential data sets. The legitimate users will then be able to log-in to the system through web and specify the analytical request as before. Two versions of the software will be developed one as an add-on to SAS, popular software used by many researchers. The second version will be a stand alone version which can be used by researchers who do not have access to SAS. The software will be developed to work under both Windows and Linux platforms. The purpose of this research is to build a statistical software system that allow biostatisticians, clinicians, epidemiologists and other public health researchers to fit statistical models that call for combining information from multiple data sources. The aim is pool information from administrative, epidemiological and clinical study databases to address public health research. The software system can be used by the Data Coordinating Centers where the analysis requests can be remotely submitted by network researchers. The system will also be useful for data repositories where data sets cannot be released due to confidentiality concerns but allow researchers to request analysis based on those data sets. The statistical software system will be distributed via web freely to the public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Infrastructure for Combining Information from Multiple Data Sources
  • 批准号:
    7815586
  • 项目类别:
  • 资助金额:
    $40.24万
  • 财政年份:
    2009
  • 负责人:
    TRIVELLORE E RAGHUNATHAN
  • 依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: