课题基金 / 基金详情

Information Explorer: a Suite of Tools for Cross-study Genetic Loci Discovery

Information Explorer: a Suite of Tools for Cross-study Genetic Loci Discovery
信息浏览器:用于交叉研究遗传位点发现的一套工具
批准号:
8145063
负责人:
Yigal Arens
金额:
$44.75万
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-07-19 至 2013-05-31

项目摘要

项目成果

Yigal Arens的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):DBGaP等数据库代表了跨多个队列收集的极其宝贵的数据资源。随着成本效益高的高通量基因分型和测序技术的日益发展,产生了大量的基因数据。虽然建立这种数据库是为了存档和分发以前进行的遗传关联分析的结果,但越来越多的研究提供了不确定的个人一级的基因和表型数据,供获得适当授权的外部研究人员使用。虽然近年来提供的数据量大幅增加,但在促进各研究之间的表型协调方面所做的工作相对较少。许多心血管疾病的遗传流行病学研究都有与任何给定表型相关的多个变量,这些变量是由不同的定义和多个测量或数据子集产生的。研究人员在这样的数据库中搜索表型和基因组合的可用性时,面临着要筛选的名副其实的变量山。这通常需要访问多个网站以获得有关数据库中列出的变量的更多信息,并检查数据分布以评估队列之间的相似性。虽然遗传变异的命名策略在很大程度上是在研究中标准化的(例如,单核苷酸多态或SNP的“Rs”数字),但对于表型变量通常不是这样。对于一项特定的研究,通常有许多不同版本的表型变量。研究人员目前必须分析和比较越来越多的变量,这些变量具有不同程度的相关文档,以获得所需的信息。这是一个耗时的过程,可能仍然会错过最合适的变量。此外,每个想要比较相同数据集的研究人员通常都需要从头开始,因为没有工具来共享表型比较结果。可利用的信息工具使表型图谱更加有效并提高其准确性,以及直观的表型查询工具,将为研究人员利用这些数据库提供主要资源。我们建议的工具将允许研究人员(1)快速获得所需的信息,以评估特定研究是否对感兴趣的假设有用;(2)排除不符合研究标准的变量;(3)确定哪些研究具有感兴趣的表型和遗传信息的组合;以及(4)更容易将研究问题从最基本的主效应扩展到更复杂的分析,如基因与环境的交互作用和包含多种表型的多变量测试。效用的增加还将使进行更大规模的荟萃分析成为可能,因为研究人员将能够更快地研磨结果、排除变量和感兴趣的协变量,从而增加检测遗传关联的统计能力。 与公共卫生的相关性:虽然基因组数据的数量(例如,GWAS、测序等)尽管近年来可获得的表型数量急剧增加,但在促进各研究之间的表型协调方面所做的工作相对较少。我们提议的工具将允许研究人员快速识别感兴趣的数据集,将研究问题从最基本的主效应扩展到更复杂的分析,如基因与环境的相互作用和包含多种表型的多变量测试,并通过磨练结果、排除变量和感兴趣的协变量来轻松执行更大的荟萃分析,并增加检测遗传关联的统计能力。
英文摘要
DESCRIPTION (provided by applicant): Databases such as dbGaP represent extremely valuable resources of data that have been assembled across multiple cohorts. The increasing development of cost-effective high-throughput genotyping and sequencing technologies are resulting in vast amounts of genetic data. While such databases were formed in order to archive and distribute the results of previously performed genetic association analyses, an increasing number of studies have provided de-identified individual-level genotypic and phenotypic data that are made available to outside researchers who have obtained the appropriate authorization. While the amount of data made available has increased dramatically in recent years, relatively little has been done in order to facilitate phenotype harmonization across studies. Many genetic epidemiologic studies of cardiovascular disease have multiple variables related to any given phenotype, resulting from different definitions and multiple measurements or subsets of data. A researcher searching such databases for the availability of phenotype and genotype combinations is confronted with a veritable mountain of variables to sift through. This often requires visiting multiple websites to gain additional information about variables that are listed on databases, and examination of data distributions to assess similarities across cohorts. While the naming strategy for genetic variants is largely standardized across studies (e.g. "rs" numbers for single nucleotide polymorphisms or SNPs), this is often not the case for phenotype variables. For a given study, there are often numerous versions of phenotypic variables. Researchers currently have to analyze and compare increasingly larger numbers of variables that have varying degrees of documentation associated with them to obtain the desired information. This is a time-consuming process that may still miss the most appropriate variables. Moreover, every researcher that wants to compare the same datasets often needs to start from scratch since there are no tools to share the phenotype comparison results. The availability of informatic tools to make phenotype mapping more efficient and improve its accuracy, along with intuitive phenotype query tools, would provide a major resource for researchers utilizing these databases. The tools we are proposing would allow researchers to (1) Quickly obtain the information needed to assess whether a specific study will be useful for the hypothesis of interest; (2) Exclude variables that do not meet research criteria; (3) Ascertain which studies have combinations of phenotype and genetic information of interest; and (4) More easily expand research questions beyond the most basic main-effects to more complex analyses such as gene-by-environment interactions and multivariate tests incorporating multiple phenotypes. The increased utility will also enable larger meta-analyses to be performed, as researchers will be able to more quickly hone in on outcomes, exclusionary variables and covariates of interest, leading to increased statistical power to detect genetic associations. PUBLIC HEALTH RELEVANCE: While the amount of genomic data (e.g., GWAS, sequencing, etc.) made available has increased dramatically in recent years, relatively little has been done in order to facilitate phenotype harmonization across studies. The tools we are proposing would allow researchers to quickly identify data sets of interest, expand research questions beyond the most basic main-effects to more complex analyses such as gene-by-environment interactions and multivariate test incorporating multiple phenotypes, and perform larger meta-analyses easily by honing in on outcomes, exclusionary variables and covariates of interest with increased statistical power to detect genetic associations.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Information Explorer: a Suite of Tools for Cross-study Genetic Loci Discovery
Information Explorer: a Suite of Tools for Cross-study Genetic Loci Discovery
海外基金