课题基金 / 基金详情

Genome-Transcription-Phenome-Wide Association: a new paradigm for association stu

Genome-Transcription-Phenome-Wide Association: a new paradigm for association stu
全基因组-转录-表型组关联:关联研究的新范式
批准号:
7845048
负责人:
Sally E Wenzel
金额:
$51.57万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-05-15 至 2014-03-31

项目摘要

项目成果

Sally E Wenzel的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):许多复杂的疾病综合征由大量高度相关而非独立的临床表型组成。这些综合征之间的差异涉及大量基因组变异的复杂相互作用,这些变异在调控网络的背景下扰乱了疾病相关基因的功能,而不是单独的。因此,揭示因果遗传变异和理解随之而来的细胞和组织转化的机制需要一种分析,共同考虑基因组(G)、转录组(T)和表型(P)内部和之间的元素和模块的epistatic、多效性和可塑性相互作用。大多数传统方法侧重于每个个体标记基因型和每个单一表型之间的关联;他们的统计能力有限,忽视了复杂的省略结构。我们提出了一个系统的尝试,对方法论发展的很大程度上未被探索,但实际上重要的问题之间的“-组”的结构关联。结构化关联分析不是单独测试每个SNP的关联,然后通过多个假设检验进行校正,而是确定实体组之间的关联,每个实体组都有自己不可忽视的复杂结构,例如具有高LD的SNP块,相同途径中的基因模块,以及属于疾病临床描述符系统的表型集群。我们将开发一个数学严谨、计算高效的机器学习平台和软件,以解决涉及揭示G、T和P组中疾病相关元素之间相互作用的方法学挑战。我们的技术创新包括用于单倍型推断的新颖统计模型和算法,重组热点检测,基因网络和表型网络推断,混合关联映射,最重要的是,一系列新的结构化回归技术,如图正则化回归,图引导融合套索和扩展,这些技术对G, T和P组中结构元素之间的关联函数进行功能近似。并对一致性和稀疏性有可证明的保证。我们设想我们所提出的研究将为复杂疾病的关联研究开辟一个新的范式,这将有助于:1)基因组内和基因组间的数据整合,以进行关联定位和疾病基因/途径的发现;2)深入探索不同基因组数据的内部结构,从而推断出由于统计能力弱而在非结构化分析中不可能检测到的隐性关联。3)联合统计推断DNA变异导致复杂性状变异如何在分子网络中流动的机制和途径,推断基因功能在分子网络中的条件特异性状态;4)开发更快、自动化的计算算法,具有更大的可扩展性和鲁棒性,可用于大规模组间分析,更方便的软件包和用户界面。所有的软件工具都将免费提供给公众。公共卫生相关性:我们提出了一个系统的方法开发尝试,以解决基因组、转录组和现象组中疾病相关元素之间的结构化关联映射,这是一个很大程度上未被探索但实际上很重要的问题。由于许多复杂的疾病涉及复杂的表型,这些表型是由复杂和相互依赖的基因组变异引起的基因调控的分子网络的复杂扰动的结果,因此在多组学水平上进行结构化关联分析不仅是需要的,而且是必要的,但它超出了传统方法的掌握范围,需要我们提出的方法创新。描述这种相互作用可以为复杂疾病提供更全面的遗传和分子观点,这可能导致识别疾病过程的基因;此外,这种方法将使我们能够制定关于这些基因在疾病发病机制中的作用的假设,并为多变量临床表型开发改进的诊断生物标志物。
英文摘要
DESCRIPTION (provided by applicant): Many complex disease syndromes consist of a large number of highly related, rather than independent, clinical phenotypes. Differences between these syndromes involve the complex interplay of a large number of genomic variations that perturb the function of disease-related genes in the context of a regulatory network, rather than individually. Thus unraveling the causal genetic variations and understanding the mechanisms of consequent cell and tissue transformation requires an analysis that jointly considers the epistatic, pleiotropic, and plastic interactions of elements and modules within and between the genome (G), transcriptome (T), and phenome (P). Most conventional methods focus on associations between every individual marker genotype and every single phenotype; they have limited statistical power and overlook the complex omit structures. We propose a systematic attempt on methodological development for the largely unexplored but practically important problem of structured associations between the "-omes". Rather than testing each SNP separately for association and then applying a correction by multiple hypothesis test, a structured association analysis identifies associations between groups of entities each with its own sophisticated structure that can not be ignored, such as blocks of SNPs with high LD, modules of genes in the same pathway, and clusters of phenotypes belong to a system of clinical descriptors of a disease. We will develop a mathematically rigorous and computationally efficient machine learning platform and software to address the methodological challenges involved with unraveling the interplay between disease-relevant elements in the G, T, and P omes. Our technical innovations include novel statistical models and algorithms for haplotype inference, recombination hotspot detection, gene network and phenotype network inference, admixture association mapping, and most importantly, a family of new structured regression techniques such as the graph-regularized regression, graph- guided fused lasso and extensions, that perform functional approximations to the association functions among structural elements in the G, T, and P omes, and have provable guarantee on consistency and sparsistency. We envisage our proposed research will open a new paradigm for association studies of complex diseases, which facilitates: 1) Intra- and inter-omic integration of data for association mapping and disease gene/pathway discovery, 2) Thorough explorations of the internal structures within different omic data, so that cryptic associations that are not possibly detectable in unstructured analysis due to their weak statistical power can be now inferred. 3) Joint statistical inference of mechanisms and pathways of how variations in DNA lead to variations in complex traits flows through molecular networks, and inference of condition-specific state of gene function in the molecular networks, and 4) Development of faster and automated computational algorithm with greater scalability and robustness to large-scale inter-omic analysis, and more convenient software package and user interface. All the software tools will be made available for free to the public. PUBLIC HEALTH RELEVANCE: We propose a systematic attempt on methodological development for the largely unexplored but practically important problem of structured association mapping between disease-relevant elements in the genome, transcriptome, and phenome. Since many complex diseases involve composite phenotypes that are the outcome of intricate perturbation of molecular network underlying gene regulatory resulted from complex and interdependent genome variations, structured association analysis at multi-omic level is not only needed, but also necessary, but it is beyond the grasp of convention methods and requires the methodological innovations we propose. Characterizing such interactions can provide a more comprehensive genetic and molecular view of complex diseases, which may lead to the identification of genes underlying disease processes; in addition, such an approach will allow us to formulate hypotheses regarding the roles of these genes with respect to disease pathogenesis, and to develop improved diagnostic biomarkers for multivariate clinical phenotypes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Type-2 or Not Type-2: That is the (Therapeutic) Question
Type-2 or Not Type-2: That is the (Therapeutic) Question
Type-2 or Not Type-2: That is the (Therapeutic) Question
Type-2 or Not Type-2: That is the (Therapeutic) Question
海外基金