Statistical methods for interpretation of genetic variants by gene regulatory networks
Statistical methods for interpretation of genetic variants by gene regulatory networks
批准号:
10710939
负责人:
Zhana Duren
金额:
$35.59万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-08 至 2028-08-31
关键词:
AddressAffectAllelesCellsComprehensionComputational BiologyDataData SetDiseaseElementsGene ExpressionGene Expression RegulationGenesGeneticGenomeGenotypeGenotype-Tissue Expression ProjectGoalsHealthHumanIndividualJointsLinkage DisequilibriumMapsMethodsModelingNetwork-basedPersonsPhenotypeQuantitative Trait LociRegulator GenesRegulatory ElementResearchSamplingSourceStatistical MethodsUntranslated RNAVariantcausal variantdisease mechanisms studyepigenomeepigenomicsgene regulatory networkgenetic varianthuman reference genomepopulation basedprecision drugsprecision medicineresponsetranscription factor
中文摘要
项目摘要
一个人的基因组通常包含数百万个变异,这些变异代表了这个人与其他人之间的差异。
基因组和参考人类基因组。解释这些变异如何导致疾病,
理解它们与表型的统计学关联的机制是
计算生物学和遗传学。这些问题并不容易解决,因为超过90%的
疾病相关的变异是在非编码区,具有高度特异性的细胞环境调控,
功能,我们对它的理解有限。这个项目的长期目标是解释
非编码遗传变异如何影响细胞环境依赖性基因调控
网络和影响表型。表达数量性状基因座(eQTL)定位和基因
调控网络(GRNs)是两种常见的解释遗传调控机制的方法,
变体。eQTL定位通过基于群体的关联将非编码区的变异与基因联系起来
study. GRNs提供了关于控制靶基因的环境特异性表达的顺式调节元件的信息,
基因,以及作用于这些元件的转录因子的信息。基于GRN的变体
解释是对eQTL作图的补充,并有可能克服eQTL的局限性
eQTL定位存在以下问题:(1)eQTL定位偏向于常见等位基因;(2)eQTL定位不能区分
(3)检测trans-eQTL的能力较低。大多数以前的监管
基于ENCODE数据的分析研究不包括个人基因分型数据,大多数eQTL定位
研究不包括监管信息。eQTL和GRN的联合建模将实现高精度
和机械变体解释。然而,这种分析所需的数据集-匹配的基因
来自相同个体的表达、表观基因组和基因分型数据-对于大型人类
sample.可用的数据集是跨个体配对基因分型和基因表达数据,例如GTEx
数据,以及跨细胞背景配对的基因表达和表观基因组学数据,例如ENCODE数据。这些
在单小区级也可获得两种类型的成对数据。为了实现我们的长期目标,我们将开发
统计方法,以整合这些不匹配的数据集(无论是散装或单细胞)从不同的来源(1)
推断高准确度上下文特异性GRN以连接变体、转录因子、顺式调节元件,以及
靶基因;和(2)检测调节靶基因的trans-eQTL。这些方法可以扩展到解释
疾病相关变异,识别因果变异,并推断个性化药物反应,以提供指导
用于精准医疗这个项目是精准医疗的基础,它将增加我们对
基因变异如何影响表型
英文摘要
Project Summary
A person’s genome typically contains millions of variants which represent the differences between this personal
genome and the reference human genome. Interpretation of how these variants cause diseases and
understanding the mechanism(s) of their statistical associations to phenotype are crucial problems in
computational biology and genetics. The problems are not straightforward to address because over 90% of
disease-associated variants are in non-coding regions that have highly specific cellular context regulatory
functions and about which we have limited comprehension. The long-term goal of this project is to explain
mechanistically how non-coding genetic variants affect cellular context-dependent gene regulatory
networks and influence phenotypes. Expression quantitative trait locus (eQTL) mapping and Gene
regulatory networks (GRNs) are two common approaches for interpreting regulatory mechanisms of genetic
variants. eQTL mapping connects variants in non-coding regions to genes by a population-based association
study. GRNs provide information on the cis-regulatory elements that control context-specific expression of target
genes, and information about the transcription factors that act on these elements. GRN-based variant
interpretation is complementary to eQTL mapping and has the potential to overcome the limitations of eQTL
mapping, which are: (1) eQTL mapping is biased for common alleles; (2) eQTL mapping cannot distinguish
variants in strong linkage disequilibrium; and (3) the power to detect trans-eQTL is low. Most previous regulatory
analysis research based on ENCODE data did not include personal genotyping data, and most eQTL mapping
research did not include regulatory information. Joint modelling of eQTLs and GRNs would enable high-accuracy
and mechanistic variant interpretation. However, the required dataset for such analysis - matched gene
expression, epigenome, and genotyping data from the same individuals - are not available for a large human
sample. Available datasets are cross-individual paired genotyping and gene expression data, such as GTEx
data, and cross-cellular-contexts paired gene expression and epigenomics data, such as ENCODE data. These
two types of paired data are also available at the single cell level. To achieve our long-term goal, we will develop
statistical methods to integrate these unmatched datasets (either bulk or single cell) from different sources to (1)
infer high accuracy context-specific GRNs to connect variants, transcription factors, cis-regulatory elements, and
target genes; and (2) detect trans-eQTLs that regulate target genes. These methods can be extended to interpret
disease-associated variants, identify causal variants, and infer personalized drug response to provide guidance
for precision medicine. This project is fundamental for precision medicine, and it will increase our understanding
of how genetic variants contribute to phenotype.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金