课题基金 / 基金详情

Development of Statistical Methods for Analyzing Whole Genome Bisulfite Sequencing Experiment Data to Identify Differentially Methylated Regions

Development of Statistical Methods for Analyzing Whole Genome Bisulfite Sequencing Experiment Data to Identify Differentially Methylated Regions
开发分析全基因组亚硫酸氢盐测序实验数据以识别差异甲基化区域的统计方法
批准号:
1615789
负责人:
Tieming Ji
金额:
$23.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-01 至 2020-05-31

项目摘要

项目成果

Tieming Ji的其他基金

相似基金

相关文献

中文摘要
翻译
DNA 甲基化是对 DNA 的一种化学修饰,它提供有关如何以及何时打开或关闭基因的信息。因此,DNA 甲基化在许多生物过程中起着至关重要的作用。在很多情况下,样本之间的 DNA 甲基化是不同的。 例如:1)不同的器官(肝脏与大脑),2)健康与疾病(癌症表现出与健康组织不同的甲基化模式),以及3)环境反应(承受高温和干旱胁迫的植物)。 这些差异称为差异甲基化。最近,已经开发出同时查询 DNA 中数百万个甲基化位点的技术。  然而,这项技术带来了计算和统计方面的挑战。生物学领域的一个紧迫问题是如何分析以这种方式生成的大量数据,从而得出生物学相关的、统计上合理的结论,而不需要昂贵的计算设备。该项目的成果将提供检测差异甲基化的统计方法。然后专家可以利用这些信息来了解细胞功能、制定治疗干预措施或解决与气候变化等环境相关的问题。 将开发免费的统计工具,可供所有对分析差异甲基化感兴趣的科学家使用。这个数学和生物科学之间的跨学科研究项目也支持这些领域学生的培训。 DNA 甲基化是一种表观遗传修饰,指导基因表达和染色质构象。脊椎动物中最常见的 DNA 甲基化形式是在胞嘧啶碱基上添加甲基,紧接着是鸟嘌呤,称为 CpG 位点。人类基因组中大约有 3000 万个 CpG 位点,具有更大的甲基化状态组合。 差异甲基化区域是基因组中两个样本组(例如疾病组与正常组)之间 CpG 平均甲基化水平不同的区域。亚硫酸氢盐测序方法通常用于测量 CpG 位点的 DNA 甲基化。这项技术的扩展允许同时查询所有 CpG 位点(称为全基因组亚硫酸氢盐测序),这带来了计算和统计方面的挑战。其中包括提高在非常大的数据集中区分信号和噪声的能力,这是现代统计学当前的焦点。项目团队将开发统计方法来检测样本组之间的差异甲基化。具体来说,该研究项目中的方法通过根据基因组位置识别和解释甲基化位点之间的相关性来改进现有方法。这种模型假设确保统计结果具有生物学意义和可解释性。此外,这些方法将借用整个基因组的信息来提高估计和统计测试的可靠性。将使用高效且可扩展的算法开发和实施具有理论依据的贝叶斯模型,以确保其适用于各种高通量甲基化数据集。方法将被实施到软件工具中,并将免费提供给生物学和统计学研究人员。
英文摘要
DNA methylation is a chemical modification to DNA that imparts information on how and when genes should be turned on or off. As such DNA methylation plays a vital role in many biological processes. There are many instances in which DNA methylation between samples is different. Examples are: 1) different organs (liver vs brain), 2) health vs disease (cancer exhibits methylation patterns that are different from those of healthy tissue), and 3) environmental responses (plants that endure heat and drought stress). These differences are known as differential methylation. Recently, technology has been developed to simultaneously query the millions of methylation sites in DNA.  This technology, however, creates computational and statistical challenges. A pressing question in the field of biology is how to analyze the massive amount of data generated in this way so as to draw biologically relevant, and statistically sound conclusions, without requiring expensive computing equipment. The output of this project will provide statistical methods for detecting differential methylation. This information can then be used by specialists to understand cell function, develop therapeutic interventions, or tackle questions associated with environment such as climate change. Freely available statistical tools that can be used by all scientists interested in analyzing differential methylation will be developed. This interdisciplinary research project between the mathematical and biological sciences also supports the training of students in these fields. DNA methylation is an epigenetic modification that directs gene expression and chromatin conformation. The most common form of DNA methylation in vertebrates is an addition of a methyl group to a cytosine base that is directly followed by a guanine, which is referred to as a CpG site. There are approximately 30 million CpG sites in the human genome with an even larger combination of states of methylation. A differentially methylated region is a region in the genome where mean methylation levels of CpGs are different between two sample groups, such as disease versus normal. Bisulfite sequencing methods are typically used for measuring DNA methylation at CpG sites. Expansion of this technology to permit simultaneous query of all CpG sites, known as whole genome bisulfite sequencing, has created computational and statistical challenges. These include improving the capability to distinguish signals from noise in very large datasets, a current focus of much modern statistics. The project team will develop statistical methods to detect differential methylation between sample groups. Specifically, the methods in this research project improve existing approaches by recognizing and accounting for correlations among methylation sites based on their genomic locations. Such model assumptions ensure that statistical results are biologically meaningful and interpretable. In addition, the methods will borrow information across the whole genome to improve estimation and statistical testing reliability. Bayesian models with theoretical justifications will be developed and implemented using efficient and scalable algorithms to ensure their applicability to a wide variety of high-throughput methylation datasets. Methods will be implemented into software tools and will be freely available for biology and statistics researchers.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1007/s10815-019-01652-1
发表时间: 2019-12
期刊: Journal of Assisted Reproduction and Genetics
影响因子: 3.1
作者: [Yahan Li;P. Tríbulo;M. Bakhtiarizadeh;L. Siqueira;Tieming Ji;R. M. Rivera;P. Hansen]
通讯作者: Yahan Li;P. Tríbulo;M. Bakhtiarizadeh;L. Siqueira;Tieming Ji;R. M. Rivera;P. Hansen
DOI: 10.1016/j.jgg.2020.02.002
发表时间: 2020-02-20
期刊: JOURNAL OF GENETICS AND GENOMICS
影响因子: 5.9
作者: [Johnson, Adam F., Hou, Jie, Birchler, James A.]
通讯作者: Birchler, James A.
DOI: 10.1186/s12859-019-3141-6
发表时间: 2019-10-28
期刊: BMC BIOINFORMATICS
影响因子: 3
作者: [Xie, Jing, Ji, Tieming, Rivera, Rocio M.]
通讯作者: Rivera, Rocio M.
Collaborative Research: Development of New Statistical Methods for Genome-Wide Association Studies
  • 批准号:
    1853556
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2019
  • 负责人:
    Tieming Ji
  • 依托单位:
海外基金