课题基金 / 基金详情

Inference of Markovian Properties of Molecular Sequences Using Shotgun Reads and Applications

Inference of Markovian Properties of Molecular Sequences Using Shotgun Reads and Applications
使用鸟枪读取和应用推断分子序列的马尔可夫性质
批准号:
1518001
负责人:
Fengzhu Sun
金额:
$60.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31

项目摘要

项目成果

Fengzhu Sun的其他基金

相似基金

相关文献

中文摘要
翻译
高通量下一代测序(NGS)技术产生了大量的片段基因组序列,彻底改变了遗传和基因组学研究。使用NGS对来自不同环境的个体生物的自然混合物组成的数千个个体基因组和宏基因组进行了测序。这些发展在了解复杂疾病的遗传基础、环境对公众健康的影响、环境变化(如全球变暖和环境污染)对环境的影响以及检测包括病毒在内的病原体方面发挥着重要作用。开发分析方法以充分利用国家地质勘探局数据,对于促进公共卫生、改善环境和加强国家安全至关重要。尽管在NGS数据分析方面取得了重大进展,但目前可用的分析工具与通过NGS数据分析可以实现的全部潜力之间仍然存在很大差距。本研究项目旨在进一步推进最近开发的统计和计算方法,用于使用NGS reads进行基因组和宏基因组的比较,而不需要组装到基因组中,避免了许多使组装成为问题的陷阱。这项研究将使计算工具更加高效和强大,并将利用它们分析宏基因组数据,研究环境因素对海洋微生物群落的影响。算法和结果都将通过网络传播。本研究结果将对多种环境下的基因组学和宏基因组学研究具有重要意义。更详细地说,将开发基于NGS短读长推断分子序列马尔可夫性质的统计和计算方法,然后将这些方法用于研究个体基因组和宏基因组样本之间的关系。首先,给出了估计阶数和转移概率矩阵及其渐近分布的方法。本文还将研究变长马尔可夫链(VLMC)的推导方法。其次,考虑到序列的马尔可夫链(MC)特性,将开发新的无比对统计来研究基因组序列之间的关系。选择单词长度的迭代方法将被开发。第三,基于NGS reads的马尔可夫链模型将用于识别宏基因组群落中的物种或菌株,并基于MC模型对宏基因组样本进行比较。最后,我们将开发一套基于NGS数据推断MCs的计算机算法,并将其应用于基因组和宏基因组数据分析。该项目的广泛影响包括基于NGS数据的基因组和宏基因组比较计算工具,以及供公众使用的软件包,统计学和生物学多个学科的研究生和本科生培训,以及面向K-12教师和学生的外联讲座。
英文摘要
High throughput next generation sequencing (NGS) technologies generate enormous amounts of fragmented genome sequences, revolutionizing genetic and genomics research. Thousands of individual genomes and metagenomes consisting of natural mixtures of individual organisms from various environments have been sequenced using NGS. These developments play essential roles in understanding the genetic basis of complex diseases, the effects of environment on public health, the impacts of environmental changes such as global warming and pollution on the environments, and the detection of pathogens including viruses. Development of analytical methods to make full use of NGS data is essential in advancing public health, improving the environment, and strengthening national security. Although significant progress has been made in the analysis of NGS data, there are still wide gaps between the current available analytical tools and the full potential that can be achieved through the analysis of NGS data. This research project aims to further advance recently-developed statistical and computational methods for the comparison of genomes and metagenomes using NGS reads, without the need for assembly into genomes, avoiding many pitfalls that make assembly problematic. The research will make the computational tools more efficient and powerful and will employ them to analyze metagenomic data to study the effects of environmental factors on marine microbial communities. Both the algorithms and results will be disseminated through the web. The results from this study will be important for both genomics and metagenomics studies under a variety of environments.In more detail, statistical and computational methods for the inference of Markovian properties of molecular sequences based on NGS short reads will be developed and the methods will then be used to study the relationships among individual genomes and metagenomic samples. Firstly, methods to estimate the order and the transition probability matrix and their asymptotic distributions will be developed. Methods to infer variable length Markov chains (VLMC) will also be developed. Secondly, new alignment-free statistics taking into account the Markov chain (MC) properties of the sequences will be developed to study the relationships among genome sequences. Iterative approaches for choosing the word length will be developed. Thirdly, Markov chain models derived from NGS reads will be used to identify species or strains in metagenomic communities and to compare metagenomic samples based on the MC models. Finally, a suite of computer algorithms related to the inference of MCs based on NGS reads and applications to genome and metagenomic data analysis will be developed. The broad impacts of the project include computational tools for genome and metagenome comparison based on NGS data together with software packages for public usage, graduate and undergraduate training across multiple disciplines of statistics and biology, and outreach lectures for K-12 teachers and students.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MIM: Machine Learning, Systems Modeling, and Experimental Approaches to Understand the Universal Rules of Life of Microbiota Using Marine Time Series Data
  • 批准号:
    2125142
  • 项目类别:
    Standard Grant
  • 资助金额:
    $250.07万
  • 财政年份:
    2022
  • 负责人:
    Fengzhu Sun
  • 依托单位:
Computational and Mathematical Study in Protein Interactions and Functions
  • 批准号:
    0241102
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $103.6万
  • 财政年份:
    2003
  • 负责人:
    Fengzhu Sun
  • 依托单位:
国内基金
海外基金
信息网络环境下Markovian跳变系统安全运行控制方法研究
  • 批准号:
    62103011
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    杨宏燕
  • 依托单位:
Semi-Markovian切换系统的动态滑模控制及逗留时间和模式依赖滑模控制器研究
  • 批准号:
    61973075
  • 项目类别:
    面上项目
  • 资助金额:
    59.0万元
  • 批准年份:
    2019
  • 负责人:
    魏延岭
  • 依托单位:
几类广义Markovian跳变系统的控制方法研究
  • 批准号:
    61603055
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2016
  • 负责人:
    李丽
  • 依托单位:
中立型Markovian跳变随机微分方程系统稳定与控制
  • 批准号:
    61573007
  • 项目类别:
    面上项目
  • 资助金额:
    51.0万元
  • 批准年份:
    2015
  • 负责人:
    陈卫民
  • 依托单位: