课题基金 / 基金详情

IMPROVE GENOME ANNOTATION USING MULTIPLE SEQUENCE ALIGNMENT RELIABILITY SCORES

IMPROVE GENOME ANNOTATION USING MULTIPLE SEQUENCE ALIGNMENT RELIABILITY SCORES
使用多序列比对可靠性评分改进基因组注释
批准号:
8229724
负责人:
Jian Ma
金额:
$19.13万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-02-22 至 2014-01-31

项目摘要

项目成果

Jian Ma的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):比较基因组学是发现人类基因组功能元素的强大工具。跨物种比较基因组学的基础是多序列比对(MSA)。尽管在过去十年中取得了进展,但MSA仍然是一项艰巨的任务,而且容易出错。比对误差会直接影响下游分析,并可能导致不正确的生物学结论。许多生物医学研究人员已经在Ensembl浏览器和UCSC基因组浏览器中使用公开可用的预先计算的msa来进行各种比较基因组分析。但是这些msa有错误。然而,用户通常不会询问对齐的可靠性,或者不知道如何定量地度量可靠性。初步研究表明,目前UCSC基因组浏览器中相当数量的保守元件可能是由不可靠的msa引入的假阳性。有问题的比对对基因组注释的影响可能比我们想象的要大得多。在这个项目中,将开发新的基于概率抽样的分数来测量多个序列比对。将采用上下文依赖的替代模型和更现实的模型来处理插入和缺失,以便将该方法应用于基因组大范围,并具有处理大量序列的深度比对的能力。此外,比对可靠性分数将用于改进基因组注释。ENCODE项目中人类基因组功能元素的数据将用于改进该模型。该方法还将被应用于拾取更多的功能元素,这些元素最初是由于对齐中的不确定性而错过的。对其他类型的基因组注释(如RNA基因、正选择)的改进也将进行探讨。这些捕获msa可靠性的新方法将大大减少比较基因组学分析中由比对错误引入的假阳性。如果成功,比较基因组学的一般方法可以得到改进,依靠基于msa的计算研究的实验室实验将更加有效。该项目的结果将被整合到UCSC基因组浏览器中,以使其他使用msa进行各种生物医学发现的研究人员受益。该方法将对ENCODE、TCGA、Genome 10K和其他大型比较基因组学项目产生重要影响。这个计算生物学的创新项目将对基因组学社区产生重要影响,并使生物医学研究取得进展。
英文摘要
DESCRIPTION (provided by applicant): Comparative genomics is a powerful tool to discover functional elements in the human genome. The foundation of cross-species comparative genomics is multiple sequence alignment (MSA). Despite of the progress in the past decade, MSA is still a difficult task and error-prone. The alignment errors can directly affect the downstream analyses and may lead to incorrect biological conclusions. Many biomedical researchers have been using publicly available, precomputed MSAs in the Ensembl Browser and the UCSC Genome Browser to conduct various comparative genomic analyses. But these MSAs have errors. However, users often do not ask how reliable the alignment is or do not know how to quantitatively measure the reliability. Preliminary study suggests that a considerable amount of conserved elements in the current UCSC Genome Browser might be false positives introduced by unreliable MSAs. The impact of problematic alignment on the genome annotation may be much greater than we thought. In this project, novel probabilistic sampling-based scores to measure multiple sequence alignment will be developed. Context- dependent substitution models and more realistic models to handle insertions and deletions will be employed in order to apply the method to the genome wide scale with the capability of dealing with deep alignments from large number of sequences. In addition, the alignment reliability scores will be used to improve genome annotation. The data of functional elements in the human genome from the ENCODE project will be used to refine the model. The method will also be applied to pick up more functional elements that are originally missed because of the uncertainty in the alignment. Improvement on other types of genome annotations (e.g. RNA gene, positive selection) will also be explored. These new methods that capture MSAs reliability will greatly reduce the false positives in comparative genomics analysis that are introduced by alignment errors. If successful, the general methodology of comparative genomics can be improved and laboratory experiments that rely on computational studies based on MSAs will be much more effective. Results from the project will be integrated into the UCSC Genome Browser to benefit other researchers who use MSAs for various biomedical discoveries for disease related signatures. The method will potentially have meaningful impact on ENCODE, TCGA, Genome 10K, and other large-scale comparative genomics projects. This innovative project in computational biology will potentially have important impact on the genomics community and enable advancement in biomedical research. PUBLIC HEALTH RELEVANCE: This project will develop novel methods for analyzing multiple sequence alignments in order to enhance our ability to annotate functional elements in the human genome. The proposed methods will help us better understand the human genome and facilitate the discovery of disease related signatures. This computational biology project will support the rapid advancement in genomics and biomedical research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Spatial omics technologies to map the senescent cell microenvironment
  • 批准号:
    10384585
  • 项目类别:
  • 资助金额:
    $35.79万
  • 财政年份:
    2021
  • 负责人:
    Jian Ma
  • 依托单位:
Spatial omics technologies to map the senescent cell microenvironment
  • 批准号:
    10907057
  • 项目类别:
  • 资助金额:
    $86.84万
  • 财政年份:
    2021
  • 负责人:
    Jian Ma
  • 依托单位:
Scalable Cancer Genomics via Nanocoding and Sequencing
Scalable Cancer Genomics via Nanocoding and Sequencing
海外基金