课题基金 / 基金详情

项目摘要

项目成果

Sreeram Kannan的其他基金

相似基金

相关文献

中文摘要
翻译
 描述(申请人提供):RNA-Seq使转录组学发生了革命性的变化,是近年来发明的最重要的高通量测序方法之一。关键的计算问题是从头组装:从数千万到数亿的短读数重建转录及其丰度。由于几个因素的共同作用,这个问题是具有挑战性的:大量不同的转录本(数万),由于选择性剪接而在转录本之间长时间重复,不同转录本之间的丰度差异很大,以及存在阅读错误。现有的汇编器大多是基于启发式考虑而设计的,并且实现了导致不可靠的转录组重建的特殊方法。一个准确的RNA-Seq组装器将使我们能够更准确地识别癌症转录本中的融合,更好地在模式生物和非模式生物中进行基因注释,并更完整地分析驱动发展和调控程序的替代剪接的动态。在这项建议中,我们提供了一种基于信息论原理的RNA-Seq汇编器设计的系统方法。我们首先确定条件数据,以保证有足够的信息来重建转录组,然后提出一个可以用最少的信息重建的组装算法。该算法最佳地使用可用的读取信息来解析重复和消除歧义异构体。信息论方法得出的一个关键见解是,不同成绩单之间的广泛不同丰度,而不是复杂的情况,实际上可以用作不同成绩单的签名,以消除它们之间的歧义。根据我们最初的想法,我们已经建立了一个初始原型,并将其与几个现有的软件进行了比较,包括真实数据和模拟数据。令人鼓舞的结果证明,我们将在资金支持期间全面开发、实施和评估的方法可以显著超过现有软件。作为拟议项目的一部分,将设计其他功能,如混合的短/长阅读组装、基因组辅助组装和对多个RNA样本的联合处理,并将其纳入软件。
英文摘要
 DESCRIPTION (provided by applicant): RNA-Seq has revolutionized transcriptomics and is one of the most important high-throughput sequencing assays invented in recent years. The key computational problem is that of de novo assembly: the reconstruction of the transcripts and their abundances from tens to hundreds of millions of short reads. The problem is challenging due to a confluence of several factors: large number of different transcripts (tens of thousands), long repeat across transcripts due to alternative splicing, widely varying abundances across transcripts, and the presence of read errors. Existing assemblers are mostly designed based on heuristic considerations and implement ad hoc methods that lead to unreliable transcriptome reconstructions. An accurate RNA-Seq assembler would enable more accurate identification of fusions in cancer transcriptomes, better gene annotations in model and non-model organisms, and more complete analyses of the dynamics of alternative splicing driving developmental and regulatory programs. In this proposal, we offer a systematic approach to the design of RNA-Seq assemblers based on information theoretic principles. We start by determining conditions data that guarantee that there enough information to reconstruct the transcriptome, and then propose an assembly algorithm that can reconstruct with the minimal information. This algorithm optimally uses the available read information to resolve repeats and disambiguate isoforms. A key insight derived from the information theoretic approach is that widely varying abundances across transcripts, rather than a complication, can actually be exploited as signatures of different transcripts to disambiguate among them. Based on our initial ideas, we have built, evaluated and compared an initial prototype with several existing software, on both real and simulated data. The encouraging results provide evidence that our approach, which we will fully develop, implement and evaluated during the funded period, can significantly outperform existing software. Additional functionalities such as mixed short/long read assembly, genome-assisted assembly and joint processing of multiple RNA samples, will be designed and incorporated into the software as part of the proposed project.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Defining causal roles of genomic variants on gene regulatory networks with spatiotemporally-resolved single-cell multiomics
  • 批准号:
    10297331
  • 项目类别:
  • 资助金额:
    $121.0万
  • 财政年份:
    2021
  • 负责人:
    Sreeram Kannan
  • 依托单位:
Defining causal roles of genomic variants on gene regulatory networks with spatiotemporally-resolved single-cell multiomics
  • 批准号:
    10474569
  • 项目类别:
  • 资助金额:
    $121.0万
  • 财政年份:
    2021
  • 负责人:
    Sreeram Kannan
  • 依托单位:
Algorithms and Software for Provably Accurate De Novo RNA-Seq Assembly
海外基金