课题基金 / 基金详情

项目摘要

项目成果

MINGFU SHAO的其他基金

相似基金

相关文献

中文摘要
翻译
项目总结/摘要 RNA-seq已成为许多生物学和生物医学实验中研究基因活性的标准程序。 RNA-seq分析的第一步通常是定量每个转录物在细胞中的表达丰度。 参考转录组。然而,研究表明,目前的转录组是不完整的,这限制了 表达定量的准确性。随着大规模RNA-seq数据的出现, 构建转录组的方法是从一组RNA-seq序列中组装全长表达的转录物, 样本,一个被称为元组装的计算问题。本建议针对这一问题,旨在 为短读段和长读段RNA-seq数据开发高效的元组装器。 作为以前的研究,我们已经开发了迄今为止最准确的单样本汇编器Scallop(自然 Biotechnology,2017;用于短读段RNA-seq)和Scallop-LR(用于长读段RNA-seq)。扇贝的核心 Scallop-LR是拼接图与定相路径一起使用,其编码跨越两个以上的读段 顶点,以表示读取对齐,以及一种新的算法,该算法分解拼接图,同时保留 所有相位路径。这种数据结构和“相位保持”的思想为我们的算法提供了算法基础。 提出了元组装算法。 元组装的关键是利用给定样本中的共享和互补信息。 我们建议在拼接图级别上联合收割机组合多个样本。具体来说,对于每个基因位点,我们构建 单个组合拼接图,通过合并各个拼接图并汇集它们的定相路径。到 将信息保持在单独的拼接图中,它们的类型学将被编码为附加的定相路径。的 因此,整个数据结构是空间有效的和无损失的,并且可以被输送到以下相位保持中 分解算法我们将专门我们现有的相位保持算法来处理双端 定相路径和长定相路径。最终,将开发统计方法来推断统计 每个单独的组装转录本的显著性,并将进行多重假设检验以控制 所有被发现的错误记录我们还提出了一种新的共识方法, 自动为不同的元装配实例选择最佳算法。 该项目的成果将是开源的,易于使用的,可复制的和准确的元汇编器, 短读段和长读段RNA-seq数据。这些元汇编器将使更准确的 新的异构体的鉴定和基因结构的注释。结合大规模RNA-seq数据, 可以构建数据驱动的转录组,有利于下游研究,如RNA-seq定量, 差异分析
英文摘要
PROJECT SUMMARY / ABSTRACT RNA-seq has become standard routine in many biological and biomedical experiments to study gene activities. A very first step of RNA-seq analysis is usually to quantify the expression abundance of each transcript in the reference transcriptome. However, studies have showed that current transcriptome is incomplete, which limits the accuracy of expression quantification. As large-scale RNA-seq data are now available, an efficient and robust way of constructing transcriptome is the assembly of the full-length expressed transcripts from a set of RNA-seq samples, a computational problem known as meta-assembly. This proposal addresses this problem and aims to develop efficient meta-assemblers for short-reads and long-reads RNA-seq data. As previous studies, we have developed so far the most accurate single-sample assemblers Scallop (Nature Biotechnology, 2017; for short-reads RNA-seq) and Scallop-LR (for long-reads RNA-seq). The core of Scallop and Scallop-LR is the use of splice graph together with phasing paths, which encode reads spanning more than two vertices, to represent reads alignment, and a novel algorithm that decomposes the splice graph while preserves all phasing paths. This data structure and idea of “phase-preserving” provides algorithmic foundations for our proposed meta-assembly algorithms. The key of meta-assembly is to take advantage of shared and complementary information in the given samples. We propose to combine multiple samples at the splice graph level. Specifically, for each gene locus, we construct a single combined splice graph, through merging individual splice graphs and pooling their phasing paths. To keep the information in individual splice graphs, their typologies will be encoded as additional phasing paths. The entire data structure is therefore space-efficient and loss-free, and can be piped into following phasing-preserving algorithms for decomposition. We will specialize our existing phasing-preserving algorithms to handle paired-end phasing paths and long phasing paths. Eventually, statistical methods will be developed to infer the statistical significance of each individual assembled transcript, and multiple hypothesis testing will be performed to control overall falsely discovered transcripts. We also propose a new consensus-approach that learns a discriminator to automatically select the optimal algorithm for different meta-assembly instances. The outcomes of this project will be open-source, easy-to-use, reproducible and accurate meta-assemblers for short-reads and long-reads RNA-seq data, respectively. These meta-assemblers will then enable more accurate identification of novel isoforms and the annotation of gene structures. Combined with large-scale RNA-seq data, data-driven transcriptomes can be constructed, benefiting downstream study such as RNA-seq quantification and differential analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational Methods for Assembling Multiple RNA-seq Samples
海外基金