Computational Methods for Assembling Multiple RNA-seq Samples
Computational Methods for Assembling Multiple RNA-seq Samples
批准号:
10350634
负责人:
MINGFU SHAO
金额:
$37.04万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-02-12 至 2026-01-31
关键词:
AddressAlgorithm DesignAlgorithmsBiologicalBiological AssayBiomedical ResearchBiotechnologyCommunitiesComplementComplicationComputing MethodologiesConsensusDataEffectivenessFoundationsGenesGenotype-Tissue Expression ProjectGraphIndividualIntronsLearningLengthMeasuresMethodsModelingNatureNeurofibrillary TanglesOutcomePhaseProtein IsoformsProtocols documentationRNA SplicingReproducibilityResearchReverse Transcriptase Polymerase Chain ReactionSamplingScallopSignal TransductionStatistical MethodsStatistical ModelsStructureTestingThe Cancer Genome AtlasTimeTissuesTranscriptTypologyUpdateVisionWeightbiological researchdata structuredesignexperimental studygenomic locusimprovedindexinglearning algorithmnovelopen sourcepreservationrepositorysimulationtranscriptometranscriptome sequencing
中文摘要
项目摘要/摘要
在许多生物和生物医学实验中,RNA-seq已成为研究基因活性的标准程序。
Rna-seq分析的第一步通常是量化每个转录本在
参考转录组。然而,研究表明,目前的转录组是不完整的,这限制了
表情量化的准确性。随着大规模的RNA-seq数据现已可用,高效和稳健的
构建转录组的方法是将一组rna-seq的全长表达转录本组装在一起。
样本,这是一种称为元组装的计算问题。这项提案解决了这一问题,旨在
为短读和长读rna-seq数据开发高效的元汇编器。
与之前的研究一样,我们开发了到目前为止最准确的单样本组装器Scallop(自然
生物技术,2017;缩写为RNA-seq)和Scallop-LR(长读为RNA-seq)。扇贝的核心
Scallop-LR是拼接图和阶段化路径的结合使用,它编码跨越两个以上的读取
顶点来表示读对齐,以及一种新的算法,该算法在保留剪接图的同时对剪接图进行分解
所有阶段化路径。这种数据结构和“保相”的思想为我们的
提出了元装配算法。
Meta-Assembly的关键是利用给定样本中共享和互补的信息。
我们建议在拼接图级别合并多个样本。具体地说,对于每个基因位点,我们构建
单个组合拼接图,通过合并各个拼接图并汇集它们的阶段化路径。至
将信息保存在单独的拼接图中,它们的类型将被编码为附加的阶段化路径。这个
因此,整个数据结构是空间高效且无损失的,并且可以通过管道传输到以下阶段保持
分解的算法。我们将专门化我们现有的相位保持算法来处理成对端
分相路径和长分相路径。最终,将开发统计方法来推断统计
每个单独组装的成绩单的重要性,并进行多个假设检验来控制
总体来说,发现了错误的成绩单。我们还提出了一种新的共识方法,即学习鉴别器来
自动为不同的元装配体实例选择最优算法。
这个项目的成果将是开源的、易于使用的、可重现的和准确的元汇编器,用于
分别是短读和长读的RNA-seq数据。这些元汇编器将实现更准确的
新亚型的鉴定和基因结构的注释。结合大规模的RNA-SEQ数据,
可以构建数据驱动的转录本,有利于下游研究,如RNA-Seq量化和
差异分析。
英文摘要
PROJECT SUMMARY / ABSTRACT
RNA-seq has become standard routine in many biological and biomedical experiments to study gene activities.
A very first step of RNA-seq analysis is usually to quantify the expression abundance of each transcript in the
reference transcriptome. However, studies have showed that current transcriptome is incomplete, which limits
the accuracy of expression quantification. As large-scale RNA-seq data are now available, an efficient and robust
way of constructing transcriptome is the assembly of the full-length expressed transcripts from a set of RNA-seq
samples, a computational problem known as meta-assembly. This proposal addresses this problem and aims to
develop efficient meta-assemblers for short-reads and long-reads RNA-seq data.
As previous studies, we have developed so far the most accurate single-sample assemblers Scallop (Nature
Biotechnology, 2017; for short-reads RNA-seq) and Scallop-LR (for long-reads RNA-seq). The core of Scallop
and Scallop-LR is the use of splice graph together with phasing paths, which encode reads spanning more than two
vertices, to represent reads alignment, and a novel algorithm that decomposes the splice graph while preserves
all phasing paths. This data structure and idea of “phase-preserving” provides algorithmic foundations for our
proposed meta-assembly algorithms.
The key of meta-assembly is to take advantage of shared and complementary information in the given samples.
We propose to combine multiple samples at the splice graph level. Specifically, for each gene locus, we construct
a single combined splice graph, through merging individual splice graphs and pooling their phasing paths. To
keep the information in individual splice graphs, their typologies will be encoded as additional phasing paths. The
entire data structure is therefore space-efficient and loss-free, and can be piped into following phasing-preserving
algorithms for decomposition. We will specialize our existing phasing-preserving algorithms to handle paired-end
phasing paths and long phasing paths. Eventually, statistical methods will be developed to infer the statistical
significance of each individual assembled transcript, and multiple hypothesis testing will be performed to control
overall falsely discovered transcripts. We also propose a new consensus-approach that learns a discriminator to
automatically select the optimal algorithm for different meta-assembly instances.
The outcomes of this project will be open-source, easy-to-use, reproducible and accurate meta-assemblers for
short-reads and long-reads RNA-seq data, respectively. These meta-assemblers will then enable more accurate
identification of novel isoforms and the annotation of gene structures. Combined with large-scale RNA-seq data,
data-driven transcriptomes can be constructed, benefiting downstream study such as RNA-seq quantification and
differential analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational Methods for Assembling Multiple RNA-seq Samples
-
批准号:10550251
-
项目类别:
-
资助金额:$36.96万
-
财政年份:2021
-
负责人:MINGFU SHAO
-
依托单位:
海外基金