课题基金 / 基金详情

Unlocking transcript diversity via differential analyses of splice graphs

Unlocking transcript diversity via differential analyses of splice graphs
通过剪接图的差异分析解锁转录本多样性
批准号:
8296802
负责人:
Jinze Liu
金额:
$42.5万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-05-23 至 2015-03-31

项目摘要

项目成果

Jinze Liu的其他基金

相似基金

相关文献

中文摘要
翻译
摘要 相同基因型和不同表型的细胞之间最基本的区别在于它们的转录组。了解两种转录本之间存在的RNA分子的差异,或特定分子丰度的变化,可以为疾病、发育和特化的分子机制提供有价值的见解。高通量测序提供了转录组的独特视图,其形式是从RNA分子中采样的数百万甚至数十亿个核苷酸序列的短读数。到目前为止,近1000个这样的RNA-seq数据集已经保存在NCBI基因表达总集中。除了衡量样本之间基因总体表达的差异外,还迫切需要在转录本水平上衡量表达的差异。可以通过rna-seq在不同人群中提取转录多样性显著变化的计算工具迫在眉睫。然而,从这些丰富的数据中重建完整范围的转录本亚型并不是一个解决的问题,因为在短阅读样本的规模上,亚型之间存在根本的歧义。 我们提出了一种新的方法来进行转录本的差异分析,该方法不依赖于全长转录本的重建,但可以准确地定位转录本的变异。我们的技术是数据驱动的,适用于任何转录组,只需要一个参考基因组,而不依赖于先验的基因结构注释。我们的研究计划建立在我们高度敏感和准确的MapSplice比对算法的基础上,从RNA-seq数据集构建表达式加权剪接图(ESG)。ESGs的大小可以比目前的RNA-seq数据集小三个数量级,但完全代表了这种数据集的实质性生物学内容。ESG表示支持高效的分析技术,可以直接识别和可视化样本之间具有统计意义的差异转录。提出了算法的一般化,以识别共同调节的剪接模式,这是生物途径分析和系统生物学分析的关键。 我们已经在包括生物学家、计算机科学家和统计学家在内的共同PI和Co-IS之间建立了一个持续的互动和协作的研究环境。建议的计算方法将使用从乳腺癌细胞系产生的RNA-SEQ数据进行测试和改进,然后进一步应用于关于肺癌发病机制、白血病干细胞和马关节软骨发育和修复(非模式哺乳动物生物)的三个精心策划的RNA-SEQ数据集。对差异表达转录异构体的实验验证将提高我们方法的准确性,并为与肺癌、白血病疾病和软骨细胞分化相关的替代异构体提出新的候选方案。 该软件将是开放源代码的,并将作为一套组件开发,这些组件可以单独使用,也可以整合到RNA-seq处理工作流程中。特别是,我们将把这些组件集成到本地服务器上托管的Galaxy云计算框架中。因此,这些方法将可供世界各地的研究人员使用。随着组件的成熟,它们可能会安装在世界各地的其他服务器上,以提供一种方便和安全的方式来分析转录本。 以适度的成本揭示转录组的动态将使细胞诊断学和生物医学研究发生革命性的变化。全基因组转录变异的测量提供了关于细胞身份和功能的详细分子信息的可能性,这将极大地扩展传统的组织学评估。基于云的方法访问可以将单个实验室变成小型基因组中心,并将使个别科学家能够在几天内评估RNA转录本之间的差异。我们的一套算法将使生物医学研究人员能够优先考虑候选基因或不同的基因本体类别,以进一步研究实验条件之间的差异转录和机制重要性。
英文摘要
Abstract A most basic difference between cells of the same genotype and different phenotype lies in their transcriptome. Understanding the difference between two transcriptomes in terms of the RNA molecules present in each, or changes in abundance of specific molecules, can offer valuable insight into the molecular mechanisms of disease, development, and specialization. High throughput sequencing provides a unique view of the transcriptome in the form of millions or even billions of short reads of nucleotide sequences sampled from the RNA molecules. To date, nearly 1000 such RNA-seq datasets have already been deposited in the NCBI Gene Expression Omnibus. Beyond measuring differences in overall expression of genes between samples, there is a critical need to measure differences in expression at the transcript level. Computational tools that can extract significant changes in transcript diversity across populations with RNA-seq are in immediate demand. However, reconstructing the full extent of transcript isoforms from this wealth of data is not a solved problem because of fundamental ambiguities between isoforms at the scale of the short read samples. We propose a novel approach to the differential analysis of transcriptomes that does not depend on the reconstruction of the full-length transcripts, and yet can accurately pinpoint the variation of transcriptomes. Our techniques are data-driven and applicable to any transcriptome, requiring only a reference genome, and do not depend on a priori gene structure annotations. Our research program builds on our highly sensitive and accurate MapSplice alignment algorithm to construct expression weighted splice graphs (ESG) from RNA-seq datasets. ESGs can be three orders of magnitude smaller in size than current RNA-seq datasets, yet fully represent the substantive biological content of such datasets. The ESG representation supports highly efficient analysis techniques that can directly identify and visualize statistically significant differential transcription between samples. Generalizations of the algorithms are proposed to identify co-regulated splicing patterns that are keys for biological pathway analyses and systems biology analyses. We have established an ongoing interactive and collaborative research environment among the co-PIs and Co-Is which include the biologists, computer scientists and statistician. The proposed computational methods will be tested and refined using RNA-seq data generated from breast cancer cell lines before being further applied to three well curated RNA-seq datasets on lung cancer pathogenesis, stem cells in leukemia, and equine articular cartilage development and repair (a non-model mammalian organism). Experimental validation of differentially expressed transcript isoforms will both improve the accuracy of our methods, as well as propose novel candidates for alternative isoforms associated with lung cancer,and leukemia diseases, and chondrocyte differentiation. The software will be open-source and will be developed as a set of components that can be used on their own or integrated into RNA-seq processing workflows. In particular we will integrate the components into the Galaxy cloud computing framework hosted on a local server. As such the methods will be available to researchers worldwide. As components mature they may be installed in other servers worldwide to provide a convenient and secure way to analyze transcriptomes. Unveiling the dynamics of the transcriptome at modest cost will revolutionize cellular diagnostics and biomedical research. Genome-wide measurement of transcription variants offers the potential for detailed molecular information about cellular identity and function that will greatly expand traditional histological assessment. Cloud-based access to the methods can turn individual laboratories into small genome centers and will enable individual scientists to assess differences among RNA transcriptomes in a matter of days. Our suite of algorithms will enable biomedical researchers to prioritize candidate genes or different gene ontology categories to investigate further for differential transcription and mechanistic importance between experimental conditions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bioinformatics Core
Unlocking transcript diversity via differential analyses of splice graphs
Unlocking transcript diversity via differential analyses of splice graphs
海外基金