Bilateral BBSRC-NSF/BIO: ABI Innovation: Data-driven hierarchical analysis of de novo transcriptomes
Bilateral BBSRC-NSF/BIO: ABI Innovation: Data-driven hierarchical analysis of de novo transcriptomes
批准号:
1564917
负责人:
Robert Patro
金额:
$31.06万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-01 至 2019-06-30
中文摘要
这个项目将决定当研究人员研究基因如何在没有完整基因组DNA序列的生物体中表达时,需要使用什么方法。基因提供了有机体在合适的条件下可以执行的功能的潜力,而基因的表达是为了确保细胞能够履行它们在环境中生存所需的功能。每个基因都可以有一个表达形式家族,如果没有基因本身的序列进行比较,就很难理清它们之间的关系,但对于大多数生物体来说,这种序列是缺失的。从这样的实验中获得的基因表达的观点通常是支离破碎的、不完整的,并且很难分析,这是本研究解决的中心问题。一旦制定出可靠的分析方法,这项研究将产生使用这些方法的软件工具,并被设计为考虑到常见的错误和数据的缺失部分。作为检查,重建的基因序列将与相关生物中的已知基因进行比较,并将使用强大的关系来指导对基因功能的预测。所有的预测都会带有一个数字,表明信息的不确定性。为了让这项研究以一个生物学上有趣的问题为基础,这些方法和软件将被用于研究使用一种不寻常的能量转换形式--C4光合作用的植物,并将其与植物使用的更常见的形式进行比较,以便研究每种植物使用的不同遗传机制,包括调节。预测将通过湿法实验室实验进行验证,然后分析方法将根据需要进行改进。通过为开展和分析类似的实验提供专家建议,将创建和培育一个与研究人员分享兴趣的在线社区。该项目将包括开发一套新的方法和一套综合的工具,用于分析从头转录。目前有一些工具旨在处理从头转录组分析管道的不同阶段(例如组装、聚类、表达量化和差异表达测试),但这些工具都不能很好地整合、有原则和有效地解决这一困难的挑战。在这个项目中开发和验证的方法将提供一个最先进的管道,用于提出和回答一系列相关的生物学问题,这些问题涉及转录本、基因和功能模块是如何差异表达和调控的;特别是在缺乏参考基因组的生物背景下。该项目将导致开发新的方法,用于在从头组装中对重叠群进行数据驱动的聚类。聚类将根据重叠群之间的序列和表达水平的相似性来决定,考虑到错误组装的已知特征,以确定来自相同基础转录本的重叠群。转录本将根据预测的、共享的外显子结构进行分组,以发现亚基因特征、基因和基因家族。这一过程的结果将是转录组的分层模型。将开发对这些层次模型进行有效量化和差异表达分析的新方法,包括将量化不确定性的测量传播到下游分析的方法健全的方法。最后,这些工具将通过大规模重新分析现有的从头转录组数据来验证,目的是阐明参与C4光合作用的遗传调控元件的身份。这次重新分析将利用在单一、多尺度模型中分析所有非模型C4RNA-SEQ实验数据所开发的方法的更高的准确性和效率。该模型的输出将是一个系统,用于对详细的分子研究的候选调控元件进行定量优先排序。该项目还包括更广泛的影响目标,将为本科生提供研究机会,并将创建一个积极维护和包容的在线社区,以从头开始转录组分析和实验设计的最佳实践为中心。有关该项目进展的更多信息,包括正在开发的相关软件,请访问https://combine-lab.github.io/txome.
英文摘要
This project will determine what methods are needed when researchers are studying how genes are expressed in organisms that do not have a completed DNA sequence of the genome. Genes provide the potential for functions an organism can carry out under the right conditions, and genes are expressed to make sure cells can perform the functions they need to live in their environment. Each gene can have a family of expressed forms, and their relationships can be very difficult to sort out without the sequence of the gene itself for comparison, but that sequence is missing for most organisms. The view of gene expression obtained from such experiments is typically fractured, incomplete, and difficult to analyze, the central problem this research addresses. Once reliable analysis methods are worked out, this research will produce software tools that use the methods and are designed to take common errors and missing parts of the data into account. As a check, the reconstructed gene sequences will be compared to known genes in related organisms, and strong relationships will be used to guide predictions of the genes' functions. All predictions will carry with them a number indicating the uncertainty of the information. To ground the research in a biologically interesting question the methods and software will be used to study plants that use an unusual form of energy conversion, C4 photosynthesis, and compare it to the more common form used by plants in order to investigate the different genetic mechanisms, including regulation, used by each type of plant. Predictions will be tested with wet-lab experiments, and analysis methods will then be improved as needed. An on-line community that shares the interests of the researchers will be created and fostered by providing expert advice for carrying out and analyzing similar experiments. This project will consist of the development of a novel collection of methods, and an integrated set of tools, for the analysis of de novo transcriptomes. There are currently a number of tools that aim to tackle different phases of the de novo transcriptome analysis pipeline (e.g. assembly, clustering, expression quantification and differential expression testing), but none of these provide a well-integrated, principled and efficient approach to this difficult challenge. The methods developed and validated in this project will provide a state-of-the-art pipeline for posing and answering a host of relevant biological questions about how transcripts, genes, and functional modules are differentially expressed and regulated; specifically, in the context of organisms for which a reference genome is lacking. The project will result in the development of novel methods for the data-driven clustering of contigs in de novo assemblies. Clustering will be decided on the basis of sequence and expression-level similarities between contigs, accounting for known hallmarks of mis-assemblies to determine contigs that arise from the same underlying transcript. Transcripts will be grouped by predicted, shared exonic structure to discover sub-genic features, genes, and gene families. The result of this process will be a hierarchical model of the transcriptome. New methods for efficient quantification and differential expression analysis of these hierarchical models will be developed, including a methodologically-sound approach for propagating measures of quantification uncertainty into downstream analysis. Finally, these tools will be validated via a large-scale reanalysis of existing de novo transcriptome data targeted at elucidating the identity of genetic regulatory elements involved in C4 photosynthesis. This reanalysis will harness the increased accuracy and efficiency of the methods developed to analyze data from all previous non-model C4 RNA-seq experiments in a single, multi-scale model. The output of the model will be a system for quantitatively prioritizing candidate regulatory elements for detailed molecular investigation. This project also includes broader impact goals that will provide research opportunities to undergraduate students, and which will create an actively-maintained and inclusive online community centered around best practices in de novo transcriptome analysis and experimental design.For further information regarding progress on this project, including the relevant software being developed, please visit https://combine-lab.github.io/txome.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1101/464222
发表时间:
2018-11
期刊:
bioRxiv
影响因子:
--
作者:
[Fatemeh Almodaresi;Prashant Pandey;Michael Ferdman;Robert C. Johnson;Robert Patro]
通讯作者:
Fatemeh Almodaresi;Prashant Pandey;Michael Ferdman;Robert C. Johnson;Robert Patro
CSR: Medium: Approximate Membership Query Data Structures in Computational Biology and Storage
-
批准号:2317838
-
项目类别:Continuing Grant
-
资助金额:$120.0万
-
财政年份:2022
-
负责人:Robert Patro
-
依托单位:
CAREER: A Comprehensive and Lightweight Framework for Transcriptome Analysis
-
批准号:2029424
-
项目类别:Continuing Grant
-
资助金额:$61.67万
-
财政年份:2020
-
负责人:Robert Patro
-
依托单位:
CAREER: A Comprehensive and Lightweight Framework for Transcriptome Analysis
-
批准号:1750472
-
项目类别:Continuing Grant
-
资助金额:$62.5万
-
财政年份:2018
-
负责人:Robert Patro
-
依托单位:
CSR: Medium: Approximate Membership Query Data Structures in Computational Biology and Storage
-
批准号:1763680
-
项目类别:Continuing Grant
-
资助金额:$120.0万
-
财政年份:2018
-
负责人:Robert Patro
-
依托单位:
海外基金