Development of reference-free algorithms for low coverage RNA-Seq characterization of cell states
Development of reference-free algorithms for low coverage RNA-Seq characterization of cell states
批准号:
RGPIN-2022-04260
负责人:
Lemieux, Sébastien
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
自微阵列早期以来,转录组学通过提供细胞内活跃分子过程的丰富和动态视图,在现代分子生物学中起着核心作用。RNA-Seq是目前大多数学术实验室可访问的常规方法,与微阵列相比,它提供了无限丰富的转录组视图。不幸的是,RNA-Seq的分析管道丢弃了原始数据中存在的宝贵观察结果,如未注释的转录本、基因剪接事件或基因重排。使用由此产生的有偏见的表达谱来训练人工智能算法,如深度神经网络,会阻止它们充分发挥潜力。我们建议对RNA-Seq数据分析进行最激烈的重组,围绕使用k-mer计数表(kct)作为纯粹的数据驱动摘要。由于k-mers是与表达值相关的短且固定长度的序列,因此它们是深度神经网络的完美输入。这种表达保留了转录组中存在的更广泛的特征,并且可以很容易地应用于基因组未注释的生物体。这种表示带来的主要挑战是它的大小,单个实验很容易达到数千万k-mers。该计划的工作将在三个方面进行。首先,我们将充分掌握使用k-mer计数表定量表示转录组的各种复杂性。特别感兴趣的将是确定最佳k-mer长度和确定适当的归一化,以考虑不同的测序深度。这第一步将用于容纳数千个样本的数据集。其次,我们将扩展我们实验室开发的基于神经网络的算法,即因式嵌入,以允许使用基于k-mer的表示作为输入。该算法具有返回一个简短的数值向量的特性,该向量尽可能地总结了在训练阶段呈现的所有定量观察结果。第三,我们将利用这些数值摘要作为第二层深度神经网络的理想输入,我们将设计和训练该网络,以便对所代表的样本进行有用的预测。通常,这些预测将是不同复杂程度的感兴趣表型。因式嵌入和这些神经网络都将在大型且完善的公共RNA-Seq数据集上进行训练。该项目将为RNA-Seq数据分析开辟一个全新的视角,为转录组学用户社区提供高性能的开源软件工具。在研究遗传特征较少的生物体的领域,这些进展尤其具有变革性。由于我们希望确认非常低深度RNA-Seq的充分性,因此该计划将开放RNA-Seq被认为无法负担的应用程序。
英文摘要
Since the early days of microarrays, transcriptomics had a central role in modern molecular biology by providing a rich and dynamic view of active molecular processes within cells. RNA-Seq is nowadays a routine methodology accessible to most academic laboratories and provides an infinitely richer view of the transcriptome when compared to microarrays. Analyses pipelines for RNA-Seq unfortunately discard precious observations present in the original data such as unannotated transcripts, gene splicing events or gene rearrangements. Using the resulting biased expression profiles to train artificial intelligence algorithms such as deep neural networks prevents them from reaching their full potential. We propose a most drastic reorganization of RNA-Seq data analysis around the use of k-mer count tables (KCTs) as purely data-driven summaries. As k-mers are short, fixed-length sequences associated with an expression value, they are the perfect input to deep neural networks. This representation retains a much wider range of features present in the transcriptome and could readily be applied to organisms in which the genome is unannotated. The main challenge brought by this representation is its size, easily reaching tens of millions of k-mers for a single experiment. Work on this program will proceed on three fronts. First, we will fully master the various intricacies of quantitatively representing transcriptomes using k-mer count tables. Of particular interest will be the determination of the optimal k-mer length and the identification of an appropriate normalization to account for the varying sequencing depth. This first step will be done to accommodate datasets of several thousands of samples. Second, we will extend a neural network-based algorithm developed in our laboratory, the factorized embedding to allow k-mer-based representations to be used as input. This algorithm has the property to return a short numerical vector that summarizes, as well as possible, all quantitative observations presented at the training stage. Third, we will take advantage of these numerical summaries as ideal input to a second layer of deep neural network that we will design and train to make useful predictions on represented samples. Typically these predictions will be phenotypes of interest of varying levels of complexity. Both the factorized embeddings and these neural networks will be trained on large and well-established, public RNA-Seq datasets. This program will open a whole new perspective on the analysis of RNA-Seq data, yielding high-performance, open source software tools for the transcriptomic user's community. These advances should be particularly transformative in fields studying less genetically characterized organisms. As we expect to confirm the sufficiency of very low depth RNA-Seq, this program will open applications where RNA-Seq has been deemed unaffordable.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
听力正常与听觉障碍人群脑中自我参照系统与环境参照系统之间的交互作用
-
批准号:31070994
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2010
-
负责人:陈骐
-
依托单位: