Development of reference-free algorithms for low coverage RNA-Seq characterization of cell states
Development of reference-free algorithms for low coverage RNA-Seq characterization of cell states
批准号:
RGPIN-2022-04260
负责人:
Lemieux, Sébastien
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Since the early days of microarrays, transcriptomics had a central role in modern molecular biology by providing a rich and dynamic view of active molecular processes within cells. RNA-Seq is nowadays a routine methodology accessible to most academic laboratories and provides an infinitely richer view of the transcriptome when compared to microarrays. Analyses pipelines for RNA-Seq unfortunately discard precious observations present in the original data such as unannotated transcripts, gene splicing events or gene rearrangements. Using the resulting biased expression profiles to train artificial intelligence algorithms such as deep neural networks prevents them from reaching their full potential. We propose a most drastic reorganization of RNA-Seq data analysis around the use of k-mer count tables (KCTs) as purely data-driven summaries. As k-mers are short, fixed-length sequences associated with an expression value, they are the perfect input to deep neural networks. This representation retains a much wider range of features present in the transcriptome and could readily be applied to organisms in which the genome is unannotated. The main challenge brought by this representation is its size, easily reaching tens of millions of k-mers for a single experiment. Work on this program will proceed on three fronts. First, we will fully master the various intricacies of quantitatively representing transcriptomes using k-mer count tables. Of particular interest will be the determination of the optimal k-mer length and the identification of an appropriate normalization to account for the varying sequencing depth. This first step will be done to accommodate datasets of several thousands of samples. Second, we will extend a neural network-based algorithm developed in our laboratory, the factorized embedding to allow k-mer-based representations to be used as input. This algorithm has the property to return a short numerical vector that summarizes, as well as possible, all quantitative observations presented at the training stage. Third, we will take advantage of these numerical summaries as ideal input to a second layer of deep neural network that we will design and train to make useful predictions on represented samples. Typically these predictions will be phenotypes of interest of varying levels of complexity. Both the factorized embeddings and these neural networks will be trained on large and well-established, public RNA-Seq datasets. This program will open a whole new perspective on the analysis of RNA-Seq data, yielding high-performance, open source software tools for the transcriptomic user's community. These advances should be particularly transformative in fields studying less genetically characterized organisms. As we expect to confirm the sufficiency of very low depth RNA-Seq, this program will open applications where RNA-Seq has been deemed unaffordable.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
听力正常与听觉障碍人群脑中自我参照系统与环境参照系统之间的交互作用
-
批准号:31070994
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2010
-
负责人:陈骐
-
依托单位: