Next Generation Machine Learning for the Accurate Detection of DNA Variations from High-Throughput Sequencing Data
Next Generation Machine Learning for the Accurate Detection of DNA Variations from High-Throughput Sequencing Data
批准号:
RGPIN-2019-04896
负责人:
Bashashati, Ali
金额:
$2.04万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31
中文摘要
脱氧核糖核酸(DNA)是人类和几乎所有其他生物体的遗传物质。“下一代测序”(NGS)技术的出现?为大规模研究DNA创造了前所未有的机会。虽然NGS功能非常强大,但它可以产生大量数据(每个癌症患者约250 GB)。这些大型数据集包含扭曲和模糊真实DNA变化(称为变异)的错误。需要在用于检测来自不同组织类型和技术的DNA变异的算法的开发方面实现重大飞跃,因为:(a)迄今为止开发的手工制作和参数化算法仍然产生数千个错误并错过真正的DNA变异;(B)虽然最近的单细胞DNA测序(SCS)技术现在能够表征每个细胞的DNA,用于检测SCS数据中的DNA变异的算法远远落后于整个数据生成;以及(c)用于存档组织材料的常见方法是福尔马林固定,其引入假DNA变异并对识别存储的组织中的真实DNA变异提出挑战。
我们的长期目标是开发计算方法来识别导致生物异常的DNA变异。在探索计划的这个周期内,我们将开发新的机器学习框架(基于深度学习和集成学习),从NGS技术测序的各种组织来源中检测DNA变异(无论它们是否导致异常)。模型数据集将被用来验证这些新的算法,并丰富他们的方法。
这一基础研究计划的成功执行将导致一类新的算法和软件,其将(a)使致力于DNA测序的大量资源的效益最大化,在后续实验中节省数百万美元,以验证从噪声NGS数据中识别的DNA变异;(B)使SCS能够更有效地表征细胞,这将对许多不同的领域具有广泛的影响,包括微生物学、神经生物学、发育,免疫学和癌症;以及(c)为有效评估福尔马林固定组织中的DNA序列打开了大门,提供了大量数据来筛选和全面评估疾病标志物。
据我们所知,该计划在加拿大是独特和新颖的,将培养2名博士和1名硕士学生,以及7名在学术界和工业界具有高度要求技能的本科生。一个详细的培训计划,使个人能够充分发挥其潜力,是在研究目标的发展,以确保高质量,互动,繁荣,和公平的研究环境HQP根据UBC的公平和多样性战略计划和政策#2集成。
英文摘要
Deoxyribonucleic acid (DNA) is the hereditary material in humans and almost all other organisms. The emergence of “next generation sequencing” (NGS) technology?has created unprecedented opportunities to study DNA in large scales. While extremely powerful, NGS produces massive quantities of data (~250GB/cancer patient). These large datasets contain errors that distort and obscure the true DNA changes (referred to as variations). Significant leap in the development of algorithms for the detection of DNA variations from different tissue types and technologies is needed, because: (a) the hand-crafted and parameterized algorithms developed so far still produce thousands of errors and miss true DNA variations; (b) while the recent single cell DNA sequencing (SCS) technologies are now able to characterize the DNA of each cell, algorithms for detection of DNA variations in SCS data lag far behind the data generation throughout; and (c) a common approach to archive tissue material is formalin fixation, which introduces false DNA variations and poses a challenge to the identification of true DNA variations in the stored tissue.
Our long-term objective is to develop computational methods to identify DNA variations that cause biological abnormalities. Within this cycle of the Discovery program, we will develop novel machine learning frameworks (based on deep learning and ensemble learning) that detect DNA variations (regardless of whether they cause abnormalities) from various sources of tissue sequenced by NGS technologies. Model datasets will be used to validate these novel algorithms and enrich their approaches.
Successful execution of this basic research program will lead to a novel class of algorithms and software that will (a) maximize the benefit of significant resources committed to DNA sequencing saving millions of dollars in follow up experiments to validate the DNA variations identified from noisy NGS data; (b) enable SCS to more effectively characterize cells which will have a broad impact on many diverse fields including microbiology, neurobiology, development, immunology and cancer; and (c) open the door for the effective assessment of DNA sequence in formalin-fixed tissue, providing an explosion in data to screen and comprehensively evaluate disease markers.
To the best of our knowledge, the proposed program is unique and novel in Canada and will train 2 PhD and 1 MSc student, as well as 7 undergraduate students with highly demanded skills in academia and industry. A detailed training plan that allows individuals to reach their full potential is integrated within the development of research objectives to ensure a high quality, interactive, flourishing, and an equitable research environment for HQP as per UBC Equity and Diversity Strategic Plan and Policy #2.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Next Generation Machine Learning for the Accurate Detection of DNA Variations from High-Throughput Sequencing Data
-
批准号:RGPIN-2019-04896
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2022
-
负责人:Bashashati, Ali
-
依托单位:
Next Generation Machine Learning for the Accurate Detection of DNA Variations from High-Throughput Sequencing Data
-
批准号:RGPIN-2019-04896
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2021
-
负责人:Bashashati, Ali
-
依托单位:
Next Generation Machine Learning for the Accurate Detection of DNA Variations from High-Throughput Sequencing Data
-
批准号:RGPIN-2019-04896
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2019
-
负责人:Bashashati, Ali
-
依托单位:
Next Generation Machine Learning for the Accurate Detection of DNA Variations from High-Throughput Sequencing Data
-
批准号:DGECR-2019-00028
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2019
-
负责人:Bashashati, Ali
-
依托单位:
Computational methods for the analysis of flow cytometry data
-
批准号:343277-2007
-
项目类别:Postdoctoral Fellowships
-
资助金额:$1.46万
-
财政年份:2009
-
负责人:Bashashati, Ali
-
依托单位:
Computational methods for the analysis of flow cytometry data
-
批准号:343277-2007
-
项目类别:Postdoctoral Fellowships
-
资助金额:$2.91万
-
财政年份:2008
-
负责人:Bashashati, Ali
-
依托单位:
Computational methods for the analysis of flow cytometry data
-
批准号:343277-2007
-
项目类别:Postdoctoral Fellowships
-
资助金额:$1.46万
-
财政年份:2007
-
负责人:Bashashati, Ali
-
依托单位:
国内基金
海外基金
Next Generation Majorana Nanowire Hybrids
-
批准号:--
-
项目类别:--
-
资助金额:20万元
-
批准年份:2020
-
负责人:Panagiotis Kotetes
-
依托单位: