Computational methods to interpret genomic variation and integrate functional genomics data in genetic analysis of human diseases
Computational methods to interpret genomic variation and integrate functional genomics data in genetic analysis of human diseases
批准号:
10623773
负责人:
Yufeng Shen
金额:
$40.69万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2028-05-31
中文摘要
摘要
我实验室的总体研究方向是开发新的计算方法,以实现遗传学的新发现。
人类疾病的研究。我们设计了基于生物直觉的机器学习模型来提取
来自大型基因组测序和功能基因组学数据集的知识。
最近大规模的人类疾病基因组和外显子组测序研究已经成功地确定了新的风险
基因和提高临床基因检测的诊断率,特别是在罕见疾病,发育障碍,
癌然而,许多重要的遗传问题仍然没有解决,不能通过遗传物质的积累来解决。
数据本身。大多数人类疾病的风险基因仍然是未知的。特别是,罕见变异的作用已经被证明是
被研究的一个主要的瓶颈是缺乏高度准确和自动化的工具来解释遗传变异。罕见
错义变体占大多数具有潜在功能影响的蛋白质编码变体;然而,
不会导致疾病。无法准确预测其功能影响是识别风险的关键障碍
基因研究中的基因,并在临床实践中消除不确定意义的变异。我们看到一个
这是一个独特的机会,在未来五年内显着改善计算方法,由于以下汇合
因素:积累大量的人口基因组序列数据,现代深度学习方法来建模基因组,
蛋白质序列和结构,跨细胞类型和发育阶段的人类功能基因组学数据,
分析遗传变异的分子效应的方法。我们将重点关注三个方面。首先是计算预测
错义变体的功能影响。我们使用深度神经网络来学习蛋白质的有效表示
预测模型中的序列和结构,并使用概率图形模型来联合估计分子水平上的影响
人口水平。第二个是功能基因组学和遗传学数据的计算集成。我们会融合
机器学习与统计遗传学,以开发疾病遗传数据建模的方法,
正常个体的表达和调节谱。该方法将提高新的危险基因的统计功效
发现和产生疾病病因学的生物学见解。第三,我们将继续开发新的生物信息学工具
为了改善大规模的拷贝数变异和嵌合突变的检测和自动化确认,
基因组学数据。
最后,我们与医学遗传学专家的合作将提供积极的反馈循环,以改进方法
并产生新的生物学见解和临床效用。我们的研究将产生新的方法来分析基因组学数据,
最终,这些方法将使疾病遗传学研究有新的发现,并提高临床治疗的产量。
基因诊断
英文摘要
Abstract
The overall research direction in my lab is to develop new computational methods to enable new discovery in genetic
studies of human diseases. We design methods using machine learning models based on biological intuitions to extract
knowledge from large genome sequencing and functional genomics data sets.
Recent large-scale genome and exome sequencing studies of human diseases have successfully identified novel risk
genes and improved diagnostic yields in clinical genetic testing, especially in rare diseases, developmental disorders, and
cancer. However, many significant genetic questions remain unsolved and cannot be solved by the accumulation of genetic
data alone. Most of the risk genes of human diseases are still unknown. In particular, the role of rare variants has been
under-studied. One major bottleneck is the lack of highly accurate and automated tools to interpret genetic variation. Rare
missense variants account for most of protein-coding variants with potential functional impact; however, most of them
do not contribute to diseases. The inability to accurately predict their functional impact is a critical hurdle to identify risk
genes in genetic research studies and to disambiguate variants of uncertain significance in clinical practice. We see a
unique opportunity to dramatically improve computational methods in next five years, due to the following confluent
factors: accumulating large population genome sequence data, modern deep learning methods to model genomic and
protein sequence and structure, human functional genomics data across cell types and developmental stages, and scalable
methods to profile molecular effect of genetic variants. We will focus on three areas. The first is computational prediction
of functional impact of missense variants. We use deep neural networks to learn effective representation of protein
sequence and structure in prediction models and use probabilistic graphical models to jointly estimate effects at molecular
and population levels. The second is computational integration of functional genomics and genetics data. We will fuse
machine learning with statistical genetics to develop methods that model disease genetic data together with single cell
expression and regulatory profiles of normal individuals. The methods will improve both statistical power of new risk gene
discovery and generate biological insights of disease etiology. Third, we will continue to develop new bioinformatics tools
to improve detecting and automated confirmation of copy number variants and mosaic mutations from large-scale
genomics data.
Finally, our collaboration with experts in medical genetics will provide positive feedback loops to improve the methods
and generate new biological insights and clinical utility. Our research will produce new methods to analyze genomics data,
and ultimately these methods will enable new discoveries in disease genetic studies and improve the yield of clinical
genetic diagnostics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational analysis of whole genome sequence data for discovering novel risk genes of structural birth defects
-
批准号:10354418
-
项目类别:
-
资助金额:$15.96万
-
财政年份:2022
-
负责人:Yufeng Shen
-
依托单位:
Computational analysis of whole genome sequence data for discovering novel risk genes of structural birth defects
-
批准号:10673600
-
项目类别:
-
资助金额:$15.96万
-
财政年份:2022
-
负责人:Yufeng Shen
-
依托单位:
Integrate cancer genomics data in genetic studies and diagnosis of developmental disorders
-
批准号:10166608
-
项目类别:
-
资助金额:$33.31万
-
财政年份:2017
-
负责人:Yufeng Shen
-
依托单位:
Integrated Genomics Core
-
批准号:10458159
-
项目类别:
-
资助金额:$25.63万
-
财政年份:2017
-
负责人:Yufeng Shen
-
依托单位:
Integrated Genomics Core
-
批准号:10647825
-
项目类别:
-
资助金额:$24.15万
-
财政年份:2017
-
负责人:Yufeng Shen
-
依托单位:
Integrate cancer genomics data in genetic studies and diagnosis of developmental disorders
-
批准号:9311160
-
项目类别:
-
资助金额:$33.27万
-
财政年份:2017
-
负责人:Yufeng Shen
-
依托单位:
Bioinformatics & Data Management
-
批准号:10176371
-
项目类别:
-
资助金额:$32.68万
-
财政年份:2013
-
负责人:Yufeng Shen
-
依托单位:
Bioinformatics & Data Management
-
批准号:10426136
-
项目类别:
-
资助金额:$32.68万
-
财政年份:2013
-
负责人:Yufeng Shen
-
依托单位:
Bioinformatics
-
批准号:8576997
-
项目类别:
-
资助金额:$33.89万
-
财政年份:2013
-
负责人:Yufeng Shen
-
依托单位:
Bioinformatics
-
批准号:8703320
-
项目类别:
-
资助金额:$42.81万
-
财政年份:--
-
负责人:Yufeng Shen
-
依托单位:
Bioinformatics
-
批准号:9284396
-
项目类别:
-
资助金额:$31.72万
-
财政年份:--
-
负责人:Yufeng Shen
-
依托单位:
海外基金