课题基金 / 基金详情

项目摘要

项目成果

DONG XU的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 真核细胞具有多样的细胞成分,包括亚细胞器和亚细胞器 隔间蛋白质对这些细胞组分的准确靶向对于建立和 维持细胞组织和功能。蛋白质的错误定位通常与代谢有关。 失调和疾病。然而,绝大多数蛋白质缺乏亚细胞/亚细胞器定位 注释。与实验方法相比,蛋白质定位的计算预测提供了一种新的方法, 为蛋白质组注释和实验设计提供了一种高效的方法。当前的预测工具 蛋白质定位还有很大的改进空间。此外,没有任何工具可以预测定位在子- 细胞器分辨率或内部定位信号。深度学习作为机器学习领域的前沿技术, 学习,为这个经典的生物信息学问题提供了一个新的机会。最近的高- 吞吐量定位数据也可以很好地训练深度学习。PI的实验室已经证明了一些成功, 一种特殊情况,即,使用深度学习预测植物的线粒体定位。 在这个项目中,PI建议开发新的方法和一个独立的工具包, 在亚细胞和亚细胞器水平的蛋白定位预测,以及表征 定位基序(包括新的内部基序)。一般的方法是设计一个半监督的深度- 利用具有已知定位的注释蛋白质序列和未注释蛋白质的学习方法 序列作为训练数据。通过实现无监督的深度学习方法, 将实现蛋白质序列的表示,表征蛋白质的局部和全局特征 序列的通过可视化和表征深度学习模型, 模式将被预测为推定的靶向肽,并与已知的定位信号进行比较。我们将 我还使用要开发的方法和要在所有蛋白质序列上训练的无监督模型, 其他基于序列的预测问题的一般框架,预测蛋白质的标签和关键 有助于标记的残基。我们将使平台高度可定制,并将其应用于三个 应用,包括泛素化蛋白预测,酶EC数预测,和蛋白质 科/亚科分类。对基于蛋白质序列的分析和预测的创新贡献 包括:(1)使用原始氨基酸序列作为训练输入,而不进行特征工程;(2)利用巨大的 在无监督深度学习中描述一般蛋白质特征的未注释数据量 (3)通过解码训练的深度表示来识别潜在的靶向信号(特别是内部基序); 学习模型,增加了复杂的注意力机制; 4)检测多个细胞器的目标 和亚细胞器定位的一种新的分层多标签架构;和(5)结合的功能, 不同的数据源通过乘法融合CNN模型。
英文摘要
Project Summary Eukaryotic cells have diverse cellular components, including subcellular organelles and sub-organelle compartments. The accurate targeting of proteins to these cellular components is crucial in establishing and maintaining cellular organizations and functions. Mis-localization of proteins is often associated with metabolic disorders and diseases. However, the vast majority of proteins lack subcellular/sub-organelle localization annotation. Compared with experimental methods, computational prediction of protein localization provides an efficient and effective way for proteome annotation and experimental design. The current prediction tools for protein localization have significant room for improvement. In addition, no tool can predict localization at the sub- organelle resolution or internal localization signals. Deep learning, as the cutting-edge technology in machine learning, presents a new opportunity for this classical bioinformatics problem. The availability of recent high- throughput localization data can also train deep learning well. The PI’s lab has demonstrated some success on a special case, i.e., predicting mitochondrial localizations for plants using deep learning. In this project, the PI proposes to develop new methods and a standalone toolkit for accurate and scalable protein localization prediction at the subcellular and sub-organelle levels, as well as for characterization of localization motifs (including novel internal motifs). The general approach is to design a semi-supervised deep- learning method that utilizes both annotated protein sequences with known localization and unannotated protein sequences as training data. Through the realization of an unsupervised deep-learning approach, a general representation of protein sequences will be implemented, characterizing both local and global features of protein sequences. By visualizing and characterizing the deep-learning models, novel, interpretable protein sequence patterns will be predicted as putative targeting peptides and compared with known localization signals. We will also use the methods to be developed and the unsupervised models to be trained on all protein sequences as a general framework for other sequence-based prediction problems that predict the label of a protein and the key residues contributing to the label. We will make the platform highly customizable and apply it to three applications, including ubiquitination protein prediction, enzyme EC number prediction, and protein family/subfamily classification. The innovative contributions to protein sequence-based analyses and predictions include: (1) using raw amino acid sequences as training inputs without feature engineering; (2) utilizing the huge amount of unannotated data in an unsupervised deep learning to characterize a general protein feature representation; (3) identifying potential targeting signals (especially internal motifs) by decoding the trained deep- learning models, augmented with sophisticated attention mechanisms; 4) detecting multiple-organelle targeting and sub-organelle localizations by a novel hierarchical multi-label architecture; and (5) combining features from different data sources by a multiplicative fused CNN model.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1016/j.csbj.2021.10.023
发表时间: 2021
期刊: Computational and structural biotechnology journal
影响因子: 6
作者: [Jiang Y, Wang D, Wang W, Xu D]
通讯作者: Xu D
DOI: 10.1016/j.csbj.2021.08.027
发表时间: 2021
期刊: Computational and structural biotechnology journal
影响因子: 6
作者: [Jiang Y, Wang D, Yao Y, Eubel H, Künzler P, Møller IM, Xu D]
通讯作者: Xu D
Multi-view self-supervised deep learning for biological sequences and beyond
  • 批准号:
    10623063
  • 项目类别:
  • 资助金额:
    $39.13万
  • 财政年份:
    2018
  • 负责人:
    DONG XU
  • 依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
  • 批准号:
    10395451
  • 项目类别:
  • 资助金额:
    $45.64万
  • 财政年份:
    2018
  • 负责人:
    DONG XU
  • 依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
Interpretable and extendable deep learning model for biological sequence analysis and prediction
  • 批准号:
    10409152
  • 项目类别:
  • 资助金额:
    $23.48万
  • 财政年份:
    2018
  • 负责人:
    DONG XU
  • 依托单位:
海外基金