课题基金 / 基金详情

Building machine learning models and neural networks trained on structural information of drug targets to predict antimicrobial resistance

Building machine learning models and neural networks trained on structural information of drug targets to predict antimicrobial resistance
构建机器学习模型和神经网络,并根据药物靶标的结构信息进行训练,以预测抗菌药物耐药性
批准号:
2597363
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
拟议的项目重点是使用相关抗生素靶标的蛋白质结构、化学和进化特征来训练机器学习模型,以预测结核分枝杆菌(Mtb)产生的抗菌素耐药性(AMR)。虽然许多研究人员正在使用遗传特征来预测耐药性,但我们之前已经证明,基于RNA聚合酶的结构和生物物理特征训练的传统机器学习模型可以稳健而准确地预测错义突变对利福平敏感性的影响。然而,这些模型本质上无法预测多个突变可能产生的影响,从而将可用的突变数据限制在可用突变数据的子集上,从而限制了模型的临床适用性。DPhil项目的主要目标就是解决这个问题。这名学生将获得由牛津大学领导的国际神秘项目积累的约7万份临床结核病样本的数据集,该项目正在通过一系列出版物报告其主要发现。CRYTIC收集了15,211个样本,每个样本都进行了全基因组测序,并使用96孔肉汤微量稀释板测量了13种抗生素的敏感性。该数据集的一个局限性是对新化合物缺乏耐药性,如贝达奎兰。其中一个隐秘的合作伙伴最近提供了约1000个广泛耐受的高价值样本,该DPhil的初始目标(Y1)是分析这一额外的数据集,包括对以前开发的机器学习模型进行再培训,以及开发严格的统计分析管道,以实现连续可靠和易于访问的性能评估和基准测试。这将促进该项目的主要目标;开发以结构和化学数据为特征的图形卷积神经网络(GCNN),以预测针对一线和二线抗结核化合物的多重突变所产生的AMR。这种方法背后的假设是,gCNN的拓扑可以准确地捕捉到来自抗性等位基因的所有信息,从而使机器学习模型能够得到有效的训练,并允许首次考虑具有高水平遗传变异性的蛋白质靶标。在时间允许的情况下,一个合乎逻辑的扩展是将从分子动力学轨迹中提取的动态数据纳入特征集,并评估其对模型性能的影响。除了gCNN是一种比传统卷积神经网络(CNN)更直观地表示结构数据的体系结构外,gCNN还保留了原子和化学键的概念,直到网络的最后一层,从而保留了药物靶标的空间信息。这允许询问原子嵌入以提高模型属性,这一概念在临床适用的分子诊断学中特别相关。虽然使用结构数据来预测AMR仍然是一种相对较新的方法,但这种方法的真正新颖性在于,到目前为止,AMR预测领域在很大程度上无法受益于基于足够大的数据集训练的神经网络,特别是关于药物靶标的结构、理化和空间信息的神经网络。此外,等变图神经网络(可以说显示出最有潜力的)是非常新的(2021年),关于结构建模问题,大多被专注于结合亲和力预测的小组采用,而不是AMR预测。该项目将属于EPSRC的下列研究主题:人工智能与数据科学、抗菌素耐药性、生物信息学、生物物理学、临床技术、软件工程
英文摘要
The proposed project focuses on training machine learning models using protein structural, chemical and evolutionary features of relevant antibiotic targets to predict antimicrobial resistance (AMR) conferred by Mycobacterium tuberculosis (Mtb). Whilst many researchers are using genetic features to predict resistance, we have previously demonstrated that traditional machine-learning models trained on structural and biophysical features of RNA polymerase can robustly and accurately predict the effect that a missense mutation confers on rifampicin susceptibility. However, these models are inherently unable to predict the effect multiple mutations can have, thereby constraining usable mutation data to a subset of the available mutation data, and thus limiting the clinical applicability of the models. The primary goal of the DPhil project is to address this. The student will have access to the dataset of around 70,000 clinical TB samples amassed by the international CRyPTIC project which was led by Oxford and is reporting its main findings through a series of publications. CRyPTIC collected 15,211 samples, each of which was whole genome sequenced and the susceptibility of 13 antibiotics measured using a 96-well broth microdilution plate. A limitations of this dataset was the lack of resistance to new compounds, such as bedaquiline. One of the CRyPTIC partners has recently provided c. 1000 high-value samples that are extensively resistant and the initial aims (Y1) of this DPhil are to analyse this additional dataset, including retraining previously developed machine learning models, as well as developing a rigorous statistical analysis pipeline to enable continuous robust and easily accessible performance assessment and benchmarking. This will facilitate the primary objective of the project; developing graph convolutional neural networks (gCNNs) featurised with structural and chemical data to predict AMR conferred by multiple mutations against first- and second-line anti-TB compounds. The hypothesis underlying this approach is that the topology of gCNNs can accurately capture all the information from a resistant allele, thereby allowing machine learning models to be efficiently trained and permitting protein targets with high levels of genetic variability to be considered for the first time. A logical extension, time permitting, would be incorporating dynamic data pulled down from molecular dynamics trajectories into the feature sets and assessing the impact on model performance. Aside from gCNNs being a more intuitive architecture to represent structural data than conventional convolutional neural networks (CNNs), gCNNs also preserve the concepts of the atom and the chemical bond until the final layers of the network, thereby preserving spatial information of the drug target. This allows for interrogation of atom embeddings to boost model attribution, a concept particularly relevant in clinically applicable molecular diagnostics. Although the use of structural data to predict AMR is still a relatively new approach, the real novelty of this methodology is that to date the field of AMR prediction has been largely unable to benefit from neural networks trained on sufficiently large datasets, and particularly neural networks trained on structural, physiochemical, and spatial information of the drug target. Furthermore, equivariant graph neural networks (which arguably show the most potential) are extremely new (2021), and with regard to structural modelling problems, have mostly been adopted by groups focussing on binding affinity prediction, not AMR prediction. This project would fall within the following EPSRC research themes: AI and data science Antimicrobial resistance Biological informatics Biophysics Clinical technologies Software engineering
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
非标准随机调度模型的最优动态策略
  • 批准号:
    71071056
  • 项目类别:
    面上项目
  • 资助金额:
    28.0万元
  • 批准年份:
    2010
  • 负责人:
    吴贤毅
  • 依托单位:
微生物发酵过程的自组织建模与优化控制
  • 批准号:
    60704036
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    21.0万元
  • 批准年份:
    2007
  • 负责人:
    高学金
  • 依托单位: