Predicting the performance of automated crystallographic model-building pipelines.

Predicting the performance of automated crystallographic model-building pipelines.
复制标题

预测自动晶体学模型构建管道的性能。

DOI:
10.1107/s2059798321010500
复制
发表时间:
2021-12-01
期刊:
Acta crystallographica. Section D, Structural biology
影响因子:
--
通讯作者:
Cowtan K
Cowtan K
中科院分区:
其他
文献类型:
--
作者:
Alharbi E;Bond P;Calinescu R;Cowtan K

文献摘要

相似文献

使用机器学习模型来预测四种晶体学模型构建管道(ARP/wARP、Buccaneer、Phenix AutoBuild和SHELXE)及其组合的性能。蛋白质是执行基本生物功能的大分子,这些功能取决于它们的三维结构。确定这种结构涉及复杂的实验室和计算工作。对于计算工作,已经开发了多个软件管道来从晶体学数据构建蛋白质结构的模型。这些流水线中的每一个根据作为输入接收的电子密度图的特性而不同地执行。确定用于蛋白质结构的最佳管道是困难的,因为管道性能从一个蛋白质结构到另一个蛋白质结构有很大差异。因此,研究人员经常选择不能从现有数据中产生最佳蛋白质模型的管道。在这里,介绍了一个软件工具,它预测的关键质量措施的蛋白质结构,一系列的管道将产生,如果提供一个给定的晶体学数据集。这些指标是基于纳入和保留的观察结果以及结构完整性的晶体学拟合质量指标。使用超过2500个数据集进行的广泛实验表明,该工具可以对实验定相数据集(分辨率在1.2和4.0 μ m之间)和分子置换数据集(分辨率在1.0和3.5 μ m之间)进行准确的预测。  因此,该工具可以向用户提供关于应该运行的管线的建议,以便最有效地进行到可沉积模型。
A machine-learning model was used to predict the performance of four crystallographic model-building pipelines (ARP/wARP, Buccaneer, Phenix AutoBuild and SHELXE) and their combinations. Proteins are macromolecules that perform essential biological functions which depend on their three-dimensional structure. Determining this structure involves complex laboratory and computational work. For the computational work, multiple software pipelines have been developed to build models of the protein structure from crystallographic data. Each of these pipelines performs differently depending on the characteristics of the electron-density map received as input. Identifying the best pipeline to use for a protein structure is difficult, as the pipeline performance differs significantly from one protein structure to another. As such, researchers often select pipelines that do not produce the best possible protein models from the available data. Here, a software tool is introduced which predicts key quality measures of the protein structures that a range of pipelines would generate if supplied with a given crystallographic data set. These measures are crystallographic quality-of-fit indicators based on included and withheld observations, and structure completeness. Extensive experiments carried out using over 2500 data sets show that the tool yields accurate predictions for both experimental phasing data sets (at resolutions between 1.2 and 4.0 Å) and molecular-replacement data sets (at resolutions between 1.0 and 3.5 Å). The tool can therefore provide a recommendation to the user concerning the pipelines that should be run in order to proceed most efficiently to a depositable model.