The predictive power of data-processing statistics.

The predictive power of data-processing statistics.
复制标题

数据处理统计的预测能力。

DOI:
10.1107/s2052252520000895
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Vollmar M
Vollmar M
中科院分区:
材料科学2区
文献类型:
--
作者:
Vollmar M

文献摘要

相似文献

本研究描述了一种方法来估计成功的可能性,在确定一个大分子结构的X射线晶体学和实验单波长异常色散(SAD)或多波长异常色散(MAD)的基础上,初始数据处理统计和样品晶体特性的定相。这种预测工具可以快速评估数据的有用性,并指导收集最佳数据集。现代大分子晶体学光束线的数据速率的增加,以及用户对实时反馈的需求,导致了对计算资源的压力和对更智能数据处理的需求。统计和机器学习方法已被应用于构建一个分类器,该分类器显示95%的准确率,用于训练和测试从440个解决的结构编译的数据集。将该分类器应用于新数据可实现79%的准确率。这些分数已经为有效使用计算资源提供了明确的指导,并为个性化数据收集助理提供了一个起点。
This study describes a method to estimate the likelihood of success in determining a macromolecular structure by X-ray crystallography and experimental single-wavelength anomalous dispersion (SAD) or multiple-wavelength anomalous dispersion (MAD) phasing based on initial data-processing statistics and sample crystal properties. Such a predictive tool can rapidly assess the usefulness of data and guide the collection of an optimal data set. The increase in data rates from modern macromolecular crystallography beamlines, together with a demand from users for real-time feedback, has led to pressure on computational resources and a need for smarter data handling. Statistical and machine-learning methods have been applied to construct a classifier that displays 95% accuracy for training and testing data sets compiled from 440 solved structures. Applying this classifier to new data achieved 79% accuracy. These scores already provide clear guidance as to the effective use of computing resources and offer a starting point for a personalized data-collection assistant.