Trust, but verify: on the importance of chemical structure curation in cheminformatics and QSAR modeling research.

Trust, but verify: on the importance of chemical structure curation in cheminformatics and QSAR modeling research.
复制标题

DOI:
10.1021/ci100176x
复制
发表时间:
2010-07-26
影响因子:
5.6
通讯作者:
Tropsha A
Tropsha A
中科院分区:
化学2区
文献类型:
--
作者:
Fourches D;Muratov E;Tropsha A

文献摘要

参考文献

被引文献

相似文献

分子建模师和化学信息学家通常分析其他科学家生成的实验数据。因此,在数据准确性方面,化学信息学家总是受到数据提供者的摆布,他们可能会无意中发布(部分)错误的数据。因此,数据集管理对于任何化学信息学分析都是至关重要的,例如相似性搜索,聚类,QSAR建模,虚拟筛选等,特别是在近年来公共领域的化学数据集的可用性急剧增加的今天。尽管这一初步步骤在任何数据集的计算分析中具有明显的重要性,但似乎没有普遍接受的化学数据管理指南或程序集。本文的主要目的是强调需要一个标准化的化学数据策展策略,应遵循在任何分子建模调查的开始。在这里,我们讨论了几个简单但重要的步骤,清理数据库中的化学记录,包括删除一小部分的数据,不能适当地处理传统的化学信息学技术。这些步骤包括去除无机和有机金属化合物、抗衡离子、盐和混合物;结构验证;环芳构化;特定化学型的标准化;互变异构形式的管理;以及重复的删除。为了强调数据策展作为数据分析中的一个强制性步骤的重要性,我们讨论了几个案例研究,其中原始“原始”数据库的化学策展使成功的建模研究(特别是QSAR分析)或导致模型预测准确性的显着提高。我们还表明,在某些情况下,严格开发的QSAR模型,甚至可以用来纠正错误的生物数据与化合物。我们相信,本文中概述的化学记录管理的良好实践对所有在分子建模、化学信息学和QSAR研究领域工作的科学家都有价值。
Molecular modelers and cheminformaticians typically analyze experimental data generated by other scientists. Consequently, when it comes to data accuracy, cheminformaticians are always at the mercy of data providers who may inadvertently publish (partially) erroneous data. Thus, dataset curation is crucial for any cheminformatics analysis such as similarity searching, clustering, QSAR modeling, virtual screening, etc., especially nowadays when the availability of chemical datasets in public domain has skyrocketed in recent years. Despite the obvious importance of this preliminary step in the computational analysis of any dataset, there appears to be no commonly accepted guidance or set of procedures for chemical data curation. The main objective of this paper is to emphasize the need for a standardized chemical data curation strategy that should be followed at the onset of any molecular modeling investigation. Herein, we discuss several simple but important steps for cleaning chemical records in a database including the removal of a fraction of the data that cannot be appropriately handled by conventional cheminformatics techniques. Such steps include the removal of inorganic and organometallic compounds, counterions, salts and mixtures; structure validation; ring aromatization; normalization of specific chemotypes; curation of tautomeric forms; and the deletion of duplicates. To emphasize the importance of data curation as a mandatory step in data analysis, we discuss several case studies where chemical curation of the original “raw” database enabled the successful modeling study (specifically, QSAR analysis) or resulted in a significant improvement of model's prediction accuracy. We also demonstrate that in some cases rigorously developed QSAR models could be even used to correct erroneous biological data associated with chemical compounds. We believe that good practices for curation of chemical records outlined in this paper will be of value to all scientists working in the fields of molecular modeling, cheminformatics, and QSAR studies.
DOI: 10.1007/s10822-007-9162-7
发表时间: 2008-02-01
影响因子: 3.5
作者:
Doweyko, Arthur M.
通讯作者: Doweyko, Arthur M.
DOI: 10.1021/ci990062c
发表时间: 1999-11-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
Brecher, J
通讯作者: Brecher, J
DOI: 10.1021/ci6003515
发表时间: 2007-03-01
影响因子: 5.6
作者:
Hou, Tingjun;Wang, Junmei;Xu, Xiaojie
通讯作者: Xu, Xiaojie
DOI: 10.1021/jm970732a
发表时间: 1998-07-02
影响因子: 7.3
作者:
Kubinyi, H;Hamprecht, FA;Mietzner, T
通讯作者: Mietzner, T
DOI: 10.1021/ci900161g
发表时间: 2009-09-01
影响因子: 5.6
作者:
Hansen, Katja;Mika, Sebastian;Mueller, Klaus-Robert
通讯作者: Mueller, Klaus-Robert