Rank order entropy: why one metric is not enough.

Rank order entropy: why one metric is not enough.
复制标题

排序熵:为什么一个指标是不够的。

DOI:
10.1021/ci200170k
复制
发表时间:
2011
影响因子:
5.6
通讯作者:
Breneman,CurtM
Breneman,CurtM
中科院分区:
化学2区
文献类型:
--
作者:
McLellan,MargaretR;Ryan,MDominic;Breneman,CurtM

文献摘要

相似文献

使用定量结构-活性关系模型来解决药物发现中的问题有一个复杂的历史,通常是由于错误应用QSAR模型,这些模型要么构建得很差,要么在其适用范围之外使用。这种情况促使了各种模型性能指标的开发(R2、PRESS R2、F检验等)。旨在增加用户对QSAR预测有效性的信心。在典型的工作流程场景中,QSAR模型是在分子的训练集上使用诸如试图评估其内部一致性的留一法或多重交叉验证法等度量方法创建和验证的。然而,目前很少有验证方法被设计来直接解决QSAR预测的稳定性,以响应训练集信息内容的变化。由于QSAR的主要目的是快速而准确地估计一组未经测试的分子的感兴趣的性质,因此手头有一种方法来正确设置用户对模型性能的期望是有意义的。事实上,对于最终用户来说,分子预测的数值往往不如根据预测的终点值知道这组分子的排名顺序重要。因此,表征预测排序稳定性的方法是预测QSAR的重要组成部分。遗憾的是,目前可用的许多验证指标中没有一个直接衡量秩次预测的稳定性,这使得开发能够量化模型稳定性的额外指标成为当务之急。为了满足这一需求,这项工作检查了由具有代表性的数据集、描述符集和建模方法创建的QSAR秩序模型的稳定性,然后使用Kendall Tau作为秩序度度量来评估这些模型,并在此基础上评估Shannon熵作为量化秩序性的手段。从训练集中随机去除数据,也称为数据截断分析(DTA),被用作一种系统地减少每个训练集的信息量的手段,同时在面对训练集数据丢失的情况下检查秩序性能和秩序稳定性。DTA ROE模型评估的前提是,模型对训练信息增量丢失的响应将指示其训练集、学习方法和描述符类型的质量和充分性,以覆盖特定的适用领域。这个过程被称为“秩序熵”评估或ROE。通过与信息论的类比,不稳定的秩序模型显示出高水平的隐熵,而在训练集约简过程中几乎保持不变的QSAR秩序模型显示出低的熵。在这项工作中,净资产收益率指标被应用于71个不同大小的数据集,发现与传统指标相比,它揭示了更多关于模型行为的信息。稳定的或持续执行的模型不一定能很好地预测排名顺序。在排名顺序上表现良好的模型在传统指标中并不一定表现良好。结果表明,净资产收益率指标表明,一些常用的定量构效关系模型应该被摒弃。ROE评估有助于辨别数据集、描述符集和建模方法的哪些组合导致优先排序方案中的可用模型,并提供对特定适用领域中特定模型的使用的信心。
The use of Quantitative Structure–Activity Relationship models to address problems in drug discovery has a mixed history, generally resulting from the misapplication of QSAR models that were either poorly constructed or used outside of their domains of applicability. This situation has motivated the development of a variety of model performance metrics (r2, PRESS r2, F-tests, etc.) designed to increase user confidence in the validity of QSAR predictions. In a typical workflow scenario, QSAR models are created and validated on training sets of molecules using metrics such as Leave-One-Out or many-fold cross-validation methods that attempt to assess their internal consistency. However, few current validation methods are designed to directly address the stability of QSAR predictions in response to changes in the information content of the training set. Since the main purpose of QSAR is to quickly and accurately estimate a property of interest for an untested set of molecules, it makes sense to have a means at hand to correctly set user expectations of model performance. In fact, the numerical value of a molecular prediction is often less important to the end user than knowing the rank order of that set of molecules according to their predicted end point values. Consequently, a means for characterizing the stability of predicted rank order is an important component of predictive QSAR. Unfortunately, none of the many validation metrics currently available directly measure the stability of rank order prediction, making the development of an additional metric that can quantify model stability a high priority. To address this need, this work examines the stabilities of QSAR rank order models created from representative data sets, descriptor sets, and modeling methods that were then assessed using Kendall Tau as a rank order metric, upon which the Shannon entropy was evaluated as a means of quantifying rank-order stability. Random removal of data from the training set, also known as Data Truncation Analysis (DTA), was used as a means for systematically reducing the information content of each training set while examining both rank order performance and rank order stability in the face of training set data loss. The premise for DTA ROE model evaluation is that the response of a model to incremental loss of training information will be indicative of the quality and sufficiency of its training set, learning method, and descriptor types to cover a particular domain of applicability. This process is termed a “rank order entropy” evaluation or ROE. By analogy with information theory, an unstable rank order model displays a high level of implicit entropy, while a QSAR rank order model which remains nearly unchanged during training set reductions would show low entropy. In this work, the ROE metric was applied to 71 data sets of different sizes and was found to reveal more information about the behavior of the models than traditional metrics alone. Stable, or consistently performing models, did not necessarily predict rank order well. Models that performed well in rank order did not necessarily perform well in traditional metrics. In the end, it was shown that ROE metrics suggested that some QSAR models that are typically used should be discarded. ROE evaluation helps to discern which combinations of data set, descriptor set, and modeling methods lead to usable models in prioritization schemes and provides confidence in the use of a particular model within a specific domain of applicability.