A novel applicability domain technique for mapping predictive reliability across the chemical space of a QSAR: reliability-density neighbourhood

A novel applicability domain technique for mapping predictive reliability across the chemical space of a QSAR: reliability-density neighbourhood
复制标题

DOI:
10.1186/s13321-016-0182-y
复制
发表时间:
2016-12-03
影响因子:
8.6
通讯作者:
Ghafourian T
Ghafourian T
中科院分区:
化学2区
文献类型:
--
作者:
Aniceto N;Freitas AA;Bender A;Ghafourian T

文献摘要

被引文献

相似文献

.定义预测模型可以安全使用的化学空间区域的能力是确保新预测可靠性的必要条件。这意味着,必须确定整个化学空间的可靠性,试图定位“安全”和“不安全”的预测区域。因此,我们设计了一个适用性域技术,解决了数据的本地,而不是处理它作为一个整体的可靠性密度邻域(RDN)。该方法的主要新奇之处在于,它根据训练集中其邻域的密度以及其个体偏差和精度来表征每个单个训练实例。通过扫描化学空间(通过迭代地增加适用性域面积),观察到新的测试化合物以与其预测性能强烈相关的方式连续地包括在适用性域区域中。这允许在训练集空间中的不同位置上映射局部可靠性,从而允许识别模型具有低可靠性的区域。该方法还显示了两个外部集之间的匹配配置文件,这表明它对新数据的执行是稳健的。该技术的另一个新颖方面是它与特定的特征选择算法配对。因此,研究了使用的特征集的影响,ReliefF选择的前20个特征产生了最佳结果,而不是像通常那样使用模型的特征或整个特征集。作为第三个新颖的方面,在这项工作中,我们提出了一个新的评分函数,以帮助评估适用性域配置文件的质量(即,准确度与所讨论的适用性域测量的曲线)。总的来说,RDN是一种很有前途的方法,可以根据预测性能正确地对新实例进行排序。因此,这种技术可以被最终用户接受,作为QSAR模型在新数据中性能的概念证明,从而促进用户对QSAR输出的信任。 本文的在线版本(doi:10.1186/s13321-016-0182-y)包含补充材料,可供授权用户使用。
. The ability to define the regions of chemical space where a predictive model can be safely used is a necessary condition to assure the reliability of new predictions. This implies that reliability must be determined across chemical space in the attempt to localize “safe” and “unsafe” regions for prediction. As a result we devised an applicability domain technique that addresses the data locally instead of handling it as a whole—the reliability-density neighbourhood (RDN). The main novelty aspect of this method is that it characterizes each single training instance according to the density of its neighbourhood in the training set, as well as its individual bias and precision. By scanning through the chemical space (by iteratively increasing the applicability domain area), it was observed that new test compounds are successively included into the applicability domain region in such a manner that strongly correlates to their predictive performance. This allows the mapping of local reliability across different locations in the training set space, and thus allows identifying regions where the model has low reliability. This method also showed matching profiles between two external sets, which is an indication that it performs robustly with new data. Another novel aspect in this technique is that it is paired with a specific feature selection algorithm. As a result, the impact of the feature set used was studied from which the top 20 features selected by ReliefF yielded the best results, as opposed to using the model’s features or the entire feature set as commonly done. As the third novel aspect, in this work we propose a new scoring function to help evaluate the quality of an applicability domain profile (i.e., the curve of accuracy vs the applicability domain measure in question). Overall, the RDN showed to be a promising method that can correctly sort new instances according to predictive performance. As a result, this technique can be received by an end-user as proof of concept for the performance of a QSAR model in new data, thus promoting the user’s trust on the QSAR output. The online version of this article (doi:10.1186/s13321-016-0182-y) contains supplementary material, which is available to authorized users.