Distinguishing between Natural Products and Synthetic Molecules by Descriptor Shannon Entropy Analysis and Binary QSAR Calculations

Distinguishing between Natural Products and Synthetic Molecules by Descriptor Shannon Entropy Analysis and Binary QSAR Calculations
复制标题

通过描述符香农熵分析和二元 QSAR 计算区分天然产物和合成分子

DOI:
10.1021/ci0003303
复制
发表时间:
2000
期刊:
Journal of chemical information and computer sciences
影响因子:
--
通讯作者:
J. Bajorath
J. Bajorath
中科院分区:
--
文献类型:
--
作者:
F. Stahura;J. Godden;L. Xue;J. Bajorath

文献摘要

被引文献

相似文献

香农熵分析,正确区分,在二进制QSAR计算,天然存在的分子和合成化合物之间的分子描述符进行了鉴定。香农熵的概念首先用于数字通信理论,直到最近才被应用于描述符分析。二元QSAR方法最初是为了将化合物的结构特征和性质与生物活性的二元制剂(即,活性的或非活性的)并且在此已经适于将分子特征与化学源相关联(即,天然的或合成的)。我们已经确定了一些分子描述符显着不同的香农熵和/或“熵分离”在天然和合成化合物数据库。此类描述符和可变分布的结构键的不同组合应用于由天然和合成分子组成的学习集,并用于推导预测性二元QSAR模型。然后,这些模型被应用于预测不同测试集的化合物的来源,这些测试集由随机收集的天然和合成分子组成,或者,具有特定生物活性的天然和合成分子集。平均而言,我们最好的模型实现了超过80%的预测准确率。对于由具有特定活性的分子组成的测试用例,实现了大于90%的准确度。从我们的分析中,确定了一些化学特征,这些化学特征在许多天然存在的分子与合成分子中存在系统性差异。
Molecular descriptors were identified by Shannon entropy analysis that correctly distinguished, in binary QSAR calculations, between naturally occurring molecules and synthetic compounds. The Shannon entropy concept was first used in digital communication theory and has only very recently been applied to descriptor analysis. Binary QSAR methodology was originally developed to correlate structural features and properties of compounds with a binary formulation of biological activity (i.e., active or inactive) and has here been adapted to correlate molecular features with chemical source (i.e., natural or synthetic). We have identified a number of molecular descriptors with significantly different Shannon entropy and/or "entropic separation" in natural and synthetic compound databases. Different combinations of such descriptors and variably distributed structural keys were applied to learning sets consisting of natural and synthetic molecules and used to derive predictive binary QSAR models. These models were then applied to predict the source of compounds in different test sets consisting of randomly collected natural and synthetic molecules, or, alternatively, sets of natural and synthetic molecules with specific biological activities. On average, greater than 80% prediction accuracy was achieved with our best models. For the test case consisting of molecules with specific activities, greater than 90% accuracy was achieved. From our analysis, some chemical features were identified that systematically differ in many naturally occurring versus synthetic molecules.