Significantly Improved HIV Inhibitor Efficacy Prediction Employing Proteochemometric Models Generated From Antivirogram Data

Significantly Improved HIV Inhibitor Efficacy Prediction Employing Proteochemometric Models Generated From Antivirogram Data
复制标题

DOI:
10.1371/journal.pcbi.1002899
复制
发表时间:
2013-02-01
影响因子:
4.3
通讯作者:
Bender, Andreas
Bender, Andreas
中科院分区:
生物学2区
文献类型:
--
作者:
van Westen, Gerard J. P.;Hendriks, Alwin;Bender, Andreas

文献摘要

被引文献

相似文献

艾滋病毒感染目前无法治愈;然而,它可以通过多种抗逆转录病毒药物的联合治疗来控制。鉴于几乎每个病人的病毒基因型都不同,现在的问题是使用哪种药物组合来实现有效的治疗。随着病毒基因型数据和临床表型数据的可用性,创建能够预测个体患者最佳治疗方案的计算模型已经成为可能。目前的模型仅基于来自病毒基因分型的序列数据;不考虑药物的化学相似性。为了探索化学相似包涵的附加价值,我们应用了蛋白质化学计量模型,将化学和蛋白质靶标特性结合在一个单一的生物活性模型中。我们的数据集是一个包含基因型和表型信息的大型临床数据库(总共约300,000个药物突变生物活性数据点,4个(NNRTI), 8个(NRTI)或9个(PI)药物,10,700个(NNRTI), 10,500个(NRTI)或27,000个(PI)突变)。我们的模型实现了低于0.5 Log Fold Change的预测误差。此外,当直接与先前发表的序列数据进行比较时,推导的模型PCM在抗性分类和对数褶皱变化预测方面表现更好(0.76对数单位对0.91)。此外,我们能够成功地从我们的数据集中确认已知和鉴定以前未发表的HIV逆转录酶(如K102Y, T216M)和HIV蛋白酶(如Q18N, N88G)的耐药突变。最后,我们将我们的模型前瞻性地应用于斯坦福大学的公共HIV耐药性数据库,在完整的集合上获得了84%的正确耐药性预测率(相比之下,之前在高质量子集上的工作为80%)。我们的结论是,蛋白质化学计量模型能够准确地预测基于基因型数据的表型抗性,即使是新的突变体和混合物。此外,我们为预测添加了一个适用域,告知用户预测的可靠性。
Infection with HIV cannot currently be cured; however it can be controlled by combination treatment with multiple anti-retroviral drugs. Given different viral genotypes for virtually each individual patient, the question now arises which drug combination to use to achieve effective treatment. With the availability of viral genotypic data and clinical phenotypic data, it has become possible to create computational models able to predict an optimal treatment regimen for an individual patient. Current models are based only on sequence data derived from viral genotyping; chemical similarity of drugs is not considered. To explore the added value of chemical similarity inclusion we applied proteochemometric models, combining chemical and protein target properties in a single bioactivity model. Our dataset was a large scale clinical database of genotypic and phenotypic information (in total ca. 300,000 drug-mutant bioactivity data points, 4 (NNRTI), 8 (NRTI) or 9 (PI) drugs, and 10,700 (NNRTI) 10,500 (NRTI) or 27,000 (PI) mutants). Our models achieved a prediction error below 0.5 Log Fold Change. Moreover, when directly compared with previously published sequence data, derived models PCM performed better in resistance classification and prediction of Log Fold Change (0.76 log units versus 0.91). Furthermore, we were able to successfully confirm both known and identify previously unpublished, resistance-conferring mutations of HIV Reverse Transcriptase (e. g. K102Y, T216M) and HIV Protease (e. g. Q18N, N88G) from our dataset. Finally, we applied our models prospectively to the public HIV resistance database from Stanford University obtaining a correct resistance prediction rate of 84% on the full set (compared to 80% in previous work on a high quality subset). We conclude that proteochemometric models are able to accurately predict the phenotypic resistance based on genotypic data even for novel mutants and mixtures. Furthermore, we add an applicability domain to the prediction, informing the user about the reliability of predictions.