Predicting HIV drug resistance with neural networks

Predicting HIV drug resistance with neural networks
复制标题

DOI:
10.1093/bioinformatics/19.1.98
复制
发表时间:
2003-01-01
期刊:
影响因子:
5.8
通讯作者:
Potter, RB
Potter, RB
中科院分区:
生物学3区
文献类型:
--
作者:
Draghici, S;Potter, RB

文献摘要

被引文献

相似文献

动机:耐药性是影响当前HIV治疗失败的一个重要因素。预测HIV蛋白酶突变体耐药性的能力可能有助于开发更有效和更持久的治疗方案。方法:预测HIV对两种现有蛋白酶抑制剂英地那韦和沙奎那韦的耐药性。这个问题是从两个角度来探讨的。首先,基于HIV蛋白酶-药物抑制剂复合物的结构特征构建预测因子。一种特殊的结构由抑制剂和蛋白酶之间的接触表来表示。其次,基于各种耐药突变体的序列数据构建分类器。在这两种情况下,首先使用自组织映射来提取重要特征,并以无监督的方式对模式进行聚类。接下来是基于训练集中已知模式的后续标记。结果:对分类器的预测性能进行了交叉验证。使用结构信息的分类器正确地分类了以前未见过的突变体,准确率在60%到70%之间。在更丰富的序列数据上测试了几种体系结构。最好的单一分类器提供了68%的准确率和69%的覆盖率。然后将多个网络组合成各种多数投票方案。最佳组合在以前未见过的数据上平均产生85%的覆盖率和78%的准确率。这比随机分类器预期的33%准确率高出两倍多。
Motivation: Drug resistance is a very important factor influencing the failure of current HIV therapies. The ability to predict the drug resistance of HIV protease mutants may be useful in developing more effective and longer lasting treatment regimens.Methods: The HIV resistance is predicted to two current protease inhibitors, Indinavir and Saquinavir. The problem was approached from two perspectives. First, a predictor was constructed based on the structural features of the HIV protease-drug inhibitor complex. A particular structure was represented by its list of contacts between the inhibitor and the protease. Next, a classifier was constructed based on the sequence data of various drug resistant mutants. In both cases, self-organizing maps were first used to extract the important features and cluster the patterns in an unsupervised manner. This was followed by subsequent labelling based on the known patterns in the training set.Results: The prediction performance of the classifiers was measured by cross-validation. The classifier using the structure information correctly classified previously unseen mutants with an accuracy of between 60 and 70%. Several architectures were tested on the more abundant sequence data. The best single classifier provided an accuracy of 68% and a coverage of 69%. Multiple networks were then combined into various majority voting schemes. The best combination yielded an average of 85% coverage and 78% accuracy on previously unseen data. This is more than two times better than the 33% accuracy expected from a random classifier.