Multi-Label Random Forest Model for Tuberculosis Drug Resistance Classification and Mutation Ranking

Multi-Label Random Forest Model for Tuberculosis Drug Resistance Classification and Mutation Ranking
复制标题

DOI:
10.3389/fmicb.2020.00667
复制
发表时间:
2020-04-22
影响因子:
5.2
通讯作者:
Clifton, David A.
Clifton, David A.
中科院分区:
生物学2区
文献类型:
--
作者:
Kouchaki, Samaneh;Yang, Yang;Clifton, David A.

文献摘要

被引文献

相似文献

耐药性预测和突变排序是结核病基因序列分析中的重要任务。由于使用一线抗生素的标准方案,耐药性共存(其中样品对多种药物具有耐药性)是常见的。因此,同时分析所有药物应能够利用反映耐药性共存的模式进行耐药性预测。在这里,多标签随机森林(MLRF)模型与单标签随机森林(SLRF)进行比较,用于从全基因组序列预测表型耐药和识别重要突变,以更好地预测13402个结核分枝杆菌分离株的数据集中的四种一线药物。结果证实,与传统临床方法(18.10%)和SLRF(0.91%)相比,MLRF可以提高性能。此外,我们确定了一系列对耐药预测重要或与耐药共现相关的候选突变。此外,我们发现,重新训练我们的分析,以一个子集的顶级突变是足以实现令人满意的性能。源代码可以在http://www.robots.ox.ac.uk/davidc/code.php上找到。
Resistance prediction and mutation ranking are important tasks in the analysis of Tuberculosis sequence data. Due to standard regimens for the use of first-line antibiotics, resistance co-occurrence, in which samples are resistant to multiple drugs, is common. Analysing all drugs simultaneously should therefore enable patterns reflecting resistance co-occurrence to be exploited for resistance prediction. Here, multi-label random forest (MLRF) models are compared with single-label random forest (SLRF) for both predicting phenotypic resistance from whole genome sequences and identifying important mutations for better prediction of four first-line drugs in a dataset of 13402 Mycobacterium tuberculosis isolates. Results confirmed that MLRFs can improve performance compared to conventional clinical methods (by 18.10%) and SLRFs (by 0.91%). In addition, we identified a list of candidate mutations that are important for resistance prediction or that are related to resistance co-occurrence. Moreover, we found that retraining our analysis to a subset of top-ranked mutations was sufficient to achieve satisfactory performance. The source code can be found at .http://www.robots.ox.ac.uk/davidc/code.php.