Robust detection of point mutations involved in multidrug-resistant Mycobacterium tuberculosis in the presence of co-occurrent resistance markers.

Robust detection of point mutations involved in multidrug-resistant Mycobacterium tuberculosis in the presence of co-occurrent resistance markers.
复制标题

DOI:
10.1371/journal.pcbi.1008518
复制
发表时间:
2020-12
影响因子:
4.3
通讯作者:
Clark TG
Clark TG
中科院分区:
生物学2区
文献类型:
--
作者:
Libiseller-Egger J;Phelan J;Campino S;Mohareb F;Clark TG

文献摘要

参考文献

被引文献

相似文献

结核病是一个主要的全球公共卫生问题,耐药结核分枝杆菌的日益流行使疾病控制更加困难。然而,全基因组测序作为一种诊断工具的越来越多的应用正在导致耐药性的分析,为临床实践和治疗决策提供信息。用于识别基因组数据中已建立的和新的赋予抗性的突变的计算方法包括全基因组关联研究(GWAS)方法、趋同进化测试和机器学习技术。这些方法可能会被广泛的共发生耐药性所混淆,其中一种药物的统计模型包括已知引起对其他药物耐药的不相关突变。在这里,我们引入了一种新的“同类相食”消除算法(“Hungry, Hungry SNPos”),试图去除这些共同发生的耐药变体。利用北京毒株(n = 3,574)的结核分枝杆菌基因组数据集,对五种药物(异烟肼、利福平、乙胺丁醇、吡嗪酰胺和链霉素)的表型耐药数据,我们证明了这种新方法比传统方法更强大,可以检测到耐药性相关的变异,这些变异太过罕见,不可能被基于相关性的技术(如GWAS)发现。结核病是最致命的传染病之一,每年造成100多万人死亡。致病细菌的耐药性越来越强,这阻碍了疾病的控制。与此同时,前所未有的细菌全基因组测序正在越来越多地为临床实践提供信息。为了检测产生耐药性的基因改变,并从基因组数据中预测耐药性状态,生物统计学方法和机器学习模型已经被采用。然而,由于多药耐药数据集中的耐药表型和基因型强烈重叠,这些基于相关性的方法的结果往往也包含与其他药物耐药相关的突变。在过去,这个问题经常被忽略,或者通过限制输入数据或在分析后筛选来部分解决——这两种策略都依赖于先验信息。在这里,我们提出了一种启发式算法,用于寻找与抗性相关的变异,并证明与传统技术相比,它对共发生抗性的鲁棒性要高得多。该软件可在https://github.com/julibeg/HHS上获得。
Tuberculosis disease is a major global public health concern and the growing prevalence of drug-resistant Mycobacterium tuberculosis is making disease control more difficult. However, the increasing application of whole-genome sequencing as a diagnostic tool is leading to the profiling of drug resistance to inform clinical practice and treatment decision making. Computational approaches for identifying established and novel resistance-conferring mutations in genomic data include genome-wide association study (GWAS) methodologies, tests for convergent evolution and machine learning techniques. These methods may be confounded by extensive co-occurrent resistance, where statistical models for a drug include unrelated mutations known to be causing resistance to other drugs. Here, we introduce a novel ‘cannibalistic’ elimination algorithm (“Hungry, Hungry SNPos”) that attempts to remove these co-occurrent resistant variants. Using an M. tuberculosis genomic dataset for the virulent Beijing strain-type (n = 3,574) with phenotypic resistance data across five drugs (isoniazid, rifampicin, ethambutol, pyrazinamide, and streptomycin), we demonstrate that this new approach is considerably more robust than traditional methods and detects resistance-associated variants too rare to be likely picked up by correlation-based techniques like GWAS. Tuberculosis is one of the deadliest infectious diseases, being responsible for more than one million deaths per year. The causing bacteria are becoming increasingly drug-resistant, which is hampering disease control. At the same time, an unprecedented amount of bacterial whole-genome sequencing is increasingly informing clinical practice. In order to detect the genetic alterations responsible for developing drug resistance and predict resistance status from genomic data, bio-statistical methods and machine learning models have been employed. However, due to strongly overlapping drug resistance phenotypes and genotypes in multidrug-resistant datasets, the results of these correlation-based approaches frequently also contain mutations related to resistance against other drugs. In the past, this issue has often been ignored or partially resolved by either restricting the input data or in post-analysis screening—with both strategies relying on prior information. Here we present a heuristic algorithm for finding resistance-associated variants and demonstrate that it is considerably more robust towards co-occurrent resistance compared to traditional techniques. The software is available at https://github.com/julibeg/HHS.
DOI: 10.1007/s10994-014-5451-2
发表时间: 2015-04
期刊: Machine learning
影响因子: 7.5
作者:
Ishwaran H
通讯作者: Ishwaran H
DOI: 10.1080/00401706.1970.10488634
发表时间: 1970-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1007/bf00994018
发表时间: 1995-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
CORTES, C;VAPNIK, V
通讯作者: VAPNIK, V
DOI: 10.1371/journal.pcbi.1005958
发表时间: 2018-03
影响因子: 4.3
作者:
Collins C;Didelot X
通讯作者: Didelot X