Locality-Sensitive Hashing-Based k-Mer Clustering for Identification of Differential Microbial Markers Related to Host Phenotype.

Locality-Sensitive Hashing-Based k-Mer Clustering for Identification of Differential Microbial Markers Related to Host Phenotype.
复制标题

DOI:
10.1089/cmb.2021.0640
复制
发表时间:
2022-07
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

微生物在人类健康和疾病的许多方面发挥着重要作用。受众多研究表明微生物组与人类疾病之间存在关联的鼓舞,最近开发了计算和机器学习方法来生成和利用微生物组特征来预测宿主表型,例如疾病与健康癌症免疫治疗应答者与无应答者。我们之前已经开发了一种减法组装方法,其重点是提取和组装来自宏基因组数据集的差异读数,这些数据集可能来自差异基因组或两组微生物组数据集(例如,健康与疾病)之间的基因。在本文中,我们通过在多个样本中使用具有相似丰度分布的k-mers组进一步改进了我们的减法组装方法。我们实现了一种支持位置敏感散列(LSH)的方法(称为kmerLSHSA),将数十亿k-mer分组为k-mer丰度组(kcag),随后用于检索用于减法组装的差分kcag。kmerLSHSA方法在模拟数据集和真实微生物组数据集上的测试表明,与利用所有基因的传统方法相比,我们的方法可以快速识别差异基因,这些差异基因可用于构建基于微生物组的宿主表型预测的有希望的预测模型。根据k-mers在多个微生物组样本中的丰度分布,我们还讨论了lsh激活k-mers聚类的其他潜在应用。
Microbial organisms play important roles in many aspects of human health and diseases. Encouraged by the numerous studies that show the association between microbiomes and human diseases, computational and machine learning methods have been recently developed to generate and utilize microbiome features for prediction of host phenotypes such as disease versus healthy cancer immunotherapy responder versus nonresponder. We have previously developed a subtractive assembly approach, which focuses on extraction and assembly of differential reads from metagenomic data sets that are likely sampled from differential genomes or genes between two groups of microbiome data sets (e.g., healthy vs. disease). In this article, we further improved our subtractive assembly approach by utilizing groups of k-mers with similar abundance profiles across multiple samples. We implemented a locality-sensitive hashing (LSH)-enabled approach (called kmerLSHSA) to group billions of k-mers into k-mer coabundance groups (kCAGs), which were subsequently used for the retrieval of differential kCAGs for subtractive assembly. Testing of the kmerLSHSA approach on simulated data sets and real microbiome data sets showed that, compared with the conventional approach that utilizes all genes, our approach can quickly identify differential genes that can be used for building promising predictive models for microbiome-based host phenotype prediction. We also discussed other potential applications of LSH-enabled clustering of k-mers according to their abundance profiles across multiple microbiome samples.
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W
DOI: 10.1093/bioinformatics/btx304
发表时间: 2017-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Kokot, Marek;Dlugosz, Maciej;Deorowicz, Sebastian
通讯作者: Deorowicz, Sebastian
人参考肠道微生物组目录,包括来自代表性不足的亚洲元基因组的新组装基因组。
DOI: 10.1186/s13073-021-00950-7
发表时间: 2021-08-27
期刊: Genome medicine
影响因子: 12.3
作者:
Kim CY;Lee M;Yang S;Kim K;Yong D;Kim HR;Lee I
通讯作者: Lee I
DOI: 10.1186/s40168-018-0451-2
发表时间: 2018-04-11
期刊: Microbiome
影响因子: 15.5
作者:
Dai Z;Coker OO;Nakatsu G;Wu WKK;Zhao L;Chen Z;Chan FKL;Kristiansen K;Sung JJY;Wong SH;Yu J
通讯作者: Yu J