Reconstruction of Viral Variants via Monte Carlo Clustering

Reconstruction of Viral Variants via Monte Carlo Clustering
复制标题

DOI:
10.1089/cmb.2023.0154
复制
发表时间:
2023-09-01
影响因子:
1.7
通讯作者:
Zelikovsky,Alex
Zelikovsky,Alex
中科院分区:
生物学4区
文献类型:
--
作者:
Juyal,Akshay;Hosseini,Roya;Zelikovsky,Alex

文献摘要

相似文献

通过聚类识别病毒变体对于了解宿主内和宿主之间病毒群体的组成和结构至关重要,这在疾病进展和流行病传播中起着至关重要的作用。本文提出并验证了新的Monte Carlo(MC)方法,通过最小化熵或汉明距离的共识聚类比对病毒序列。我们在四个基准上验证了这些方法:两个SARS-CoV-2宿主间数据集和两个HIV宿主内数据集。我们的工具的并行化版本是可扩展的非常大的数据集。我们表明,熵和汉明距离为基础的MC聚类识别有意义的信息从测序数据。所提出的聚类方法在不同的运行中始终收敛到类似的聚类。最后,我们表明,MC聚类提高了从测序数据重建宿主内病毒群体。
Identifying viral variants through clustering is essential for understanding the composition and structure of viral populations within and between hosts, which play a crucial role in disease progression and epidemic spread. This article proposes and validates novel Monte Carlo (MC) methods for clustering aligned viral sequences by minimizing either entropy or Hamming distance from consensuses. We validate these methods on four benchmarks: two SARS-CoV-2 interhost data sets and two HIV intrahost data sets. A parallelized version of our tool is scalable to very large data sets. We show that both entropy and Hamming distance-based MC clusterings discern the meaningful information from sequencing data. The proposed clustering methods consistently converge to similar clusterings across different runs. Finally, we show that MC clustering improves reconstruction of intrahost viral population from sequencing data.