Predicting demographic group structures based on DNA sequence data

Predicting demographic group structures based on DNA sequence data
复制标题

DOI:
10.1093/molbev/msg128
复制
发表时间:
2003-07-01
影响因子:
10.7
通讯作者:
Mullins, JI
Mullins, JI
中科院分区:
生物学1区
文献类型:
--
作者:
Anderson, JP;Learn, GH;Mullins, JI

文献摘要

被引文献

相似文献

通过搜索它们的进化历史或通过比较它们的序列相似性来推断序列组之间的关系的能力可以是假设检验中的关键步骤。解释人类免疫缺陷病毒1型(HIV-1)序列的关系可能具有挑战性,因为它们的基因组迅速演变,但它也可能导致更好地了解潜在的生物学。一些研究集中在HIV-1的进化上,但很少有信息将HIV-1的序列相似性和进化历史与感染个体的流行病学信息联系起来。我们的目标是将HIV-1遗传多样性模式与流行病学信息(包括风险和人口统计学因素)相关联。这些相关性随后被用来通过分析HIV-1序列的短片段来预测流行病学信息。使用标准的系统发育和表型技术对100个HIV-1亚型B序列,我们能够显示病毒序列和感染的地理区域以及男性与男性发生性关系的风险之间的一些相关性。为了帮助识别病毒序列之间更微妙的关系,进行多维标度(MDS)方法。该方法确定了病毒序列与男男性行为者和与注射毒品使用者发生性关系或自己使用注射毒品者的风险因素之间的统计学显著相关性。使用树结构,MDS和新开发的似然分配方法对我们测序的原始100个样本,以及一组设盲样本,我们能够预测人口统计学/风险组成员资格,其统计学比率优于仅凭偶然。这种方法可以通过仅检查HIV-1基因组的一小部分来鉴定属于特定人口群体的病毒变体。这种基于序列信息的人口流行病学预测在为感染个体分配不同的治疗方案方面可能变得有价值。
The ability to infer relationships between groups of sequences, either by searching for their evolutionary history or by comparing their sequence similarity, can be a crucial step in hypothesis testing. Interpreting relationships of human immunodeficiency virus type 1 (HIV-1) sequences can be challenging because of their rapidly evolving genotnes, but it may also lead to a better understanding of the underlying biology. Several studies have focused on the evolution of HIV-1, but there is little information to link sequence similarities and evolutionary histories of HIV-1 to the epidemiological information of the infected individual. Our goal was to correlate patterns of HIV-1 genetic diversity with epidemiological information, including risk and demographic factors. These correlations were then used to predict epidemiological information through analyzing short stretches of HIV-1 sequence. Using standard phylogenetic and phenetic techniques on 100 HIV-1 subtype B sequences, we were able to show some correlation between the viral sequences and the geographic area of infection and the risk of men who engage in sex with men. To help identify more subtle relationships between the viral sequences, the method of multidimensional scaling (MDS) was performed. That method identified statistically significant correlations between the viral sequences and the risk factors of men who engage in sex with men and individuals who engage in sex with injection drug users or use injection drugs themselves. Using tree construction, MDS, and newly developed likelihood assignment methods on the original 100 samples we sequenced, and also on a set of blinded samples, we were able to predict demographic/risk group membership at a rate statistically better than by chance alone. Such methods may make it possible to identify viral variants belonging to specific demographic groups by examining only a small portion of the HIV-1 genome. Such predictions of demographic epidemiology based on sequence information may become valuable in assigning different treatment regimens to infected individuals.