HIV sequence compendium 2002

HIV sequence compendium 2002
复制标题

HIV序列简编2002

DOI:
10.2172/1184349
复制
发表时间:
2002
期刊:
Electroencephalography and clinical neurophysiology
影响因子:
--
通讯作者:
B. Korber
B. Korber
中科院分区:
--
文献类型:
--
作者:
C. Kuiken;B. Foley;E. Freed;B. Hahn;P. Marx;F. McCutchan;J. Mellors;Steven Wolinsky;B. Korber

文献摘要

被引文献

相似文献

本简编是艾滋病毒序列数据库所载数据的年度印刷摘要。在这些概要中,我们试图以一种对艾滋病毒研究人员最有用的方式来呈现数据的明智选择。传统上,我们以比对的形式呈现序列数据本身:第二部分,一组HIV-1/SIVcpz全长基因组的比对(例如,许多类似LAI的序列被省略,因为它们太相似,以至于偏离了比对);第三部分,组合的HIV-1/HIV-2/SIV全基因组比对;第四-第六部分,HIV-1/SIV-CPZ、HIV-2/SIV和SIVagm的氨基酸比对。HIV-2/SIV和SIVagm的氨基酸比对是分开的,因为这些组之间的遗传距离如此之大,以至于将它们放在一个比对中会使其非常拉长,因为必须插入大量的间隙。像往常一样,从文献中收集的大量背景信息的表格伴随着整个基因组比对。数据库中的全基因序列集合现在已经足够大,我们有丰富的代表大多数亚型。对于许多亚型,特别是B亚型,大量跨越整个基因的序列并没有更多地包括在打印的比对中,以节省空间。所有路线的更完整版本可在我们的网站http://hiv-web.lanl.gov/content/hiv-db/ALIGN_CURRENT/ALIGN-INDEX.html.上找到重要的是,所有这些比对都经过编辑,根据为所有人创建的系统发育树和文献,只包括一个人的序列。由于现有序列的数量,我们决定今年根据亚型的流行病学重要性使用不同的选择原则。A-D亚型和CRF 01和02亚型是迄今为止分布最广泛的变异体,对于这些变异体(如果有),我们在比对中包括了8-10个代表。其他亚型和CRF的重要性较小,在这4-5个亚型和CRF中,每个或尽可能多的亚型都被包括在内。在比对中,我们还包括了“循环重组形式”,即具有流行病学意义的马赛克基因组。有关CRF的更多信息,请参阅1999年的命名法(http://hiv-web.lanl.gov/content/hiv-db/REVIEWS/nomenclature/Nomen.html)回顾,有关已知CRF的模式概述,请参阅。氨基酸比对章节以注释表开始,该注释表包括序列名称、登录号、代表的基因组区域、作者和参考文献。我们已经作出努力,使艾滋病毒-2/SIV和SIVagm的比对也是最新的。《更少》
This compendium is an annual printed summary of the data contained in the HIV sequence database. In these compendia we try to present a judicious selection of the data in such a way that it is of maximum utility to HIV researchers. Traditionally, we present the sequence data themselves in the form of alignments: Section II, an alignment of a selection of HIV-1/SIVcpz full-length genomes (a lot of LAI-like sequences, for example, have been omitted because they are so similar that they bias the alignment); Section III, a combined HIV-1/HIV-2/SIV whole genome alignment; Sections IV–VI, amino acid alignments for HIV-1/SIV-cpz, HIV-2/SIV, and SIVagm. The HIV-2/SIV and SIVagm amino acid alignments are separate because the genetic distances between these groups are so great that presenting them in one alignment would make it very elongated because of the large number of gaps that have to be inserted. As always, tables with extensive background information gathered from the literature accompany the whole genome alignments. The collection of whole-gene sequences in the database is now large enough that we have abundant representation of most subtypes. For many subtypes, and especially for subtype B, a large number of sequences that span entire genes were not more » included in the printed alignments to conserve space. A more complete version of all alignments is available on our website, http://hiv-web.lanl.gov/content/hiv-db/ALIGN_CURRENT/ALIGN-INDEX.html. Importantly, all these alignments have been edited to include only one sequence per person, based on phylogenetic trees that were created for all of them, as well as on the literature. Because of the number of sequences available, we have decided to use a different selection principle this year, based on the epidemiological importance of the subtypes. Subtypes A–D and CRFs 01 and 02 are by far the most widespread variants, and for these (when available) we have included 8–10 representatives in the alignments. The other subtypes and CRFs are of lesser importance, and of these 4–5 each, or as many as are available, were included. In the alignments we have also included the ‘Circulating Recombinant Forms’, mosaic genomes that have epidemiological significance. See the 1999 review of nomenclature (http://hiv-web.lanl.gov/content/hiv-db/REVIEWS/nomenclature/Nomen.html) for more on CRFs, and see for an overview of the patterns of known CRFs. Amino acid alignment chapters begin with an annotation table that includes sequence names, accession numbers, genomic region represented, author, and references. We have made an effort to bring the HIV-2/SIV and SIVagm alignments up-to-date as well. « less