COMPILATION, ALIGNMENT, AND PHYLOGENETIC-RELATIONSHIPS OF DNA-POLYMERASES
COMPILATION, ALIGNMENT, AND PHYLOGENETIC-RELATIONSHIPS OF DNA-POLYMERASES
复制标题
DOI:
10.1093/nar/21.4.787
复制
发表时间:
1993-02-25
影响因子:
14.9
通讯作者:
ITO, J
中科院分区:
文献类型:
--
作者:
BRAITHWAITE, DK;ITO, J
SEQUENCE ALIGNMENT The multiple alignments ofthe amino acid sequences for this update were performed in most cases by merely adding on to our original alignments (1) where possible. Due to the large number of sequences added to the alignment for Family B we have changed the original alignment in some areas between obvious blocks of conserved sequences. Thenewer sequences were added by aligning each to the closest related sequence already aligned, or in many cases to the closest related group of sequences already aligned. A more recent addition to the UWGCG (University of Wisconsin Genetic Computer Group) program package, PILEUP, a multiple alignment program, was used extensively to try and locate significant homology in groups of closely related sequences. These newly formed groups of highly related sequences were then regapped to conform with the entire alignment based upon the previous alignment of those sequences in the new group from the original alignment. As in the previous paper, all the final adjustments had to be made by eye and, as stated above, in Family B the added sequences led to some improvements to the original alignment that became evident to the eye when they were being combined with the entire alignment by hand.GENERATION OF PHYLOGENETIC TREES FOR THE DNA POLYMERASE DOMAINS Using Felsenstein's PHYLIP program package (71), specifically the programs named in the outline below, we generated phylogenetic trees for the 9 Family A DNA polymerases (Figures 2A and 2B) and for the 47 Family B DNA polymerases (Figures 3A and 3B). The trees for Family A were created from the alignment in Figure 1A using the most conserved regions found at the following positions: 798 to 814, 877 to 998, 1047 to 1090, 1104 to 1123, 1131 to 1158, 1175 to 1206, 1236 to 1251, 1284 to 1305, 1322 to 1340, and 1365 to 1379. These conserved regions were recombined and 100 bootstrap samples were generated using SEQBOOT program. Using the DNADIST program, we turned the samples into distance matrices using the Kimura-2 parameter method. The resulting matrices were then input to the NEIGHBOR program using the UPGMA method to produce approximately 100 trees. Finally those trees were reduced to a single tree using the CONSENSE program. This final tree was then plotted for publication using two different methods. The trees in Figures 2A and 3A were created by the DRAWGRAM program setup to produce a phenogram type tree and the trees in Figures 2B and 3B were created by the DRAWTREE program. The trees for Family B were created from the alignment in Figure 1B, according to the same procedure, using the most conserved regions found at the following positions: 1407 to 1760, 1885 to 1901, 1956 to 1990, 2081 to 2100, 2181 to 2210, and 2280 to 2320. The Family B DNA polymerases can be subdivided into two subfamilies, the protein-primed DNA polymerase subfamily and the RNA-primed DNA polymerase subfamily.