Novel phylogenetic studies of genomic sequence fragments derived from uncultured microbe mixtures in environmental and clinical samples

Novel phylogenetic studies of genomic sequence fragments derived from uncultured microbe mixtures in environmental and clinical samples
复制标题

DOI:
10.1093/dnares/dsi015
复制
发表时间:
2005-10-31
期刊:
影响因子:
4.1
通讯作者:
Ikemura, Toshimichi
Ikemura, Toshimichi
中科院分区:
生物学2区
文献类型:
--
作者:
Abe, Takashi;Sugawara, Hideaki;Ikemura, Toshimichi

文献摘要

被引文献

相似文献

自组织映射图(SOM)被开发为一种新的生物信息学策略,用于从环境和临床样本中未培养微生物的混合基因组样本中获得序列片段的系统发育分类。这种系统发育分类是可能的,既没有同源序列集,也没有序列比对。我们首先构建了从1502个原核生物中获得的210 000个5kb序列片段中的四核苷酸频率的SOM,其中至少10kb的基因组序列已经存储在公共DNA数据库中。这些序列可以主要根据系统发育群进行分类,而没有关于物种的信息。我们使用SOM方法对来自马尾藻海环境样品和酸性矿山废水中生长的嗜酸生物膜的序列片段进行了分类。在一张地图上有效地显示了环境序列的系统发育多样性。来自单个基因组但独立克隆的序列可以在电子计算机中重新关联。长期以来,G+C%一直被用作微生物系统发育分类的基本参数,但G+C%显然是一个过于简单的参数,无法区分广泛的已知物种。寡核苷酸频率可以用来区分物种,因为寡核苷酸频率在它们的基因组中差异很大。
A self-organizing map (SOM) was developed as a novel bioinformatics strategy for phylogenetic classification of sequence fragments obtained from pooled genome samples of uncultured microbes in environmental and clinical samples. This phylogenetic classification was possible without either orthologous sequence sets or sequence alignments. We first constructed SOMs for tetranucleotide frequencies in 210 000 5 kb sequence fragments obtained from 1502 prokaryotes for which at least 10 kb of genomic sequence has been deposited in public DNA databases. The sequences could be classified primarily according to phylogenetic groups without information regarding the species. We used the SOM method to classify sequence fragments derived from environmental samples of the Sargasso Sea and of an acidophilic biofilm growing in acid mine drainage. Phylogenetic diversity of the environmental sequences was effectively visualized on a single map. Sequences that were derived from a single genome but cloned independently could be reassociated in silico. G + C% has been used for a long period as a fundamental parameter for phylogenetic classification of microbes, but the G + C% is apparently too simple a parameter to differentiate a wide variety of known species. Oligonucleotide frequency can be used to distinguish the species because oligonucleotide frequencies vary significantly among their genomes.