Comparative analysis of environmental sequences: potential and challenges

Comparative analysis of environmental sequences: potential and challenges
复制标题

DOI:
10.1098/rstb.2005.1809
复制
发表时间:
2006-03-29
影响因子:
6.3
通讯作者:
Bork, P
Bork, P
中科院分区:
生物学1区
文献类型:
--
作者:
Foerstner, KU;von Mering, C;Bork, P

文献摘要

被引文献

相似文献

环境测序,也被称为宏基因组学,越来越多地被用于深入了解不同栖息地的生物群落,并在生物技术和医学方面具有各种可预见的潜在应用。第一批公开的大规模数据已经提供了大量隐藏在这些环境中未知物种的大量DNA片段中的信息。比较序列分析对于这些数据的解释是必不可少的。然而,每个样本固有的不同层次的复杂性需要建立一些基线进行比较:如何规范系统发育和功能多样性的差异,如何避免不完整数据的偏差,以及如何处理物种优势或基因组大小的差异?在这里,我们讨论了这些项目中的一些,并描绘了一些简单的判别序列的性质为四个不同的栖息地。
Environmental sequencing, also dubbed metagenomics, is increasingly being used to obtain insights into organismal communities in diverse habitats, and has a variety of potential applications foreseeable in biotechnology and medicine. The first public large-scale data provide already a wealth of information hidden in vast amounts of fragmented pieces of DNA from unknown species residing in these environments. Comparative sequence analysis is essential for the interpretation of such data. However, different layers of complexity that are intrinsic to each sample require the establishment of some baselines for comparison: how to normalize for the differences in phylogenetic and functional diversity how to avoid biases from incomplete data, and how to deal with differences in species dominance or genome sizes? Here we discuss a few of these items and delineate some simple discriminative sequence properties for four distinct habitats.