Artificial Functional Difference Between Microbial Communities Caused by Length Difference of Sequencing Reads

Artificial Functional Difference Between Microbial Communities Caused by Length Difference of Sequencing Reads
复制标题

测序读数长度差异引起的微生物群落之间的人为功能差异

DOI:
10.1142/9789814366496_0025
复制
发表时间:
2011
影响因子:
--
通讯作者:
Yuzhen Ye
Yuzhen Ye
中科院分区:
--
文献类型:
--
作者:
Quan Zhang;T. Doak;Yuzhen Ye

文献摘要

相似文献

基于同源性的方法通常用于微生物群落的注释,提供用于表征和比较微生物群落的内容和功能的功能概况。元基因组读数是这些研究的开始数据,然而,即使对于相同的微生物群落,从不同测序技术产生的测序读数建立的功能图谱之间也存在相当大的差异。利用模拟实验,我们表明这种功能差异可能是由读取长度的实际差异引起的,而不是测序技术的采样偏差的结果。此外,来自不同测序技术的功能差异不能完全用读数偏倚来解释,即1)无注释较短读数的较高比例(即,“读长度问题”),以及2)不同功能类别中不同长度的蛋白质。相反,我们在这里表明,特定的功能类别没有得到充分的注释,因为基于相似性搜索的功能注释工具往往会错过更多来自包含保守程度较低的基因/蛋白质的功能类别的读取。此外,不同功能的短读的功能注释的准确性不同,进一步扭曲了功能图谱。为了解决这些问题,我们提出了一种简单而有效的方法来改进不同功能类别在元基因组功能谱中的频率估计,该方法基于对完整微生物基因组的模拟阅读的功能注释。
Homology-based approaches are often used for the annotation of microbial communities, providing functional profiles that are used to characterize and compare the content and the functionality of microbial communities. Metagenomic reads are the starting data for these studies, however considerable differences are observed between the functional profiles-built from sequencing reads produced by different sequencing techniques-for even the same microbial community. Using simulation experiments, we show that such functional differences are likely to be caused by the actual difference in read lengths, and are not the results of a sampling bias of the sequencing techniques. Furthermore, the functional differences derived from different sequencing techniques cannot be fully explained by the read-count bias, i.e. 1) the higher fraction of unannotated shorter reads (i.e., "read length matters"), and 2) the different lengths of proteins in different functional categories. Instead, we show here that specific functional categories are under-annotated, because similarity-search-based functional annotation tools tend to miss more reads from functional categories that contain less conserved genes/proteins. In addition, the accuracy of functional annotation of short reads for different functions varies, further skewing the functional profiles. To address these issues, we present a simple yet efficient method to improve the frequency estimates of different functional categories in the functional profiles of metagenomes, based on the functional annotation of simulated reads from complete microbial genomes.