Using reads to annotate the genome: influence of length, background distribution, and sequence errors on prediction capacity

Using reads to annotate the genome: influence of length, background distribution, and sequence errors on prediction capacity
复制标题

DOI:
10.1093/nar/gkp492
复制
发表时间:
2009-08-01
影响因子:
14.9
通讯作者:
Rivals, Eric
Rivals, Eric
中科院分区:
生物学2区
文献类型:
--
作者:
Philippe, Nicolas;Boureux, Anthony;Rivals, Eric

文献摘要

被引文献

相似文献

超高通量测序被用来在全基因组范围内以前所未有的深度分析转录组或相互作用组。这些技术产生短序列读数,然后将其映射到基因组序列上,以预测推测的转录或蛋白质相互作用区域。我们认为,背景分布、序列误差和读数长度等因素对序列普查实验的预测能力有影响。在这里,我们建议一种计算方法来衡量这些因素,并分析它们对转录和表观基因组分析的影响。这项调查为方法学和生物学问题提供了新的线索。例如,通过分析染色质免疫沉淀阅读集,我们估计4.6%的阅读受到SNPs的影响。我们发现,尽管核苷酸错误概率很低,但它随着序列中位置的增加而显著增加。选择19个碱基对以上的读取长度实际上消除了找到不相关位置的风险,而选择20个碱基对以上的唯一映射读取的数量减少。使用我们的程序,我们在基因组位置之间获得了0.6%的假阳性。因此,即使是罕见的签名也应该识别生物相关的区域,如果它们被映射到基因组上的话。这表明,数字转录学可能有助于描述尚未发现的低丰度转录本的丰富特征。
Ultra high-throughput sequencing is used to analyse the transcriptome or interactome at unprecedented depth on a genome-wide scale. These techniques yield short sequence reads that are then mapped on a genome sequence to predict putatively transcribed or protein-interacting regions. We argue that factors such as background distribution, sequence errors, and read length impact on the prediction capacity of sequence census experiments. Here we suggest a computational approach to measure these factors and analyse their influence on both transcriptomic and epigenomic assays. This investigation provides new clues on both methodological and biological issues. For instance, by analysing chromatin immunoprecipitation read sets, we estimate that 4.6% of reads are affected by SNPs. We show that, although the nucleotide error probability is low, it significantly increases with the position in the sequence. Choosing a read length above 19 bp practically eliminates the risk of finding irrelevant positions, while above 20 bp the number of uniquely mapped reads decreases. With our procedure, we obtain 0.6% false positives among genomic locations. Hence, even rare signatures should identify biologically relevant regions, if they are mapped on the genome. This indicates that digital transcriptomics may help to characterize the wealth of yet undiscovered, low-abundance transcripts.