Reference-free validation of short read data.

Reference-free validation of short read data.
复制标题

DOI:
10.1371/journal.pone.0012681
复制
发表时间:
2010-09-22
期刊:
影响因子:
3.7
通讯作者:
Zobel J
Zobel J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Schröder J;Bailey J;Conway T;Zobel J

文献摘要

参考文献

被引文献

相似文献

高通量DNA测序技术提供了快速且廉价地对材料(例如整个基因组)进行测序的能力。然而,通过这些技术产生的短读段数据在测序过程的几个阶段可能会有偏差或受损;这些偏差中的一些的来源和性质并不总是已知的。准确的偏倚评估是实验质量控制、基因组组装和覆盖率结果解释所必需的。另一个挑战是,对于来自未鉴定来源的新基因组或材料,可能没有可用于检查读数的参考。我们提出了一种分析方法,用于识别短读段集合中的偏差,而无需参考。这些与现有方法结合,构成了可用于量化一组读数的质量的方法。我们的方法涉及使用三种不同的措施:分析的基础调用;分析的k-mer;和分析的分布k-mer。我们将我们的方法应用于广泛的短读数据,并显示,令人惊讶的是,强烈的偏见似乎存在。这些包括一些多碱基序列的过度表达,对某些碱基的每位置偏差,以及对某些起始位置的明显偏好。短读数据中存在偏倚是已知的,但它们似乎比以前文献中确定的更大,更多样化。对一组短读段的统计分析可以帮助在组装或重新测序之前识别问题,并且应该有助于指导用于偏差校正的化学或统计方法。
High-throughput DNA sequencing techniques offer the ability to rapidly and cheaply sequence material such as whole genomes. However, the short-read data produced by these techniques can be biased or compromised at several stages in the sequencing process; the sources and properties of some of these biases are not always known. Accurate assessment of bias is required for experimental quality control, genome assembly, and interpretation of coverage results. An additional challenge is that, for new genomes or material from an unidentified source, there may be no reference available against which the reads can be checked. We propose analytical methods for identifying biases in a collection of short reads, without recourse to a reference. These, in conjunction with existing approaches, comprise a methodology that can be used to quantify the quality of a set of reads. Our methods involve use of three different measures: analysis of base calls; analysis of k-mers; and analysis of distributions of k-mers. We apply our methodology to wide range of short read data and show that, surprisingly, strong biases appear to be present. These include gross overrepresentation of some poly-base sequences, per-position biases towards some bases, and apparent preferences for some starting positions over others. The existence of biases in short read data is known, but they appear to be greater and more diverse than identified in previous literature. Statistical analysis of a set of short reads can help identify issues prior to assembly or resequencing, and should help guide chemical or statistical methods for bias rectification.
DOI: 10.1186/gb-2009-10-8-r83
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Kircher M;Stenzel U;Kelso J
通讯作者: Kelso J
亚洲个体的二倍体基因组序列
DOI: 10.1038/nature07484
发表时间: 2008-11-06
期刊: NATURE
影响因子: 64.8
作者:
Wang, Jun;Wang, Wei;Li, Ruiqiang;Li, Yingrui;Tian, Geng;Goodman, Laurie;Fan, Wei;Zhang, Junqing;Li, Jun;Zhang, Juanbin;Guo, Yiran;Feng, Binxiao;Li, Heng;Lu, Yao;Fang, Xiaodong;Liang, Huiqing;Du, Zhenglin;Li, Dong;Zhao, Yiqing;Hu, Yujie;Yang, Zhenzhen;Zheng, Hancheng;Hellmann, Ines;Inouye, Michael;Pool, John;Yi, Xin;Zhao, Jing;Duan, Jinjie;Zhou, Yan;Qin, Junjie;Ma, Lijia;Li, Guoqing;Yang, Zhentao;Zhang, Guojie;Yang, Bin;Yu, Chang;Liang, Fang;Li, Wenjie;Li, Shaochuan;Li, Dawei;Ni, Peixiang;Ruan, Jue;Li, Qibin;Zhu, Hongmei;Liu, Dongyuan;Lu, Zhike;Li, Ning;Guo, Guangwu;Zhang, Jianguo;Ye, Jia;Fang, Lin;Hao, Qin;Chen, Quan;Liang, Yu;Su, Yeyang;San, A.;Ping, Cuo;Yang, Shuang;Chen, Fang;Li, Li;Zhou, Ke;Zheng, Hongkun;Ren, Yuanyuan;Yang, Ling;Gao, Yang;Yang, Guohua;Li, Zhuo;Feng, Xiaoli;Kristiansen, Karsten;Wong, Gane Ka-Shu;Nielsen, Rasmus;Durbin, Richard;Bolund, Lars;Zhang, Xiuqing;Li, Songgang;Yang, Huanming;Wang, Jian
通讯作者: Wang, Jian
DOI: 10.1101/gr.079053.108
发表时间: 2009-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Chaisson, Mark J.;Brinza, Dumitru;Pevzner, Pavel A.
通讯作者: Pevzner, Pavel A.
DOI: 10.1038/nmeth.1226
发表时间: 2008-07-01
期刊: NATURE METHODS
影响因子: 48
作者:
Mortazavi, Ali;Williams, Brian A.;Wold, Barbara
通讯作者: Wold, Barbara
DOI: 10.1186/1471-2105-9-431
发表时间: 2008-10-13
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Rougemont, Jacques;Amzallag, Arnaud;Naef, Felix
通讯作者: Naef, Felix