PCAP: A whole-genome assembly program

PCAP: A whole-genome assembly program
复制标题

DOI:
10.1101/gr.1390403
复制
发表时间:
2003-09-01
期刊:
影响因子:
7
通讯作者:
Hillier, L
Hillier, L
中科院分区:
生物学1区
文献类型:
--
作者:
Huang, XQ;Wang, JM;Hillier, L

文献摘要

被引文献

相似文献

我们描述了一个名为PCAP的全基因组组装程序,用于处理数千万个读数。PCAP程序有几个功能来解决装配中的效率和精度问题。多个处理器用于执行汇编中最耗时的计算。使用更灵敏的方法来避免由测序错误引起的缺失重叠。基于与其他读段的许多重叠而不是与其他读段的许多较短的字匹配来检测读段的重复区域。识别并去除读数的污染末端区域。重叠群的共有序列的生成是基于重叠群中读段的比对,其中碱基质量值和覆盖度信息都用于确定每个共有碱基。PCAP程序在包含3000万个读数的小鼠全基因组数据集和包含170万个读数的人类20号染色体数据集上进行了测试。该计划是免费提供学术使用。
We describe a whole-genome assembly program named PCAP for processing tens of millions of reads. The PCAP program has several features to address efficiency and accuracy issues in assembly. Multiple processors are used to perform most time-consuming computations in assembly. A more sensitive method is used to avoid missing overlaps caused by sequencing errors. Repetitive regions of reads are detected oil the basis of many overlaps with other reads, instead of many shorter word matches with other reads. Contaminated end regions of reads are identified and removed. Generation of a consensus sequence for a contig is based on an alignment of reads in the contig, in which both base quality values and coverage information are used to determine every consensus base. The PCAP program was tested on a mouse whole-genome data set of 30 million reads and a human Chromosome 20 data set of 1.7 million reads. The program is freely available for academic use.