A strategy for assembling the maize (Zea mays L.) genome

A strategy for assembling the maize (Zea mays L.) genome
复制标题

DOI:
10.1093/bioinformatics/bth017
复制
发表时间:
2004-01-22
期刊:
影响因子:
5.8
通讯作者:
Schnable, PS
Schnable, PS
中科院分区:
生物学3区
文献类型:
--
作者:
Emrich, SJ;Aluru, S;Schnable, PS

文献摘要

被引文献

相似文献

由于玉米(Zea mays L.)基因组由重复序列组成,测序工作正瞄准其“基因丰富”的部分。传统的组装程序是不足以为这种方法,因为它们是优化的基因组的统一采样,并固有地缺乏区分高度相似的旁系同源物的能力。结果:我们报告的生物信息学工具的发展,玉米基因组的准确组装。该软件,这是基于创新的并行算法,以确保可扩展性,组装730 974基因组调查序列片段在4小时内使用64奔腾III 1.26 GHz处理器的商品集群。在不牺牲质量的情况下,使用化学创新来显著减少成对比对的数量。使用克隆对信息来估计用于改进多态性与测序错误的区分的错误率。该组装还用于评估各种过滤策略的有效性,从而提供可用于集中后续测序工作的信息。
Summary: Because the bulk of the maize (Zea mays L.) genome consists of repetitive sequences, sequencing efforts are being targeted to its 'gene-rich' fraction. Traditional assembly programs are inadequate for this approach because they are optimized for a uniform sampling of the genome and inherently lack the ability to differentiate highly similar paralogs.Results: We report the development of bioinformatics tools for the accurate assembly of the maize genome. This software, which is based on innovative parallel algorithms to ensure scalability, assembled 730 974 genomic survey sequences fragments in 4 h using 64 Pentium III 1.26 GHz processors of a commodity cluster. Algorithmic innovations are used to reduce the number of pairwise alignments significantly without sacrificing quality. Clone pair information was used to estimate the error rate for improved differentiation of polymorphisms versus sequencing errors. The assembly was also used to evaluate the effectiveness of various filtering strategies and thereby provide information that can be used to focus subsequent sequencing efforts.