COBRE: UID: PILOT: ALGORITHMIC IMPROVEMENT OF CALLS AND READS
COBRE: UID: PILOT: ALGORITHMIC IMPROVEMENT OF CALLS AND READS
批准号:
8359581
负责人:
ROBERT B HECKENDORN
金额:
$3.69万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-02-01 至 2012-01-31
关键词:
AddressCenters of Research ExcellenceComputer softwareConsensusDataDideoxy Chain Termination DNA SequencingEvolutionFundingGenomeGrantLeadLengthMeasuresModelingNational Center for Research ResourcesPrincipal InvestigatorProcessReadingResearchResearch InfrastructureResourcesSourceTechnologyTimeUnited States National Institutes of Healthbasecomputerized data processingcostimprovedmarkov model
中文摘要
这个子项目是许多利用资源的研究子项目之一
由NIH/NCRR资助的中心拨款提供。子项目的主要支持
而子项目的主要调查员可能是由其他来源提供的,
包括其它NIH来源。 列出的子项目总成本可能
代表子项目使用的中心基础设施的估计数量,
而不是由NCRR赠款提供给子项目或子项目工作人员的直接资金。
454测序的引入显著提高了测序通量,并大大降低了测序成本。比较估计其通量比桑格测序好100倍,成本显著降低。因此,它代表了测序技术的明显进步。
然而,454测序仍然受到几个重大弱点的限制。与桑格测序相比,读取长度相对较短-尽管技术的改进导致更长的读取。此外,454读取具有相对较高的错误率。这些弱点是密切相关的,因为错误率随着读取长度的增加而显著增加,使得错误是读取长度的主要限制因素之一-超过某个点,读取数据充满错误以至于无用。
更短且更容易出错的读段使得更难以产生较长(例如非细菌)基因组的成功、准确序列。这一弱点可以通过提高覆盖率来部分克服。
多个重叠读段用于鉴定错误读段并构建“共有”读段。然而,产生额外的覆盖增加了成本和时间,减少了454测序的优势。
我们建议通过以下方式解决读取错误的问题:1)通过测量454个威尔斯孔的测量强度之间的相关性,并使用这些相关性来建模和校正调用错误,以及2)通过
使用原始强度数据(而不是像目前所做的那样,仅仅是所谓的碱基)结合隐马尔可夫模型来产生更准确的共有读段。所得到的误差校正算法将被封装成易于插入到当前454数据处理流水线中的软件。这项研究的结果将提高生成的454序列的质量,并有助于最大限度地发挥454测序的高通量和低成本的优势,通过限制对冗余读段进行纠错。
英文摘要
This subproject is one of many research subprojects utilizing the resources
provided by a Center grant funded by NIH/NCRR. Primary support for the subproject
and the subproject's principal investigator may have been provided by other sources,
including other NIH sources. The Total Cost listed for the subproject likely
represents the estimated amount of Center infrastructure utilized by the subproject,
not direct funding provided by the NCRR grant to the subproject or subproject staff.
The introduction of 454 Sequencing has lead to significant improvements in sequencing throughput and dramatically reduced sequencing costs. Comparisons estimate that its throughput is 100 fold better than Sanger sequencing at significantly reduced cost. Thus, it represents a clear advance in sequencing technologies.
However, 454 Sequencing is still limited by several significant weaknesses. Compared to Sanger sequencing, the read lengths are relatively short - although improvements in the technology are leading to longer reads. Additionally, 454 reads have comparitively high error rates. These weaknesses are closely related as error rates increase significantly as read lengths increase making errors are one of the major limiting factors on read length - past a certain point the read data is so error-filled as to be useless.
Shorter, and more error prone, reads makes it more difficult to generate successful, accurate sequences of longer, e.g. non-bacterial, genomes. This weakness can be partially overcome through higher coverage.
Multiple overlapping reads are used to identify erroneous reads and build 'consensus' reads. However, generating additional coverage increases cost and time, reducing the advantages of 454 Sequencing.
We propose to address the problem of read errors by 1) by measuring the correlations between measured intensities of 454 wells and using those correlations to model and correct calling errors and 2) by
using the raw intensity data (rather than just the called bases, as is currently done) in combination with Hidden Markov models to produce more accurate consensus reads. The resulting error correction agorithms will be packaged into software that is easy to insert into the current 454 data processing pipline. The results of this research would both improve the quality of generated 454 sequences and help to maximize 454 Sequencing's advantages of high throughput and low cost, by limiting the need for redundant reads for error correction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文