课题基金 / 基金详情

Improving overlap-finding techniques for whole genome shotgun data

Improving overlap-finding techniques for whole genome shotgun data
改进全基因组鸟枪数据的重叠查找技术
批准号:
0312360
负责人:
James Yorke
金额:
$9.94万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-07-15 至 2005-06-30

项目摘要

项目成果

James Yorke的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Yorke A genome (the DNA in a cell) can be represented by asequence of letters called "bases." A large genome can consistof billions of bases. Chemical techniques allow scientists toread only a few hundred bases at a time. The whole genome shotgun(WGS) assembly technique creates a draft of the sequence of awhole genome by selecting such short fragments at random from thegenome, determining the sequence of the fragments, and thencomputationally re-assembling millions of these fragments. Twofragments are said to "overlap" if it is plausible that they comefrom the same part of the genome, based on a comparison of theirsequences. The goal of this project is to focus efforts onproducing an extremely robust set of overlaps, using acombination of sophisticated error-correction techniques, as wellas "localizing" fragments to validate overlaps by ensuring thatboth fragments come from the same vicinity of the genome.Several issues complicate the determination of which pairs offragments overlap. First, most genomes contain many "repeatregions," i.e., two or more almost identical copies of longstretches of sequence. Thus, two fragments that do not actuallyoverlap may look like they do. Second, the random samplingtechnique results in many base errors --- bases can be mis-reador missed entirely. These errors, combined with the fact thatrepeat regions usually differ slightly, make it very difficult todistinguish a spurious overlap from a true overlap in which oneor both fragments contain read errors. Thus, if extreme care isnot taken, it is easy to use a spurious overlap and therebymistakenly connect distant parts of the genome. Preliminaryresults in collaboration with Celera Genomics, the Baylor Collegeof Medicine, and The Institute for Genomic Research (TIGR) havedemonstrated that the investigator's current techniques canalready produce more sequence at higher quality. The goal isimprove these techniques and make them widely available. The determination and interpretation of genetic informationis one of the great challenges of the twenty-first century. Thegenome, i.e., all the DNA in a cell, is the molecular basis ofdiversity and the cornerstone of genetic information. Draftgenomes have been obtained for human, mouse, and some insects,fish, plants, and bacteria. This is a start, but a fullunderstanding of biological processes cannot be had by studyingthe genomes of only a handful of species. The federal governmentis spending about 100 million dollars per year generatingsequence data. Millions of small pieces of a genome are sampledfrom the genome. The second stage is called "assembly," whenthese pieces are re-assembled on a computer like a giant jigsawpuzzle. The puzzle is complicated by two facts: first, many ofthe puzzle pieces have small errors that make them mis-fitagainst pieces that they SHOULD fit with; and second, many piecesthat should NOT go together actually fit together quite well.This makes it extremely difficult to correctly assemble a genome.There are two ways to decrease the ambiguities: first, one couldgenerate more pieces. However, each new piece costs about $2,and one would need to generate millions of new pieces to have asignificant effect on assembly quality. The investigators use asecond route. They attempt to squeeze as much information out ofthe existing pieces as possible. The latter route issubstantially cheaper, and there is still much room forimprovement here over existing techniques. The investigators areusing sophisticated mathematics to help discern with extremeprecision those pairs of pieces that do, and those that do not,fit together. Preliminary results of the investigators -- incollaboration with several large sequencing centers -- havedemonstrated that using their techniques to "pre-process" thepieces can produce more of the genome, with fewer errors. Thisproject aims at extending these ideas further and making themfreely accessible to all investigators. The impact on the federalgenome (biotechnology) projects is potentially great.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Mathematical Modeling of DNA Repeats and HIV Epidemics
Applications of Nonlinear Dynamics
Chaos with Multiple Positive Lyapunov Exponents
Mathematical Sciences: "Chaos with Multiple Positive Lyapunov Exponents
  • 批准号:
    9423843
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $28.36万
  • 财政年份:
    1995
  • 负责人:
    James Yorke
  • 依托单位:
海外基金