The Terabase Search Engine
The Terabase Search Engine
批准号:
8688406
负责人:
Steven L. Salzberg
金额:
$35.5万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-01 至 2017-04-30
关键词:
AccelerationAffectAlgorithmsArchivesChromosome StructuresCodeCommunitiesComplexComputational algorithmComputer softwareComputersCoupledDNA SequenceDNA Sequence DatabasesDataData CompressionData SetDatabasesDepositionDiseaseGalaxyGenesGenomeGoalsHealthHumanHuman GeneticsHuman GenomeInfectious Diseases ResearchInvestigationModelingMolecularMutationPositioning AttributeProcessReaction TimeReadingReal-Time SystemsResearchResearch PersonnelResourcesRetrievalRunningScientistSequence AnalysisServicesSiteSolutionsSorting - Cell MovementSpeedSystemTimeValidationVariantWritingbasecloud baseddesigngenome sequencinghuman DNA sequencinghuman diseaseindexinginstrumentmicrobialnext generation sequencingnovel strategiesopen sourceprogramstooltraituser-friendlyweb interface
中文摘要
描述(申请人提供):我们建议创建一个新的系统,Terabase搜索引擎,使生物医学研究人员能够搜索所有已测序并保存在公共档案中的人类DNA序列。人类DNA序列的巨大且不断增长的资源为科学发现和结果验证提供了丰富的机会,但数据集的规模已经远远超过了大多数研究人员使用它们的能力。二十多年来,遗传学家和遗传学家一直依赖DNA序列数据库进行广泛的科学工作,包括发现新基因和新突变,调查物种内部和物种之间的进化变化,影响染色体结构和变化的作用力,以及许多其他分子和进化过程。使用BLAST和类似程序搜索所有已知基因和基因组的能力一直被认为是有能力的,世界各地的序列搜索引擎都提供这种能力。然而,从下一代测序(NGS)项目中涌出的原始数据已经超出了我们提供快速访问它的能力。一台NGS仪器可以在一次运行中产生60亿个读数,包括6000亿个碱基,而且这一能力仍在增长。像BLAST这样的传统比对程序不能在合理的时间内对这些数据进行排序。更新、更快的程序,如Bowtie(由我们团队开发),可以更快地将NGS读数与基因组进行比对,但今天数据集的大小--现在超过1万亿次读数--远远超过大多数计算机存储它的能力。和
即使是当今最快的比对程序也不能在合理的时间内搜索所有这些数据。为了向研究界提供这些巨大而极具价值的DNA序列,需要一种新的方法。Terabase搜索引擎将是一个新的、高效的系统,可以实时搜索数万亿个数据库。使用分级搜索策略和广泛的预处理来加快响应时间,TSE将允许科学家将任何序列,无论是人类的还是非人类的,与所有公开可用的人类序列读数进行比对。与人类基因组匹配的读数将被索引并存储在非常高速的磁盘上,以便快速检索。与微生物序列匹配的读数将被捕获并单独存储,用于微生物组和传染病研究。该系统将通过用户友好的网络界面提供,本地数据库将存储每个用户的结果,以便在东京证交所网站上进一步分析或下载到当地网站。这个系统将使
这是有史以来第一次,任何科学家都可以将一个序列与完整的人类DNA序列进行比对,并检索匹配的所有东西,而不需要编写特殊用途的程序或使用复杂的基于云的软件界面。这个项目的所有软件都将在开放源码模式下开发,允许其他人不受限制地使用、修改、共享和重新分发代码。
英文摘要
DESCRIPTION (provided by applicant): We propose to create a new system, the Terabase Search Engine that will make it possible for biomedical researchers to search all human DNA sequences that have been sequenced and deposited in public archives. The vast and growing resource of human DNA sequences provides a wealth of opportunities for scientific discovery and for validation of results, but the size of the data sets has already far exceeded the ability o most researchers to use them. For more than two decades, geneticists and geneticists have relied on DNA sequence databases for a wide range of scientific endeavors, including the discovery of new genes and new mutations, the investigation of evolutionary changes within and between species, the forces affecting chromosomal structure and change, and many other molecular and evolutionary processes. The ability to search all known genes and genomes using BLAST and similar programs has long been assumed, and sequence search engines throughout the world provide this ability. However, the raw data pouring out of next-generation sequencing (NGS) projects has exceeded our ability to provide rapid access to it. A single NGS instrument can generate six billion reads encompassing 600 billion bases in a single run, and this capacity is still growing. Traditional alignment programs like BLAST cannot sort through this data in a reasonable amount of time. Newer, faster programs such as Bowtie (developed by our group) allow far faster alignment of NGS reads to the genome, but today the size of the data sets, now in excess of 1 trillion reads, far exceeds the ability of most computers to store it. And
even the fastest alignment programs today could not search all this data in a reasonable amount of time. A new approach is required in order to serve up these huge and hugely valuable DNA sequences to the research community. The Terabase Search Engine will be a new, highly efficient system for searching trillions of bases in real time. Using a hierarchical search strategy with extensive pre-processing to speed up response time, the TSE will allow a scientist to align any sequence, human or non-human, to all publicly-available human sequence reads. Reads that match the human genome will be indexed and stored on very high-speed disks for rapid retrieval. Reads that match microbial sequences will be captured and stored separately for use in micro biome and infectious disease research. The system will be made available through a user-friendly web interface, and a local database will store each user's results for further analysis on the TSE site or for download to a local site. This system will make
it possible, for the first time ever, for any scientist to align a sequence to the complete set of human DNA sequences and to retrieve everything that matches, without the need to write special-purpose programs or to use complex cloud-based software interfaces. All of the software for this project will be developed under an open-source model that will permit others to use, modify, share, and re-distribute the code without restriction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Comprehensive Human Expressed Sequences in Brain (CHESS-BRAIN) and their roles in neuropsychiatric illness
-
批准号:10541887
-
项目类别:
-
资助金额:$61.81万
-
财政年份:2021
-
负责人:Steven L. Salzberg
-
依托单位:
Comprehensive Human Expressed Sequences in Brain (CHESS-BRAIN) and their roles in neuropsychiatric illness
-
批准号:10362615
-
项目类别:
-
资助金额:$55.87万
-
财政年份:2021
-
负责人:Steven L. Salzberg
-
依托单位:
Comprehensive Human Expressed Sequences in Brain (CHESS-BRAIN) and their roles in neuropsychiatric illness
-
批准号:10205617
-
项目类别:
-
资助金额:$44.52万
-
财政年份:2021
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Methods for Microbial and Microbiome Sequence Analysis
-
批准号:10331733
-
项目类别:
-
资助金额:$40.34万
-
财政年份:2019
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Methods for Microbial and Microbiome Sequence Analysis
-
批准号:10550160
-
项目类别:
-
资助金额:$40.34万
-
财政年份:2019
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Methods for Microbial and Microbiome Sequence Analysis
-
批准号:10083744
-
项目类别:
-
资助金额:$40.34万
-
财政年份:2019
-
负责人:Steven L. Salzberg
-
依托单位:
The Terabase Search Engine
-
批准号:8882493
-
项目类别:
-
资助金额:$34.61万
-
财政年份:2014
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Gene Modeling and Genome Sequence Assembly
-
批准号:8329127
-
项目类别:
-
资助金额:$10.9万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Alignment Software for Second-Generation Sequencing
-
批准号:8068060
-
项目类别:
-
资助金额:$70.77万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Alignment Software for Second-Generation Sequencing
-
批准号:8464182
-
项目类别:
-
资助金额:$66.02万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Alignment Software for Second-Generation Sequencing
-
批准号:8296503
-
项目类别:
-
资助金额:$69.38万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Gene Modeling and Genome Sequence Assembly
-
批准号:7864735
-
项目类别:
-
资助金额:$14.1万
-
财政年份:2009
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8314380
-
项目类别:
-
资助金额:$23.66万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:9206164
-
项目类别:
-
资助金额:$24.3万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8637538
-
项目类别:
-
资助金额:$24.3万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:7779520
-
项目类别:
-
资助金额:$27.4万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8034822
-
项目类别:
-
资助金额:$5.22万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8829867
-
项目类别:
-
资助金额:$24.3万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:7427231
-
项目类别:
-
资助金额:$27.68万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:7591226
-
项目类别:
-
资助金额:$27.68万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
海外基金