Computational Methods for Microbial and Microbiome Sequence Analysis
Computational Methods for Microbial and Microbiome Sequence Analysis
批准号:
10331733
负责人:
Steven L. Salzberg
金额:
$40.34万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-02-01 至 2024-01-31
关键词:
AddressAdoptedAlgorithmsAnthrax diseaseArchaeaAreaBacteriaBacterial GenomeCholeraComputer softwareComputing MethodologiesCustomDataData SetDevelopmentDiagnosisEnvironmentEukaryotaExperimental DesignsExplosionGenesGenomeHuman MicrobiomeInfectionLicensingLyme DiseaseMetagenomicsMethodsOrganismPathogenicityPerformanceProblem SolvingPublicationsPublishingScientistSequence AnalysisShotgun SequencingSyphilisSystemTechnologyThird Generation SequencingTimeTuberculosisUpdateVirusWorkbacterial genome sequencingcost effectiveexperimental studygenome databaseimprovedmicrobialmicrobiomemicroorganismnext generation sequencingopen sourcetoolwhole genome
中文摘要
项目摘要
这个项目将支持我们在微生物序列分析的计算方法方面的工作,包括基因
发现、全基因组比对、基因组组装和宏基因组序列分析。多年来我们
已经开发了多种系统来解决这些领域的问题,其中一些系统被广泛使用。这些
工具需要不断更新和改进,以跟上测序技术的变化,
以及不断增长的基因组测序数量。其中一个系统是微光,
在细菌、病毒、古生菌和简单的真核生物中寻找基因的计算方法。格丽默群岛
高准确性,发现超过99%的基因在大多数原核基因组。它已经被成千上万的
世界各地的科学家和过去大多数已发表的细菌基因组测序项目中,
十年描述Glimmer的三种主要出版物总共被引用了4,700多次,
仅在2016-17年就有超过700次引用。近年来,Glimmer的使用有所增加,
下一代测序项目的爆炸式增长,这对细菌基因组尤其具有成本效益。一
第二个系统MUMmer是一个高效的全基因组比对器,用于基因组之间的比较
并比较基因组组合以检测大小变化。MUMmer及其组件,
特别是Nucmer,已经被广泛使用并并入其他系统,包括多基因组比对器
和几个基因组组装包。已经引用了三个描述MUMmer的主要出版物
超过3,600次,包括2016-17年的750次引用。近年来,我们把工作重点放在
开发分析宏基因组数据的方法,生产几种新的工具,包括Kraken
还有阿吉。这两种系统都试图为宏基因组学中的每个读段分配物种标识符
数据集由于Kraken算法不仅准确,而且比早期的方法快得多,它很快就被
在它发布后不久就被许多实验室采用,并且它的使用量还在继续增长。更新更大的空间-
一个高效的数据采集系统也取得了很大的成功,最近被纳入了分析
新的第三代测序公司之一的包。我们将继续努力,
这两种算法的性能,这个项目将允许我们扩展它们来处理最新的长读
这些数据越来越多地被用于宏基因组学实验。最后,实验室的一个新方向是使用
宏基因组鸟枪测序来诊断感染,我们不仅修改了我们的
算法,但也建立定制的基因组数据库,我们严格筛选基因组,以确定
并去除产生假阳性的污染物和低复杂性序列。正如我们为许多人所做的那样,
在未来的几年里,我们将在一个开源的平台上免费发布这个项目产生的所有软件和数据。
许可证,允许其他科学家使用,修改和重新发布它们,而不受任何限制。
英文摘要
Project Summary
This project will support our work on computational methods for microbial sequence analysis, including gene
finding, whole-genome alignment, genome assembly, and metagenomic sequence analysis. Over the years we
have developed multiple systems to solve problems in these areas, some of which are very widely used. These
tools need continued updates and improvements to keep pace with changes in sequencing technology, changes
in experimental design, and the ever-growing number of sequenced genomes. One of these systems is Glimmer,
a computational method for finding genes in bacteria, viruses, archaea, and simple eukaryotes. Glimmer is
highly accurate, finding over 99% of the genes in most prokaryotic genomes. It has been used by thousands of
scientists around the world and in the majority of published bacterial genome sequencing projects over the past
decade. Collectively the three main publications describing Glimmer have been cited over 4,700 times,
including >700 citations in 2016-17 alone. Usage of Glimmer has been increased in recent years due to the
explosion in next-generation sequencing projects, which are particularly cost-effective for bacterial genomes. A
second system, MUMmer, is an efficient whole-genome aligner that is used to compare genomes to one another
and to compare genome assemblies to detect changes, both large and small. MUMmer and its components,
especially Nucmer, have been widely used and incorporated in other systems, including multi-genome aligners
and several genome assembly packages. The three main publications describing MUMmer have been cited
over 3,600 times including >750 citations in 2016-17. In recent years we have focused our efforts on
developing methods for the analysis of metagenomics data, producing several newer tools, including Kraken
and Centrifuge. Both of these systems attempt to assign a species identifier to every read in a metagenomics
data set. Because the Kraken algorithm is not only accurate but far faster than earlier methods, it was rapidly
adopted by many labs soon after its release, and its usage continues to grow. The even newer and more space-
efficient Centrifuge system has also been highly successful and was recently incorporated into the analysis
package of one of the new third-generation sequencing companies. We continue to work on improving the
performance of both algorithms, and this project will allow us to extend them to handle the newest long-read
data that is increasingly being used for metagenomics experiments. Finally, a new direction of the lab is the use
of metagenomic shotgun sequencing to diagnose infections, for which we are not only modifying our
algorithms, but also building customized genome databases where we rigorously screen the genomes to identify
and remove contaminants and low-complexity sequences that create false positives. As we have done for many
years, we will release all of the software and data generated by this project for free under an open source
license, allowing other scientists to use, modify, and redistribute them without restrictions of any kind.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Comprehensive Human Expressed Sequences in Brain (CHESS-BRAIN) and their roles in neuropsychiatric illness
-
批准号:10541887
-
项目类别:
-
资助金额:$61.81万
-
财政年份:2021
-
负责人:Steven L. Salzberg
-
依托单位:
Comprehensive Human Expressed Sequences in Brain (CHESS-BRAIN) and their roles in neuropsychiatric illness
-
批准号:10362615
-
项目类别:
-
资助金额:$55.87万
-
财政年份:2021
-
负责人:Steven L. Salzberg
-
依托单位:
Comprehensive Human Expressed Sequences in Brain (CHESS-BRAIN) and their roles in neuropsychiatric illness
-
批准号:10205617
-
项目类别:
-
资助金额:$44.52万
-
财政年份:2021
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Methods for Microbial and Microbiome Sequence Analysis
-
批准号:10550160
-
项目类别:
-
资助金额:$40.34万
-
财政年份:2019
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Methods for Microbial and Microbiome Sequence Analysis
-
批准号:10083744
-
项目类别:
-
资助金额:$40.34万
-
财政年份:2019
-
负责人:Steven L. Salzberg
-
依托单位:
The Terabase Search Engine
-
批准号:8882493
-
项目类别:
-
资助金额:$34.61万
-
财政年份:2014
-
负责人:Steven L. Salzberg
-
依托单位:
The Terabase Search Engine
-
批准号:8688406
-
项目类别:
-
资助金额:$35.5万
-
财政年份:2014
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Gene Modeling and Genome Sequence Assembly
-
批准号:8329127
-
项目类别:
-
资助金额:$10.9万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Alignment Software for Second-Generation Sequencing
-
批准号:8068060
-
项目类别:
-
资助金额:$70.77万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Alignment Software for Second-Generation Sequencing
-
批准号:8464182
-
项目类别:
-
资助金额:$66.02万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Alignment Software for Second-Generation Sequencing
-
批准号:8296503
-
项目类别:
-
资助金额:$69.38万
-
财政年份:2011
-
负责人:Steven L. Salzberg
-
依托单位:
Computational Gene Modeling and Genome Sequence Assembly
-
批准号:7864735
-
项目类别:
-
资助金额:$14.1万
-
财政年份:2009
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8314380
-
项目类别:
-
资助金额:$23.66万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8637538
-
项目类别:
-
资助金额:$24.3万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:9206164
-
项目类别:
-
资助金额:$24.3万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:7779520
-
项目类别:
-
资助金额:$27.4万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8034822
-
项目类别:
-
资助金额:$5.22万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:8829867
-
项目类别:
-
资助金额:$24.3万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:7427231
-
项目类别:
-
资助金额:$27.68万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
Bioinformatics Software for Analyzing Microbial Genomes
-
批准号:7591226
-
项目类别:
-
资助金额:$27.68万
-
财政年份:2008
-
负责人:Steven L. Salzberg
-
依托单位:
海外基金