K-mer indexing for pan-genome reference annotation
K-mer indexing for pan-genome reference annotation
批准号:
9905108
负责人:
Hanlee P Ji
金额:
$37.61万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-02-01 至 2023-01-31
关键词:
AddressAlgorithmsArchitectureBRCA1 geneBiologicalBiomedical ResearchChromosomesClinVarClinicalClinical assessmentsCloud ComputingCodeCollectionCommunitiesComplexDataData SetDatabasesDevelopmentDiploidyDiseaseElementsFoundationsFrequenciesGene FrequencyGenesGenetic AnnotationGenetic CodeGenetic PolymorphismGenetic VariationGenomeGenomicsGoalsHaplotypesHumanHuman BiologyHuman GeneticsHuman GenomeIndividualInfrastructureIntuitionLengthLinkLocationMemoryMetadataMethodsMutationNatureNucleotidesOncogenesPerformancePersonsPhasePopulationPrivacyProcessResearchResearch PersonnelResolutionSamplingSavingsSchemeSequence AnalysisSpeedStructureSystemTimeUpdateVariantWorkbaseclinical applicationclinically relevantcloud basedcostdata sharingdesignflexibilityfootgenetic variantgenome analysisgenome sciencesgenome sequencinggenomic datahuman diseasehuman reference genomeimprovedindexinginsertion/deletion mutationnext generationnext generation sequencingnovelpan-genomepopulation basedpreservationreference genomeweb portal
中文摘要
摘要:
人类基因组和参考序列研究是人类基因组学的重要基础研究之一,特别是在人类基因组学的背景下。
下一代基因组测序技术(NGS)的研究分析。这一参考文献已经使我们能够在生物医学和生物医学研究中获得新的发现。
特别是在人类遗传病和基因识别方面发挥了重要作用。然而,人类基因组的研究并不是最重要的。
它的局限性在于它的静态特性和线性特性。具体地说,就是当前的参考模型缺乏最大的特色和更好的语境。
灵活性能够代表人类基因组变异的最大广度。个体基因组的一些重要元素也不是。
遗漏了数据或错误地表示了数据。作为一种有效的解决方案,它将成为连接下一代数据引用和程序集之间的桥梁。
人口、基因组和测序是一种研究,我们可能已经发展出一种基于K-mer的基因索引方法。
在计算上更高效,它在不同人口的背景下提供了更准确的数据表示,并促进了数据的更新。
分析人类基因组的多样性。我们的目标是更好地利用这一战略来开发一种更强大的计算能力。
这一体系结构表示,他们将在构建泛基因组的新背景下,对大量的生物基因组集合进行编码和注释。
参考文献。
首先,我们计划开发一种可扩展的、高效的单倍型/分阶段的大型数据集合体的K-mer表示方法。
参考基因组,由1)以一种简单的方式生成人类所有参考基因组GRCh38中所有K-MERS的索引列表。
这样就可以高效地存储各种不同的信息,如元数据,然后(2)增量地更新K-mer数据索引。
包括从正在进行的人口和测序工作中衍生出来的所有新的K-MERS,同时还包括制定新的K-MERS计划。
直接分析压缩后的基因组数据。
第二,我们计划在1)之前将K-mer表示法应用于基因组DNA分析,从而提供所有已知的信息。
人类的遗传变异出现在一个综合指数中,该指数在计算上非常高效,而且很容易理解。
开发这些功能有助于建立我们最新的泛基因组数据库索引,该索引可以支持一些超快速的数据查询,例如一些具有临床重要性的数据。
变体、序列和序列3)将常规基因组序列与信息流相关联,以将泛基因组序列索引中的元数据添加到基因组序列索引中。
允许对遗传基因变异进行注释,以提供特定的基因组参考。
第三,我们将通过使用云计算,为人类泛基因组创建一个更好的在线研究网络门户网站,以实现人类效用的最大化。
在我们的方法中,我们希望促进社区的参与,并鼓励社区研究和社区的贡献。
我们可以预计,这些目标的完成将提供:一个高度可扩展的计算环境体系结构,其中包括。
不断添加各种不同的数据信息,而不会损失分辨率或准确性;;的快速数据查询速度比它所能做到的更快。
随着云数据库的增长,它几乎保持不变;;是一个全球通用的可访问的门户网站,使用的是云计算。
这项工作将有助于更好地解决多个组件的问题,也将有助于提高研究人员的理解能力。
基因变异与疾病的关系密切,同时也为基础设施的长期发展提供了巨大的成本节约。
以及更高的计算成本。
英文摘要
ABSTRACT
The human genome reference sequence is one of the foundations of genome sciences, especially in the context
of next-generation sequencing (NGS) analysis. The reference has enabled discoveries in biomedical research
and been particularly instrumental in human disease gene identification. However, the human genome reference
is limited by its static and linear nature. Specifically, the current reference lacks the featural and contextual
flexibility to represent the breadth of human variation. Important elements of individual genomes are either
missed or incorrectly represented. As a solution that will bridge the next generation of reference assemblies with
population genome sequencing studies, we have developed a K-mer-based indexing approach. This method is
more efficient computationally, provides accurate representation in the context of populations and facilitates the
analysis of diverse human genomes. Our goal is to use this strategy in developing a robust computational
architecture that will encode and annotate large collections of genomes in the context of a pan-genome
reference.
First, we plan to develop a scalable, efficient K-mer representation of a large collection of haplotype/phased
reference genomes, by 1) generating an index of all K-mers in human reference genome GRCh38 in a manner
that can efficiently store variant information as metadata, and then 2) incrementally updating the K-mer index to
include all novel K-mers derived from ongoing population sequencing efforts, while 3) developing schemes for
directly analyzing compressed genomic data.
Second, we plan to apply K-mer representation to genomic analysis by 1) providing the entirety of known
human genetic variation in an aggregated index that is computationally efficient and easy to understand, 2)
developing functions for our pan-genomic index that supports ultra-rapid queries, such as of clinically important
variants, and 3) linking conventional coordinate information to the K-mer metadata in the pan-genome index to
allow annotating genetic variation to a particular genome reference.
Third, we will create an online web portal for the pan-genome, using cloud computing, to maximize the utility
of our approach, to promote community engagement and to enabling contribution from the research community.
We expect that completion of these aims will provide: a scalable computational architecture which incorporates
the continuous addition of variant information without loss of resolution or accuracy;; rapid query speeds that will
remain nearly constant as the database grows;; a universally accessible portal using cloud computing.
This work will help solve the issues of multiple assemblies. It will improve researchers’ ability to understand
the relationship of variants and disease, while also providing great savings over the long-term in infrastructure
and computing costs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
K-mer indexing for pan-genome reference annotation
-
批准号:10793082
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10813237
-
项目类别:
-
资助金额:$7.64万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Integrating cancer genomics and spatial architecture of tumor infiltrating lymphocytes
-
批准号:10637960
-
项目类别:
-
资助金额:$44.52万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Single cell modeling of cancer mutations
-
批准号:10612689
-
项目类别:
-
资助金额:$37.53万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Project 1 - Molecular and Cellular Determinants of High Risk Gastric Precancerous Lesions
-
批准号:10715762
-
项目类别:
-
资助金额:$36.89万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Core A: Administrative
-
批准号:10715765
-
项目类别:
-
资助金额:$15.47万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10706493
-
项目类别:
-
资助金额:$34.77万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10272359
-
项目类别:
-
资助金额:$34.58万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Multimodal iterative sequencing of cancer genomes and single tumor cells
-
批准号:10363694
-
项目类别:
-
资助金额:$37.66万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Multimodal iterative sequencing of cancer genomes and single tumor cells
-
批准号:10112576
-
项目类别:
-
资助金额:$36.6万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Multimodal iterative sequencing of cancer genomes and single tumor cells
-
批准号:10576304
-
项目类别:
-
资助金额:$37.13万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10927525
-
项目类别:
-
资助金额:$14.42万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
K-mer indexing for pan-genome reference annotation
-
批准号:10093116
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2020
-
负责人:Hanlee P Ji
-
依托单位:
K-mer indexing for pan-genome reference annotation
-
批准号:10328233
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2020
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:8495566
-
项目类别:
-
资助金额:$94.06万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Oligonucleotide-Selective Sequencing for integrated and rapid cancer genome analy
-
批准号:8472073
-
项目类别:
-
资助金额:$35.85万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:8658063
-
项目类别:
-
资助金额:$89.35万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:8856176
-
项目类别:
-
资助金额:$85.46万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Oligonucleotide-Selective Sequencing for integrated and rapid cancer genome analy
-
批准号:8655833
-
项目类别:
-
资助金额:$37.46万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:9062864
-
项目类别:
-
资助金额:$90.98万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
海外基金