K-mer indexing for pan-genome reference annotation
K-mer indexing for pan-genome reference annotation
批准号:
10093116
负责人:
Hanlee P Ji
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-02-01 至 2023-01-31
关键词:
AddressAlgorithmsArchitectureBRCA mutationsBiologicalBiomedical ResearchChromosomesClinVarClinicalClinical assessmentsCloud ComputingCodeCollectionCommunitiesComplexDataData SetDatabasesDevelopmentDiploidyDiseaseElementsFoundationsFrequenciesGene FrequencyGenesGenetic AnnotationGenetic CodeGenetic PolymorphismGenetic VariationGenomeGenomicsGoalsHaplotypesHumanHuman BiologyHuman GeneticsHuman GenomeIndividualInfrastructureIntuitionLengthLinkLocationMemoryMetadataMethodsNatureNucleotidesOncogenesPerformancePersonsPhasePopulationPrivacyProcessResearchResearch PersonnelResolutionSamplingSavingsSchemeSequence AnalysisSpeedStructureSystemTimeUpdateVariantWorkbaseclinical applicationclinically relevantcloud basedcommunity engagementcostdata sharingdesignflexibilityfootgenetic variantgenome analysisgenome sciencesgenome sequencinggenomic datahuman diseasehuman reference genomeimprovedindexinginsertion/deletion mutationnext generationnext generation sequencingnovelpan-genomepopulation basedpreservationreference genomeweb portal
中文摘要
点击翻译按钮获取中文摘要
英文摘要
ABSTRACT
The human genome reference sequence is one of the foundations of genome sciences, especially in the context
of next-generation sequencing (NGS) analysis. The reference has enabled discoveries in biomedical research
and been particularly instrumental in human disease gene identification. However, the human genome reference
is limited by its static and linear nature. Specifically, the current reference lacks the featural and contextual
flexibility to represent the breadth of human variation. Important elements of individual genomes are either
missed or incorrectly represented. As a solution that will bridge the next generation of reference assemblies with
population genome sequencing studies, we have developed a K-mer-based indexing approach. This method is
more efficient computationally, provides accurate representation in the context of populations and facilitates the
analysis of diverse human genomes. Our goal is to use this strategy in developing a robust computational
architecture that will encode and annotate large collections of genomes in the context of a pan-genome
reference.
First, we plan to develop a scalable, efficient K-mer representation of a large collection of haplotype/phased
reference genomes, by 1) generating an index of all K-mers in human reference genome GRCh38 in a manner
that can efficiently store variant information as metadata, and then 2) incrementally updating the K-mer index to
include all novel K-mers derived from ongoing population sequencing efforts, while 3) developing schemes for
directly analyzing compressed genomic data.
Second, we plan to apply K-mer representation to genomic analysis by 1) providing the entirety of known
human genetic variation in an aggregated index that is computationally efficient and easy to understand, 2)
developing functions for our pan-genomic index that supports ultra-rapid queries, such as of clinically important
variants, and 3) linking conventional coordinate information to the K-mer metadata in the pan-genome index to
allow annotating genetic variation to a particular genome reference.
Third, we will create an online web portal for the pan-genome, using cloud computing, to maximize the utility
of our approach, to promote community engagement and to enabling contribution from the research community.
We expect that completion of these aims will provide: a scalable computational architecture which incorporates
the continuous addition of variant information without loss of resolution or accuracy;; rapid query speeds that will
remain nearly constant as the database grows;; a universally accessible portal using cloud computing.
This work will help solve the issues of multiple assemblies. It will improve researchers’ ability to understand
the relationship of variants and disease, while also providing great savings over the long-term in infrastructure
and computing costs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
K-mer indexing for pan-genome reference annotation
-
批准号:10793082
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10813237
-
项目类别:
-
资助金额:$7.64万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Integrating cancer genomics and spatial architecture of tumor infiltrating lymphocytes
-
批准号:10637960
-
项目类别:
-
资助金额:$44.52万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Single cell modeling of cancer mutations
-
批准号:10612689
-
项目类别:
-
资助金额:$37.53万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Project 1 - Molecular and Cellular Determinants of High Risk Gastric Precancerous Lesions
-
批准号:10715762
-
项目类别:
-
资助金额:$36.89万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Core A: Administrative
-
批准号:10715765
-
项目类别:
-
资助金额:$15.47万
-
财政年份:2023
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10706493
-
项目类别:
-
资助金额:$34.77万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10272359
-
项目类别:
-
资助金额:$34.58万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Multimodal iterative sequencing of cancer genomes and single tumor cells
-
批准号:10363694
-
项目类别:
-
资助金额:$37.66万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Multimodal iterative sequencing of cancer genomes and single tumor cells
-
批准号:10112576
-
项目类别:
-
资助金额:$36.6万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Multimodal iterative sequencing of cancer genomes and single tumor cells
-
批准号:10576304
-
项目类别:
-
资助金额:$37.13万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
Determine the mechanisms of acquired brain-tropism
-
批准号:10927525
-
项目类别:
-
资助金额:$14.42万
-
财政年份:2021
-
负责人:Hanlee P Ji
-
依托单位:
K-mer indexing for pan-genome reference annotation
-
批准号:9905108
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2020
-
负责人:Hanlee P Ji
-
依托单位:
K-mer indexing for pan-genome reference annotation
-
批准号:10328233
-
项目类别:
-
资助金额:$30.0万
-
财政年份:2020
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:8495566
-
项目类别:
-
资助金额:$94.06万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Oligonucleotide-Selective Sequencing for integrated and rapid cancer genome analy
-
批准号:8472073
-
项目类别:
-
资助金额:$35.85万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:8658063
-
项目类别:
-
资助金额:$89.35万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:8856176
-
项目类别:
-
资助金额:$85.46万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Oligonucleotide-Selective Sequencing for integrated and rapid cancer genome analy
-
批准号:8655833
-
项目类别:
-
资助金额:$37.46万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
Functional Analysis of Oncogenic Networks in Primary Organoids
-
批准号:9062864
-
项目类别:
-
资助金额:$90.98万
-
财政年份:2013
-
负责人:Hanlee P Ji
-
依托单位:
海外基金