Direct whole-genome haplotype-resolved assembly using sequence graphs
Direct whole-genome haplotype-resolved assembly using sequence graphs
批准号:
10015321
负责人:
Shilpa Garg
金额:
$4.8万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-10 至 2021-01-03
关键词:
AwarenessBioinformaticsBiologyCentromereCharacteristicsChromosomesChurchClinicalCommunitiesComplementComplexComputational BiologyDNA SequenceDataData ScienceData SetDevelopmentDiploidyDoctor of PhilosophyGenerationsGenomeGenomicsGoalsGraphHaplotypesHigh-Throughput Nucleotide SequencingHomologous GeneHumanHuman GenomeHybridsImmunoglobulinsIn SituIndividualInformaticsInheritedInstitutesKiller CellsLinkMajor Histocompatibility ComplexMedicalMedical GeneticsMentorsMetagenomicsParentsPhasePopulationPopulation GeneticsPositioning AttributeProductionReceptor CellRecombinant DNARecording of previous eventsResearchResearch Project GrantsScientistSupervisionTechniquesTechnologyTestingTimeTrainingVariantbasecareerclinically relevantcomputer sciencecomputerized toolscostcost effectivedesigngenetic pedigreehuman genome sequencinghuman reference genomeinnovationinsightmedical schoolsnovelopen sourcepreservationprofessorreceptorreconstructiontoolwhole genome
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Abstract
The lack of complete, high-quality sequencing of human genomes is a major bottleneck for accurate and
complete analyses in population and medical genetics. Advances in a variety of sequencing technologies have
created enormous opportunities to yield full assemblies of every chromosome and its homologue (called as
haplotypes). The reconstruction of haplotype sequences from sequencing data is known as diploid assembly or
haplotype-aware de novo assembly. Standard de novo assemblers are limited in their ability to combine mixed
data types, and also collapse haplotype sequences, resulting in expensive, discontinuous, and inaccurate
assemblies. Our interim goal is a finished human genome that would not only reveal the last remaining regions
of the genome, but also benefit downstream analyses by providing an unbiased reference for comparison and
mapping, as well as the complete phased sequencing of several human and non-human genomes for specific
research projects.
This project will develop a novel computational toolkit WHdenovo, that can optimally combine various
sequencing data types to generate phased assemblies of single individuals and pedigrees. In aim 1 (K99
phase), I will provide computationally efficient tools that are easy-to-use, open-source and are production level
for generating diploid assemblies of pedigrees at minimal cost. In aim 2 (R00 phase), I will develop novel
computational tools for generating pedigree-independent diploid assemblies of single individuals over whole
genomes including centromeres. In aim 3 (R00 phase), the tools developed during aims 1 and 2 will be applied
to generating diploid assemblies of diverse human and non-human genomes, and of clinically relevant regions
such as the histocompatibility complex (MHC) and killer cell immunoglobulin-like receptor (KIR) region. My goal
is to design tools that will be useful to large consortiums such as Genome in a Bottle, High Quality Human
Reference Genomes, and the Personal Genome Project.
My extensive background in computational biology puts me in a unique position to accomplish this proposal,
which requires a seamless integration between data science and genomics. Career and Training: I received
my PhD in Computer Science at Max Planck Institute for Informatics, and started postdoctoral research in the
lab of Professor George Church at Harvard Medical School. During the K99 phase, I will continue to be
mentored by Professor Church. Under the supervision of co-mentor Heng Li, I will advance my expertise in
making computational tools efficient in practice, and how to tune them for upcoming novel high throughput
sequencing (HTS) datasets. This proposed plan would prepare me to be an independent bioinformatics
research scientist.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金