Revealing new short tandem repeat variation in the human population across sequencing technologies: towards rare disease diagnosis and discovery
Revealing new short tandem repeat variation in the human population across sequencing technologies: towards rare disease diagnosis and discovery
批准号:
10572951
负责人:
Harriet Dashnow
金额:
$18.54万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-01 至 2025-01-31
关键词:
AddressAllelesCatalogsChildCommunitiesDNA SequenceDataData SetDatabasesDetectionDiagnosisDiseaseDropoutExclusionFrequenciesFundingFutureGene ExpressionGene FrequencyGene OrderGenetic DiseasesGenomeGenomicsGenotypeHospitalsHumanHuman GenomeIncidenceIndividualInvestigationMedical ResearchMethodsMutationNational Human Genome Research InstitutePathogenicityPatientsPhenotypePopulationPopulation HeterogeneityRare DiseasesRepetitive SequenceResearchResearch PersonnelResourcesSchemeShort Tandem RepeatSingle Nucleotide PolymorphismStretchingStutteringTechnologyTrans-Omics for Precision MedicineUnderserved PopulationUniversitiesUntranslated RNAVariantWashingtonWorkcohortcontigdetection methoddisease diagnosisempowermentexperiencegenetic disorder diagnosisgenome sequencinggenome-widehuman diseaseinnovationinsertion/deletion mutationnanoporenovelrare genetic disorderwhole genome
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Abstract
Short tandem repeats (STRs) are 1–6 bp repetitive and highly polymorphic DNA sequences. Expansions in
dozens of STRs are associated with genetic disease. However, STRs are challenging to sequence and interpret,
meaning that individuals with STR disease often go undiagnosed. In rare disease studies, it is now standard to
prioritize candidate pathogenic SNVs, indels and SVs by excluding variants that have high allele frequencies in
population-scale databases such as gnomAD. However, there is no such genome-wide database available for
large STR expansions. I will produce a publicly available STR variation community resource, stratified by
ancestry, to enable prioritization of candidate pathogenic STR expansions.
Long-read sequencing technologies from PacBio and Nanopore have been heralded as the solution to accurately
genotype long repeats because their reads can span the repetitive region. However, there are several challenges
when genotyping STRS in long-reads that are not adequately addressed by existing approaches. I will develop
a method to genotype STRs from long-read Oxford Nanopore sequencing data. It will discover informative reads
using a combination of alignment and identifying repetitive regions in reads. It will then infer the genotype by
integrating evidence from multiple reads, informed by my investigation of biases in these technologies.
Drawing together new short and long-read computational approaches to calling STR expansions, and my
population-scale STR catalog, with an emphasis on diverse and under-served populations, this proposal will
establish a genetic diagnosis for hundreds of patients, while searching for new STR disease loci. I will analyze
patient cohorts enriched for phenotypes associated with STRs from the UDN, University of Washington, Harry
Perkins Institute of Medical Research and Children’s Mercy Hospital to solve cases and discover new disease-
associated STRs in both short and long-read sequencing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金