课题基金 / 基金详情

项目摘要

项目成果

Vineet Bafna的其他基金

相似基金

相关文献

中文摘要
翻译
项目总结 全基因组测序(WGS)有可能描绘出所有临床相关的遗传变异 同时。然而,临床变异发现管道主要集中在编码单核苷酸上 变体(SNV),在较小程度上也适用于监管SNV和小型Indels,忽略了更复杂的 致病变异,如重复或结构重排。 重复可以采取多种形式,但我们考虑三类重复:短串联重复(STR), 可变数目串联重复序列(VnTRs)和低拷贝重复序列或节段性复制 占人类基因组的8%以上。这些不同的类被牵连到许多 孟德尔式的疾病。30多种疾病,主要是神经退行性疾病,是由STR扩张引起的, 包括亨廷顿病、脆性X综合征、ALS/FTD和遗传性共济失调。同样,车主们也有 与一系列精神和其他特征有关,包括髓样囊性肾病和1型 糖尿病。在许多情况下,疾病的进展与生殖系重复计数有关,但序列 个体重复单位内的变异和重复长度的体细胞不稳定性也已被证明是 在某些情况下致病。最后,100多个重复基因的突变被认为与罕见疾病有关。 孟德尔障碍和癌症,包括林奇综合征的PMS2和听力损失的STRC。已被占用 与这些重复课程相关的疾病加在一起,影响着全球数百万人。 尽管它们与疾病相关,但这些重复类型通常不会出现在序列分析中 由于它们带来了生物信息学方面的挑战,因此需要建立新的管道系统。在过去的几年里,我们和其他人 在开发从短读数中分析临床相关重复序列的方法方面取得了重大进展。 然而,重要的挑战仍然存在,包括对长的、复杂的、不完美的或富含GC的基因分型的能力 重复,以推断临床上相关的体细胞变异,以及现有方法的计算负担。 此外,现有的预测单个SNV或INDELs致病性的框架不适用于 大多数重复,因此需要优先排序方法来预测新的重复变体的影响。 这个项目的目标是使重复分析成为现有孟德尔分析的标准组成部分 变量调用管道。为此,我们将开发新的方法来分析来自Long的重复变异 Reads(目标1),扩展我们现有的短读取方法以考虑更复杂的变体类型(目标2), 并建立一个确定致病重复突变优先顺序的框架(目标3)。
英文摘要
PROJECT SUMMARY Whole genome sequencing (WGS) has the potential to profile all clinically relevant genetic variants simultaneously. However, clinical variant discovery pipelines have focused largely on coding single nucleotide variants (SNVs), and to a lesser extent on regulatory SNVs and small indels, ignoring more complex classes of pathogenic variants such as repeats or structural rearrangements. Repeats can take many forms, but we consider three classes of repeats: short tandem repeats (STRs), variable number tandem repeats (VNTRs), and low-copy repeats or segmental duplications, together accounting for more than 8% of the human genome. These variant classes have been implicated in a number of Mendelian diseases. More than 30 disorders, primarily neurodegenerative, are caused by STR expansions, including Huntington’s Disease, Fragile X Syndrome, ALS/FTD, and hereditary ataxias. Similarly, VNTRs have been implicated in a range of psychiatric and other traits including medullary cystic kidney disease and type 1 diabetes. In many cases, the disease progression is correlated with germline repeat counts, but sequence variation within individual repeat units, and somatic instability of repeat length, has also been shown to be pathogenic in some cases. Finally, mutations in more than 100 duplicated genes have been implicated in rare Mendelian disorders and cancer, including PMS2 in Lynch Syndrome and STRC in hearing loss. Taken together, diseases associated with these repeat classes affect millions of individuals worldwide. Despite their relevance to disease, these repeat types are typically absent from sequence analysis pipelines due to the bioinformatics challenges they present. Over the last several years, we and others have made significant progress in developing methods to analyze clinically relevant repeats from short reads. However, important challenges remain, including the ability to genotype long, complex, imperfect, or GC rich repeats, to infer clinically relevant somatic variation, and the computational burden of existing methods. Further, existing frameworks for predicting the pathogenicity of individual SNVs or indels are not applicable to most repeats, and thus there is a need for prioritization methods to predict the impact of new repeat variants. The goal of this project is to make repeat analysis a standard component of existing Mendelian variant calling pipelines. To this end, we will develop novel methods for profiling repeat variants from long reads (Aim 1), extend our existing methods for short reads to consider more complex variant types (Aim 2), and establish a framework for prioritization of pathogenic repeat mutations (Aim 3).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
eDyNAmiC - UCSD
eDyNAmiC - UCSD
Software and algorithms for elucidating the structure, function, and evolution of extrachromosomal DNA
Graduate Training Program in Bioinformatics
海外基金