Discovery and analysis of structural variation in whole genome sequences
Discovery and analysis of structural variation in whole genome sequences
批准号:
9118280
负责人:
RYAN E MILLS
金额:
$38.06万
依托单位国家:
美国
项目类别:
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-09-13 至 2018-07-31
关键词:
AddressAlgorithmsAllelesAreaBenignChromosomal RearrangementClinicalCommunitiesComplexDNA Sequence AlterationDNA Sequence RearrangementDataData SetDatabasesDetectionDiagnosticDiseaseEventFrequenciesFutureGeneticGenetic VariationGenomeGenomicsGenotypeGoalsHealthHuman GenomeIndividualInheritedKaryotype determination procedureLengthMachine LearningMedicalMedical GeneticsMethodologyMethodsModelingNatureOrganismPathogenicityPopulationPublishingReadingRecordsReportingResearchResearch PersonnelResolutionScanningSeedsSourceSpecificityStatistical ModelsStructureSystemTechniquesTechnologyTestingTrainingVariantWorkbaseclinical Diagnosisclinical applicationcohortdirect applicationdisease phenotypegenetic variantgenome sequencinggenomic variationimprovedinterestmarkov modelrare variantstructural genomicstoolvirtualwhole genome
中文摘要
描述(由申请人提供):
大量个体的全基因组测序正迅速成为研究人员调查许多疾病表型遗传基础的常用工具。主要目标是发现导致或促成这些疾病的潜在遗传变异,并在诊断环境中正确识别这些变异。这些差异通常都包括单碱基改变(SNPs),但也可以包括更大、更复杂的结构变异(SV)形式的染色体重排,即使使用现代测序技术也更难检测到。已经发表了许多研究这个问题的方法,但即使是最大规模的努力也只关注删除事件,报告的敏感度为70%。复杂的染色体重排研究更是少之又少。因此,最重要的是开发出准确的方法,从序列数据中以高特异性检测所有类型的SVS。这项建议旨在提高研究人员从全基因组序列中识别和分析遗传变异的整体能力。SV发现的一个重要且经常被忽视的方面是,典型的成对端、读深度和分离读方法将以不同程度的准确度识别不同的非重叠变体集。在目标1中,我们将开发一个统一的SV发现算法,该算法可以以概率的方式整合所有这些不同的信息源。这种方法将有助于研究,特别是识别罕见的变异,以及临床应用,这些应用需要很高的精确度,到目前为止仅限于较旧的核型和微阵列方法。这将识别大多数结构变体,然而基因组序列中有许多区域本质上是复杂的,定义为由多个相邻或重叠的染色体重排组成,用典型的SV检测方法难以解决。在目标2中,我们提出了解决这些复杂区域的方法,并评估了它们的频率和影响。此外,医学遗传学的关键一步是将已识别的基因突变与已知致病和良性变异的数据库进行比较。这在目前的SVS中是有问题的,因为它们最初经常被报告具有不同程度的断点解析,这可能会阻碍对变量的正确分配。这个问题在具有多个断点的更复杂的区域中进一步复杂化,对于这些区域,简单的比较方法不能很好地工作。在目标3中,我们将开发和实现一个系统,该系统描述和利用变异特征来识别个人的序列数据是否包含感兴趣的变体。总体而言,该项目将促进我们对人类基因组的理解,并为一般研究和临床社区提供使用工具。
英文摘要
DESCRIPTION (provided by applicant):
The whole genome sequencing of large cohorts of individuals is quickly becoming a common tool for researchers to investigate the genetic basis of many disease phenotypes. The primary goals are to discover the underlying genetic variation that cause or contribute to these diseases as well as to correctly identify these variants in a diagnostic setting. These differences typicall consist of single base changes (SNPs), but can also encompass larger, more complex chromosomal rearrangements in the form of structural variation (SV) which are much more difficult to detect even with modern sequencing technologies. A number of approaches have been published that have studied this problem, but even the largest scale endeavors have only focused on deletion events and reported a sensitivity of <70%. Complex chromosomal rearrangements are even less well studied. Thus, it is paramount that accurate methods are developed which can detect all types of SVs at high specificity from sequence data. This proposal aims to improve the overall ability of researchers to identify and analyze genetic variation from whole genome sequences. An important, and often overlooked, aspect of SV discovery is the fact that typical paired-end, read depth, and split read approaches will identify different sets of non-overlapping variants at varying degrees of accuracy. In Aim 1, we will develop a unified SV discovery algorithm that can incorporate all of these different sources of information in a probabilistic fashion. Such a method would be useful for research, in particular with the identification of rare variants, as well as clinical applications which require a great del of accuracy and have thus far been limited to older karyotyping and microarray approaches. This would identify the majority of structural variants, however there are many regions in genomic sequences which are complex in nature, defined as consisting of multiple neighboring or overlapping chromosomal rearrangements that are challenging to resolve with typical SV detection approaches. In Aim 2, we propose methods to resolve these complex regions and assess their frequency and impact. Furthermore, a crucial step in medical genetics is the comparison of identified genetic mutations to databases of known pathogenic and benign variants. This is currently problematic with SVs, as they have often been originally reported with varying degrees of breakpoint resolution that can hamper the correct assignment of the variant. This issue is compounded further in more complex regions with multiple breakpoints, for which simplistic comparison methods do not work well. In Aim 3, we will develop and implement a system that describes and utilizes variant profiles to identify whether an individual's sequence data contains a variant of interest. Overall, this project will advance our understanding of the human genome as well as provide tools for use in the general research and clinical communities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Discovery and analysis of structural variation in whole genome sequences
-
批准号:8733748
-
项目类别:
-
资助金额:$37.47万
-
财政年份:2013
-
负责人:RYAN E MILLS
-
依托单位:
Discovery and analysis of structural variation in whole genome sequences
-
批准号:8528145
-
项目类别:
-
资助金额:$38.27万
-
财政年份:2013
-
负责人:RYAN E MILLS
-
依托单位:
Improving INDEL Identification in Genomic Sequences
-
批准号:7222429
-
项目类别:
-
资助金额:$4.4万
-
财政年份:2006
-
负责人:RYAN E MILLS
-
依托单位:
Improving INDEL Identification in Genomic Sequences
-
批准号:7296903
-
项目类别:
-
资助金额:$4.6万
-
财政年份:2006
-
负责人:RYAN E MILLS
-
依托单位:
Improving INDEL Identification in Genomic Sequences
-
批准号:7488007
-
项目类别:
-
资助金额:$0.91万
-
财政年份:2006
-
负责人:RYAN E MILLS
-
依托单位:
海外基金