Robust Methods for the Efficient Analysis and Integration of DNA Sequence Data
Robust Methods for the Efficient Analysis and Integration of DNA Sequence Data
批准号:
7692191
负责人:
ANDREW S ALLEN
金额:
$23.4万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-26 至 2011-06-30
关键词:
AccountingAddressBase SequenceCommunitiesComplexComputer softwareDNA SequenceDNA Sequence AnalysisDataData SetDevelopmentDiseaseDisease AssociationDisease ProgressionDocumentationEvolutionFutureGeneticGenetic ResearchGenetic VariationGenomeGenotypeGoalsHaplotypesHuman GeneticsIndividualInformation NetworksInternetInvestigationLocalized DiseaseMajor Depressive DisorderMethodologyMethodsPerformancePopulationPopulation GeneticsProceduresProductionPropertyResearchResearch PersonnelResearch Project GrantsRestRoleSamplingScientistSignal TransductionSingle Nucleotide PolymorphismSoftware ToolsSource CodeStatistical MethodsStratificationStructureTestingTrustVariantWorkbasecase controlcostdatabase of Genotypes and Phenotypesfallsfollow-upgenetic associationgenetic variantgenome sequencinggenome wide association studyhuman diseasenovelresearch studyresponsesimulationstatisticstooluser friendly software
中文摘要
描述(由申请人提供):人类遗传学研究正处于如何捕获遗传变异的重大转变的尖端-从基于标记的方法到基于个体基因组测序的完整表征的方法。这是一个令人兴奋的前景,但并非没有挑战。即将产生的大量序列数据提出了如何最好地利用这些数据的几个问题。例如,由于数据的绝对规模,将序列变异与人类疾病联系起来的统计方法需要在统计和计算方面都是有效的。此外,近期的大多数遗传关联实验将不会仅仅依赖于序列数据,而是将有具有序列数据的个体的子样本,而其余样本将保持未测序,但将包含基因型信息。或者,序列数据可以在单独的外部样品上获得。因此,重要的是开发统计方法,将这些不同类型的数据适当地整合到一个统一的推理框架中。本研究项目将通过提出开发一类新的基于序列的单倍型共享统计来解决这些问题,该统计利用DNA序列进化在检测变异/疾病关联中的意义(具体目标1)。此外,我们建议开发一个统计框架,允许对DNA序列和基因型数据进行统一分析(具体目标2)。在整个过程中,我们将利用我们以前的工作,为单倍型推断开发强大的方法,以开发计算和统计有效的程序,这些程序对群体遗传假设保持强大。将强调分层分析方法,以允许调整由于人口分层造成的混淆。有效的蒙特卡罗程序将提出,以说明大量的序列变异调查。我们将开发一套软件工具,完全实现所开发的方法,并使其免费提供给一般研究界(具体目标3)。最后,使用这些工具,我们将分析一个公开可用的DNA序列数据集,目的是更好地定位疾病相关序列变异(具体目标4)。通过该提案开发的方法代表了一个统一的、统计上严谨的框架,用于开发强大的测试,利用DNA序列之间的进化关系,同时允许将不同的数据类型合并到统一的分析中。这些程序将为研究人员提供更精细地定位疾病相关序列变异的工具,从而通过功能研究对变异进行更好的优先排序。人类遗传学研究正处于如何捕获遗传变异的重大转变的风口浪尖——从基于标记的方法到基于测序的个人基因组完整特征的方法。然而,即将产生的大量测序数据导致了有关统计分析和纳入更大实验的问题。为了解决这些问题,我们提出了一个统一的、统计严谨的框架,用于开发强大的测试,利用DNA序列之间的进化关系,并允许将不同的数据类型合并到统一的分析中。
英文摘要
DESCRIPTION (provided by applicant): Human genetics research is on the cusp of a major transformation in how genetic variation is captured-from a marker-based approach to one based on a complete characterization of an individual's genome by sequencing. This is an exciting prospect but not without its challenges. The imminent production of large amounts of sequence data raises several issues on how best to use these data. For example, because of the sheer scale of the data, statistical approaches for associating sequence variants with human disease need to be efficient, both statistically and computationally. In addition, most genetic association experiments in the near term will not rely solely on sequence data but instead will have sub-samples of individuals with sequence data while the rest of the sample will remain unsequenced but will contain genotype information. Alternatively, sequence data may be available on a separate, external sample. Thus it will be important to develop statistical methods that can appropriately integrate these various types of data into a unified inferential framework. This research project will address these issues by proposing to develop a novel class of sequence- based haplotype sharing statistics that exploit the implications of DNA sequence evolution in testing for variant/disease association (specific aim 1). Further, we propose to develop a statistical framework that allows for the unified analysis of DNA sequence and genotype data (specific aim 2). Throughout we will leverage our previous work developing robust methods for haplotype inference to develop computationally and statistically efficient procedures that remain robust to population genetic assumptions. A stratified analytic approach will be emphasized to allow for adjustment for confounding due to population stratification. Efficient Monte Carlo procedures will be proposed to account for the large number of sequence variants investigated. We will develop a suite of software tools that fully implement the methodology developed and make them freely available to the general research community (specific aim 3). Finally, using these tools, we will analyze a publicly available DNA sequence dataset with the goal of better localizing disease- associated sequence variants (specific aim 4). The methods developed through this proposal represent a unified and statistically rigorous framework for developing powerful tests that exploit evolutionary relationships between DNA sequences while allowing for disparate data types to be incorporated into a unified analysis. These procedures will give researchers the tools to more finely localize disease-associated sequence variants, allowing variants to be better prioritized for subsequent investigation via functional studies. Human genetics research is on the cusp of a major transformation in how genetic variation is captured-from a marker based approach to one based on a complete characterization of an individual's genome by sequencing. The imminent production of large amounts of sequencing data, however, leads to questions concerning their statistical analysis and incorporation into the larger experiment. We address these questions by proposing a unified and statistically rigorous framework for developing powerful tests that exploit evolutionary relationships between DNA sequences and that allow for disparate data types to be incorporated into a unified analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Design, prediction, and prioritization of systematic perturbations of the human genome
-
批准号:10665666
-
项目类别:
-
资助金额:$72.98万
-
财政年份:2021
-
负责人:ANDREW S ALLEN
-
依托单位:
Design, prediction, and prioritization of systematic perturbations of the human genome
-
批准号:10473740
-
项目类别:
-
资助金额:$72.98万
-
财政年份:2021
-
负责人:ANDREW S ALLEN
-
依托单位:
Design, prediction, and prioritization of systematic perturbations of the human genome
-
批准号:10295506
-
项目类别:
-
资助金额:$35.37万
-
财政年份:2021
-
负责人:ANDREW S ALLEN
-
依托单位:
The Duke FUNCTION Center: Pioneering the comprehensive identification of combinatorial noncoding causes of disease
-
批准号:10271500
-
项目类别:
-
资助金额:$248.99万
-
财政年份:2020
-
负责人:ANDREW S ALLEN
-
依托单位:
Quantifying the genetic diversity of human regulatory element activity
-
批准号:10404498
-
项目类别:
-
资助金额:$76.48万
-
财政年份:2019
-
负责人:ANDREW S ALLEN
-
依托单位:
Robust Methods for the Efficient Analysis and Integration of DNA Sequence Data
-
批准号:8064557
-
项目类别:
-
资助金额:$20.99万
-
财政年份:2008
-
负责人:ANDREW S ALLEN
-
依托单位:
Robust Methods for the Efficient Analysis and Integration of DNA Sequence Data
-
批准号:7892941
-
项目类别:
-
资助金额:$23.4万
-
财政年份:2008
-
负责人:ANDREW S ALLEN
-
依托单位:
Advanced Haplotype Analyses in Coronary Artery Disease
-
批准号:6934516
-
项目类别:
-
资助金额:$14.21万
-
财政年份:2004
-
负责人:ANDREW S ALLEN
-
依托单位:
Advanced Haplotype Analyses in Coronary Artery Disease
-
批准号:7437286
-
项目类别:
-
资助金额:$14.21万
-
财政年份:2004
-
负责人:ANDREW S ALLEN
-
依托单位:
Advanced Haplotype Analyses in Coronary Artery Disease
-
批准号:7279291
-
项目类别:
-
资助金额:$14.21万
-
财政年份:2004
-
负责人:ANDREW S ALLEN
-
依托单位:
Advanced Haplotype Analyses in Coronary Artery Disease
-
批准号:6815671
-
项目类别:
-
资助金额:$14.21万
-
财政年份:2004
-
负责人:ANDREW S ALLEN
-
依托单位:
Advanced Haplotype Analyses in Coronary Artery Disease
-
批准号:7094069
-
项目类别:
-
资助金额:$14.21万
-
财政年份:2004
-
负责人:ANDREW S ALLEN
-
依托单位:
海外基金