课题基金 / 基金详情

Statistical Methods for Next-Gen Sequencing in Disease Association Studies

Statistical Methods for Next-Gen Sequencing in Disease Association Studies
疾病关联研究中下一代测序的统计方法
批准号:
7853195
负责人:
Eden R. Martin
金额:
$50.0万
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-30 至 2011-07-31

项目摘要

项目成果

Eden R. Martin的其他基金

相似基金

相关文献

中文摘要
翻译
疾病关联研究中下一代测序的统计方法 通过这个项目,我们建议开发用于基因型识别的统计方法和软件, 下一代序列数据中的关联测试。该领域是由分子的进步,允许 经济实惠的大规模平行测序下一代统计方法的快速发展 疾病研究中的序列数据是跟上先进分子技术步伐所必需的。下一篇: 代测序基于随机、短读技术;因此任何核苷酸的覆盖范围都是 易变且容易出错。区分随机误差与真正可变的位点是“SNP-1”所必需的。 呼唤”。除此之外的一个步骤是确定个体在该位点的实际基因型。这是一个高度 统计问题,我们还没有看到这个问题在统计上严格的方式处理。 我们提出的解决方案,以及使我们的方法新颖的原因,假设我们有一个样本, 每个人都有下一代的序列数据。我们预计,测序可能最终取代 用于疾病关联研究的GWAS SNP阵列。虽然这可能是几年后的全基因组 测序,对足够多的人进行单独测序,进行小关联研究, 目标捕捉阵列我们可以利用来自下一代个体样本的信息 序列数据来更准确地估计个体的基因型和位置特异性错误率。我们 方法是在一个似然框架中表达基因型概率和错误率。我们可以使用 标准的统计理论来帮助我们识别基因型。此方法的性能应优于调用 如目前所做的,基于任意过滤器一次针对单个个体的基因型。 这种统计框架的一个明显优点是,基因型调用中的不确定性可以被 直接结合到我们的疾病关联测试中(例如,病例对照和罕见变异分析)。在这 我们将增加关联测试的能力,减少由于错误或系统性缺失而导致的偏差。 将下一代序列数据合并到关联测试中提供了完整的分析管道 从序列到关联。
英文摘要
Statistical Methods for Next-Generation Sequencing in Disease Association Studies Through this project we propose to develop statistical approaches and software for genotype calling and association testing in next-generation sequence data. The field is driven by molecular advances that allow for affordable, massively parallel sequencing. The rapid development of statistical methods for next-generation sequence data in disease studies is necessary to keep pace with the advancing molecular technology. Next- generation sequencing is based on random, short-read technology; thus the coverage of any nucleotide is highly variable and subject to error. Distinguishing random error from truly variable sites is required for "SNP- calling". One step beyond this is identifying the individual's actual genotype at the site. This is a highly statistical problem and we have yet to see this problem addressed in a statistically rigorous manner. The solution that we propose, and what makes our approach novel, assumes that we have a sample of individuals, each with next-generation sequence data. We anticipate that sequencing may ultimately replace GWAS SNP arrays for disease-association studies. While this may be several years away for whole-genome sequencing, sequencing enough people individually for a small association study is already becoming practical with target capture arrays. We can leverage the information from a sample of individuals with next-generation sequence data to more accurately estimate an individual's genotype and the position-specific error rate. Our approach is to express the genotype probabilities and error rate in a likelihood framework. We can then use standard statistical theory to help us call genotypes. This approach should perform better than calling genotypes for a single individual at a time based on an arbitrary filter as is currently done. A distinct advantage of this statistical framework is that the uncertainty in the genotype calls can be incorporated directly into our disease-association tests (e.g., case-control and rare variant analysis). In this way we will increase power of our association tests and reduce bias due to error or systematic missingness. Incorporation of next-generation sequence data into the association tests provides a complete analysis pipeline from sequence to association.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
GLASS-AD: Global Latinos Sequencing Study for Alzheimer's Disease
Female Sexual Orientation GWAS
Female Sexual Orientation GWAS
Female Sexual Orientation GWAS
海外基金