课题基金 / 基金详情

项目摘要

项目成果

DAVID M UMBACH的其他基金

相似基金

相关文献

中文摘要
翻译
我们正在开发方法,使用从几个人的集合样本中测量的暴露,以及每个人单独测量的基因类型,来研究基因与环境的相互作用。假设一个人进行了病例对照研究,并在一组SNPs(单核苷酸多态)上对每个个体进行了基因分型。假设一个人也有来自相同个人的生物标本(例如,血清或尿液),但缺乏对每个个体样本进行感兴趣暴露的检测的预算。汇集标本并对汇集的标本进行分析,不仅可以节省化验成本,还可以保存标本体积以备将来使用。在过去,我们开发了分析病例对照研究的方法,这些研究的暴露是在混合样本中测量的。这些方法合理地假设集合样本上的测量值是单个样本的值的平均值。使用这些方法,在单个SNP上测试基因与环境的相互作用需要在具有相同SNP基因的个体的层级内创建样本库。为了研究一组SNPs的基因-环境相互作用,我们以前的方法需要为每个研究的SNP创建新的池样本,潜在的分析成本节省将消失。我们正在开发的方法将单个测量结果视为丢失的数据,并以原则性的方式使用汇集的样本来推算这些丢失的数据。有了一组给定的输入数据,我们就可以使用病例对照数据的标准统计方法来估计基因与环境的交互作用。在实践中,我们使用多重归因法:创建多个归因组数据,对每个组进行病例对照分析,并将多重分析的结果结合起来。这一方法显示了一些希望,但仍有一些问题有待解决。解决这个问题的工作正在进行中。 当单个SNPs的边际效应很小时,在全基因组研究中识别致病SNPs可能是具有挑战性的,因为测试阈值必须反映研究中的大量SNPs。对于复杂的疾病,特定的SNPs组合可能会显著增加一种上位性或基因-基因交互作用的风险。我们目前正在研究使用机器学习技术在病例父母数据中发现一组共同导致疾病的SNPs(致病SNPs)。首先,我们设计了一种方法,使用实际的病例-亲本三联体基因类型来创建模拟的全基因组数据集,该数据集反映了现实的连锁不平衡结构,并以已知的致病SNPs集为种子。这份手稿正在审查中,计算机代码已经公开。我们目前正在努力更好地描述以这种方式模拟的种群的遗传特性。第二,我们实现了一个现有的随机搜索算法(称为GA-KNN),它基于进化算法来寻找多组k个SNP,这些SNP可以预测疾病(这里的k是一个小数字,比如2或4)。通过对那些在预测疾病的集合中出现频率最高的SNPs进行分类,我们希望发现导致疾病的SNPs集合。在模拟数据的初步试验中,我们的方法显示了希望。在正在进行的工作中,我们正在尝试加快算法的速度,并看看在更复杂的情况下是否能保持有希望的性能。 (另见Z01 ES040007;Pi Clare Weinberg;施敏也是这个项目的实验室内合作者;她的时间分配在Weinberg的项目中,但不在这个项目中。)
英文摘要
We are in the process of developing methods for using exposures measured in pooled specimens from several individuals, together with genotypes measured separately on each individual, to study gene-environment interactions. Suppose one has case-control study and genotyped each individual at a panel of SNPs (single nucleotide polymorphisms). Suppose that one also has biological specimens (e.g., serum or urine) from the same individuals but lacks the budget to assay each individual specimen for an exposure of interest. Pooling specimens and assaying the resulting pooled specimens will not only save assay costs but preserve specimen volume for future uses. In the past, we have developed methods for analyzing case-control studies with exposures measured in pooled specimens. Those methods assume, reasonably, that the measured value on the pooled specimen is the average of the values for the individual specimens. With those methods, testing gene-environment interactions at a single SNP required creating specimen pools within strata of individuals who all had the same genotype for that SNP. To study gene-environment interactions for a panel of SNPs, our previous methods would require creating new pooled specimens for each SNP studied and the potential savings in assay costs would disappear. The approach that we are developing regards the individual measurements as missing data and uses the pooled specimens in a principled way to impute those missing data. With a give set of imputed data in hand, we can use standard statistical methods for case-control data to estimate gene-environment interactions. In practice, we use a multiple-imputation approach: creating multiple sets of imputed data, doing a case-control analysis for each set, and combining the results from the multiple analyses. This approach has shown some promise but some problems remain to be resolved. Work on this problem is ongoing. Identification of causative SNPs in a genome-wide study can be challenging when individual SNPs have small marginal effects because testing thresholds must reflect the large number of SNPs under study. For complex diseases, particular combinations of SNPs may dramatically increase risk a kind of epistasis or gene-gene interaction. We are currently investigating the use of a machine learning technique for the discovery of sets of SNPs that together cause disease (causative SNPs) in case-parents data. First, we devised a way to use actual case-parent triad genotypes to create simulated genome-wide data sets that reflect realistic linkage disequilibrium structure and are seeded with known sets of causative SNPs. This manuscript is under review, and the computer code is publicly available. We are currently working to better characterize the genetic properties of populations simulated in this way. Second, we implemented an existing stochastic search algorithm (called GA-KNN) that is based on an evolutionary algorithm to find multiple sets of k SNPs that are predictive of disease (here k is a small number, say 2 or 4). By cataloguing those SNPs which appear most frequently among the sets that are predictive of disease, we hope to uncover the sets of causative SNPS. In preliminary trials on simulated data seeded with two interacting sets of four SNPs each, our approach shows promise. In ongoing work, we are attempting to speed up the algorithm and to see whether the promising performance is maintained in more complex situations. (see also Z01 ES040007; PI Clare Weinberg; Min Shi is also a within-lab collaborator on this project; her time is allocated in Weinberg's project but not in this one.)
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
STATISTICAL METHODS FOR MISMEASURED OR MISSING DATA
STATISTICAL METHODS IN HUMAN DEVELOPMENT/CLINICAL STUDIES
Statistical Methods In Human Development/Clinical Study
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
海外基金