Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
批准号:
9352092
负责人:
DAVID M UMBACH
金额:
$3.14万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AlgorithmsAllelesBiologicalBiological AssayBudgetsCase-Control StudiesCatalogingCatalogsComplexControl GroupsDataData AnalysesData SetDiseaseEnvironmentFutureGene FrequencyGenesGeneticGenetic EpistasisGenetic Predisposition to DiseaseGenotypeHandIndividualInvestigationLinkage DisequilibriumMachine LearningMeasurementMeasuresMethodsModificationNational Institute of Environmental Health SciencesParentsPerformancePopulationPredispositionProcessPropertyResearch PersonnelRiskSavingsSerumSiblingsSingle Nucleotide PolymorphismSisterSpecimenSpeedStatistical MethodsStructureTechniquesTestingTimeTriad Acrylic ResinUrineWorkbasecase controlcopingcostdesigndisorder riskgene environment interactiongene interactiongenetic makeupgenetic variantgenome-widegenome-wide analysisinterestmalignant breast neoplasmmethod developmenttool
中文摘要
我们正在开发一种方法,利用从几个个体收集的标本中测量的暴露量,以及在每个个体上单独测量的基因型,来研究基因与环境的相互作用。假设有一个病例对照研究,并在单核苷酸多态性面板上对每个个体进行基因分型。假设也有来自同一个体的生物标本(如血清或尿液),但缺乏预算对每个个体标本进行检测以获得感兴趣的暴露。汇集标本并分析由此产生的汇集标本不仅可以节省分析成本,还可以保存标本体积以供将来使用。在过去,我们已经开发了分析病例对照研究的方法,在汇集的标本中测量暴露。这些方法合理地假设,混合标本上的测量值是单个标本值的平均值。使用这些方法,在单个SNP上测试基因-环境相互作用需要在具有相同SNP基因型的个体地层中创建标本池。为了研究一组SNP的基因-环境相互作用,我们以前的方法需要为研究的每个SNP创建新的汇集标本,并且在分析成本方面的潜在节省将消失。我们正在开发的方法将单个测量视为缺失的数据,并以一种有原则的方式使用汇集的标本来估算这些缺失的数据。有了一组给定的输入数据,我们可以使用病例对照数据的标准统计方法来估计基因与环境的相互作用。在实践中,我们使用多重输入方法:创建多组输入数据,对每组数据进行病例对照分析,并将多个分析的结果结合起来。这种方法显示出一些希望,但仍有一些问题有待解决。解决这个问题的工作正在进行中。
英文摘要
We are in the process of developing methods for using exposures measured in pooled specimens from several individuals, together with genotypes measured separately on each individual, to study gene-environment interactions. Suppose one has case-control study and genotyped each individual at a panel of SNPs (single nucleotide polymorphisms). Suppose that one also has biological specimens (e.g., serum or urine) from the same individuals but lacks the budget to assay each individual specimen for an exposure of interest. Pooling specimens and assaying the resulting pooled specimens will not only save assay costs but preserve specimen volume for future uses. In the past, we have developed methods for analyzing case-control studies with exposures measured in pooled specimens. Those methods assume, reasonably, that the measured value on the pooled specimen is the average of the values for the individual specimens. With those methods, testing gene-environment interactions at a single SNP required creating specimen pools within strata of individuals who all had the same genotype for that SNP. To study gene-environment interactions for a panel of SNPs, our previous methods would require creating new pooled specimens for each SNP studied and the potential savings in assay costs would disappear. The approach that we are developing regards the individual measurements as missing data and uses the pooled specimens in a principled way to impute those missing data. With a give set of imputed data in hand, we can use standard statistical methods for case-control data to estimate gene-environment interactions. In practice, we use a multiple-imputation approach: creating multiple sets of imputed data, doing a case-control analysis for each set, and combining the results from the multiple analyses. This approach has shown some promise but some problems remain to be resolved. Work on this problem is ongoing.
Identification of causative SNPs in a genome-wide study can be challenging when individual SNPs have small marginal effects because testing thresholds must reflect the large number of SNPs under study. For complex diseases, particular combinations of SNPs may dramatically increase risk a kind of epistasis or gene-gene interaction. We are currently investigating the use of a machine learning technique for the discovery of sets of SNPs that together cause disease (causative SNPs) in case-parents data. First, we devised a way to use actual case-parent triad genotypes to create simulated genome-wide data sets that reflect realistic linkage disequilibrium structure and are seeded with known sets of causative SNPs. We are currently working to better characterize the genetic properties of populations simulated in this way. Second, we implemented an existing stochastic search algorithm (called GA-KNN) that is based on an evolutionary algorithm to find multiple sets of k SNPs that are predictive of disease (here k is a small number, say 2 or 4). By cataloguing those SNPs which appear most frequently among the sets that are predictive of disease, we hope to uncover the sets of causative SNPS. In preliminary trials on simulated data seeded with two interacting sets of four SNPs each, our approach shows promise. In ongoing work, we are attempting to speed up the algorithm and to see whether the promising performance is maintained in more complex situations.
(see also Z01 ES040007; PI Clare Weinberg; Min Shi is also a within-lab collaborator on this project; her time is allocated in Weinberg's project but not in this one.)
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
STATISTICAL METHODS FOR MISMEASURED OR MISSING DATA
-
批准号:6106658
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
STATISTICAL METHODS IN HUMAN DEVELOPMENT/CLINICAL STUDIES
-
批准号:6106670
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods In Human Development/Clinical Study
-
批准号:6507351
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
-
批准号:8336546
-
项目类别:
-
资助金额:$5.58万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods In Human Development/clinical Studie
-
批准号:7007381
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And
-
批准号:7168878
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
-
批准号:8734072
-
项目类别:
-
资助金额:$2.71万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
-
批准号:8553699
-
项目类别:
-
资助金额:$4.86万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
-
批准号:9550027
-
项目类别:
-
资助金额:$3.34万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods In Human Development/clinical Studies
-
批准号:9143429
-
项目类别:
-
资助金额:$11.99万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods In Human Development/clinical Studie
-
批准号:6672942
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods In Human Development/clinical Studies
-
批准号:8553706
-
项目类别:
-
资助金额:$9.72万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And Genetic Susceptibility
-
批准号:8149006
-
项目类别:
-
资助金额:$6.41万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods In Human Development/clinical Studie
-
批准号:7169654
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
STATISTICAL METHODS FOR MISMEASURED OR MISSING DATA
-
批准号:6289962
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/environment Interaction And
-
批准号:7327678
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
STATISTICAL METHOD FOR GENE/ENVIRONMENT INTERACTION AND GENETIC SUSCEPTIBILITY
-
批准号:6432301
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical Methods For Gene/Environment Interaction And
-
批准号:6504694
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical techniques applied to environmental health sciences
-
批准号:8553817
-
项目类别:
-
资助金额:$4.86万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
Statistical techniques applied to environmental health sciences
-
批准号:8734180
-
项目类别:
-
资助金额:$6.78万
-
财政年份:--
-
负责人:DAVID M UMBACH
-
依托单位:
海外基金