课题基金 / 基金详情

Methods for Epidemiology Studies

Methods for Epidemiology Studies
流行病学研究方法
批准号:
8177710
负责人:
Nilanjan Chatterjee
金额:
$350.55万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Nilanjan Chatterjee的其他基金

相似基金

相关文献

中文摘要
翻译
病例对照数据中隐藏的群体亚结构可能会扭曲遗传关联的CochranArmitage趋势检验(CATTs)的表现。提出了一种统一的方法,用于推导一般族中任意CATT在三种情况下的偏差和方差失真。我们的结果提供了对偏差和方差失真的一些拟议修正的特性的见解,并说明了为什么它们可能不能完全纠正总体子结构的影响。比较了几种组合全基因组关联数据的方法,包括在控制全基因组显著性水平的情况下检测疾病相关SNP的能力,以及检测概率,即根据低p值选择的T个最有希望的SNP之一的特定疾病相关SNP的概率。专注于单一自由度的元分析方法比全局方法具有更高的能力和检测概率。研究了原发性疾病罕见、继发表型和遗传标记二分的情况。仅基于对照的遗传标记与二级表型之间的关联分析是有效的。开发了一种自适应加权方法,将病例和对照数据结合起来研究关联,同时仅在有强烈证据表明相互作用时才减少到对照分析。开发工具来估计易感位点的数量和基于现有GWAS发现的一个性状的效应大小分布,然后预测统计能力和风险预测效用研究,整合效应大小的估计分布。利用已报道的身高、克罗恩病和乳腺癌、前列腺癌和结直肠癌(BPC)的GWAS结果来估计,在低渗透的常见变异谱中,每一种特征都可能含有额外的位点,这些位点加在一起可以解释至少15-20%的已知遗传力。通过进行足够大的研究,将有可能发现这些特征的大量额外位点。然而,对于具有适度家族聚集性的疾病,如BPC癌,预测的发现不太可能导致临床上重要风险模型的高歧视性。使用来自CGEMS病例对照乳腺癌全基因组关联研究的数据,从经验上证明,由于种群分层(PS)的存在,导致基因组中的远程连锁不平衡,仅病例和相关方法有可能产生大规模假阳性。我们表明,可以通过考虑假设非连锁位点之间的基因-基因独立性的方法来消除偏差,而不是在整个群体中,而只是在均匀地层中,可以根据适当大的PS标记面板的主要成分来定义。我们提出了参数策略和非参数策略来利用这种条件基因-基因独立性假设。提出假阴性报告概率(FNRP),类似于Sholom Wacholder的假阳性报告概率(FPRP),用于同时分析多个遗传变异。建议使用FPRP和FNRP背后的贝叶斯因子作为任何变异-疾病关系的纯粹统计证据,并表明它们的比率接近真实的贝叶斯因子,从而将FPRP/FNRP框架与传统的贝叶斯统计统一起来。风险建模提出并研究两个标准,以评估预测疾病发病率的模型对筛查和预防的有用性,或评估疾病诊断后管理的预后模型的有用性。采用PCF的病例比例(q)和需要随访的比例(PNF)。PCF(q)评估一个项目的有效性,该项目跟踪了100%的高危人群。PNF(p)通过指出必须跟踪多少高危人群来评估覆盖100%病例的可行性。给出了这两个准则与Lorenz曲线及其逆曲线的关系,并给出了PCF和PNF估计的分布理论。开发基于影响函数的新方法,用于单个风险模型的推断,以及比较两个风险模型的pcf和pnf,这两个模型都是在相同的验证数据中进行评估的。提出了一类glmm平均结构的GOF检验方法。我们的检验统计量是由协变量空间的分区定义的单元格中观测值与估计模型下期望值之间差的二次形式。当模型参数用极大似然估计时,我们证明了该检验统计量具有渐近卡方分布,并在分析和模拟中研究了它在局部替代下的功率。对于线性混合模型,我们还研究了用最小二乘法和矩量法估计参数时的整定问题。几个数据示例说明了这些方法。进行了广泛的模拟,比较了2度(FP2)和4度(FP4)的FPs以及使用广义交叉验证(GCV)和限制最大似然(REML)进行平滑参数选择的p样条的两种变体。使用不同的样本量和信噪比,评估了p样条和FPs恢复连续、二元和生存结果与线性、二次和更复杂的非线性函数之间关联的真实“函数形式”的能力。对于更弯曲的函数FP2,目前实现中用于拟合FPs的默认设置,与基于样条的估计器和FP4相比,显示出相当大的偏差和更高的均方误差(MSE),后者在大多数模拟设置中表现同样良好。然而,由于特定的原点选择,FPs容易产生伪影,而基于GCV的p样条曲线有时会显示出摇摆不定的估计,特别是对于小样本量。暴露评估、暴露测量误差和暴露数据缺失比较了亚抽样下两种诊断试验的诊断准确性和一致性。采用两阶段半参数Cox比例风险回归模型,开发了纳入多种疾病特征数据的队列研究中的疾病发病率分析方法,该模型允许通过不同疾病特征的水平来检查协变量影响的异质性。提出了一种处理竞争风险数据中缺失失效原因的估计方程方法。证明了在一般缺失随机假设下估计方程方法的渐近无偏性,提出了一种新的基于影响函数的夹心方差估计量。这些方法通过模拟研究和涉及癌症预防研究(CPS-II)营养队列的真实数据应用来说明。描述性流行病学研究方法确定死亡率最高和最低的地区,并制作相应的彩色编码地图,有助于流行病学家确定有希望进行分析病因学研究的地区。基于带有协变量的两阶段PoissonGamma模型,我们使用已知危险因素(如吸烟率)的信息来调整死亡率,并揭示可能反映先前被掩盖的病因学关联的相对风险的剩余变异。除了协变量调整外,我们还研究了基于smr、经验贝叶斯(EB)估计和后验百分位排名(PPR)方法的排名,并指出了需要更复杂程序的情况,以便获得高概率正确分类相对风险高和低的地区。
英文摘要
Methods for Genetic Epidemiology Hidden population substructure in casecontrol data can distort the performance of CochranArmitage trend tests (CATTs) for genetic associations. A unified approach is presented for deriving the bias and variance distortion under three scenarios for any CATT in a general family. Our results provide insight into the properties of some proposed corrections for bias and variance distortion and show why they may not fully correct for the effects of population substructure . Compared several procedures to combine geonome-wide association data both in terms of the power to detect a disease-associated SNP while controlling the genome-wide significance level, and in terms of the detection probability, namely the probability that a particular disease-associated SNP will be among the T most promising SNPs selected on the basis of low p-values. Meta-analytic approaches that focus on a single degree of freedom had higher power and detection probablities than global methods . Studied the situation in which the primary disease is rare and the secondary phenotype and genetic markers are dichotomous. An analysis of the association between a genetic marker and the secondary phenotype based on controls only is valid. Developed an adaptively weighted method that combines the case and control data to study the association, while reducing to the controls only analysis if there is strong evidence of an interaction . Developed tools to estimate the number of susceptibility loci and the distribution of their effect sizes for a trait based on discoveries from existing GWAS and then to project statistical power and risk prediction utility studies integrating over estimated distributions of effect sizes. Used reported GWAS findings for height, Crohns disease and cancers of breast, prostate and colorectum (BPC) to estimate that each of the traits is likely to harbor additional loci within the spectrum of low penetrant common variants that together could explain at least 15-20% of known heritability. By conducting sufficiently large studies, it will be possible to discover a large set of additional loci for these traits. However, for diseases with modest familial aggregation, like BPC cancers, the projected discoveries are unlikely to lead to high discriminatory power for clinically important risk models. Used data from the CGEMS case-control genome-wide association study of breast cancer to demonstrate empirically that the case-only and related methods have the potential to create large-scale false positives due to the presence of population stratification (PS) that creates long-range linkage disequilibrium in the genome. We show that the bias can be removed by considering methods that assume gene-gene independence between unlinked loci, not in the entire population, but only within homogeneous strata that can be defined based on the principal components of a suitably large panel of PS markers. We propose both parametric and non-parametric strategies for exploiting such conditional gene-gene independence assumptions. Proposed a False Negative Report Probability (FNRP), analogously to Sholom Wacholder's False Positive Report Probability (FPRP), for analyzing multiple genetic variants simultaneously. Propose use of the Bayes Factors underlying FPRP and FNRP as the pure statistical evidence for any variant-disease relationship, and show that their ratio approximates the true Bayes Factor, thus unifying the FPRP/FNRP framework with traditional Bayesian statistics. Risk modeling Propose and study two criteria to assess the usefulness of models that predict risk of disease incidence for screening and prevention, or the usefulness of prognostic models for management following disease diagnosis. The proportion of cases followed PCF(q) and the proportion needed to follow-up, PNF(p). PCF(q) assesses the effectiveness of a program that follows 100q% of the population at highest risk. PNF(p) assess the feasibility of covering 100p% of cases by indicating how much of the population at highest risk must be followed. Showed the relationship of those two criteria to the Lorenz curve and its inverse, and present distribution theory for estimates of PCF and PNF. Develop new methods, based on influence functions, for inference for a single risk model, and also for comparing the PCFs and PNFs of two risk models, both of which were evaluated in the same validation data. Proposed a class of GOF tests for the mean structure of GLMMs. Our test statistic is a quadratic form of the difference between observed values and the values expected under the estimated model in cells defined by a partition of the covariate space. We show that this test statistic has an asymptotic chi-squared distribution when model parameters are estimated by maximum likelihood, and study its power under local alternatives both analytically and in simulations. For the case of linear mixed models, we also study the setting when parameters are estimated by least squares and method of moments. Several data examples illustrate the methods. Conducted extensive simulations to compare FPs of degree 2 (FP2) and degree 4 (FP4) and two variants of P-splines that used generalized cross validation (GCV) and restricted maximum likelihood (REML) for smoothing parameter selection. The ability of P-splines and FPs to recover the true'' functional form of the association between continuous, binary and survival outcomes and exposure for linear, quadratic and more complex, non-linear functions, using different sample sizes and signal to noise ratios was evaluated. For more curved functions FP2, the current default setting in implementations for fitting FPs, showed considerable bias and consistently higher mean squared error (MSE) compared to spline-based estimators and FP4, that performed equally well in most simulation settings. FPs however, are prone to artefacts due to the specific choice of the origin, while P-splines based on GCV reveal sometimes wiggly estimates in particular for small sample sizes. Exposure Assessment, Errors in Exposure Measurements, and Missing Exposure Data Compared the diagnostic accuracy and agreement of two diagnostic tests under subsampling. Developed methods for analysis of disease incidence in cohort studies incorporating data on multiple disease traits using a two-stage semiparametric Cox proportional hazards regression model that allows one to examine the heterogeneity in the effect of the covariates by the levels of the different disease traits. Proposed an estimating-equation approach for handling missing cause of failure in competing-risk data. Proved asymptotic unbiasedness of the estimating-equation method under a general missing-at-random assumption and propose a novel influence-function based sandwich variance estimator. The methods are illustrated using simulation studies and a real data application involving the Cancer Prevention Study (CPS-II) nutrition cohort. Methods for descriptive epidemiologic studies Identifying regions with the highest and lowest mortality rates and producing the corresponding color-coded maps help epidemiologists identify promising areas for analytic etiological studies. Based on a two-stage PoissonGamma model with covariates, we used information on known risk factors, such as smoking prevalence, to adjust mortality rates and reveal residual variation in relative risks that may reflect previously masked etiological associations. In addition to covariate adjustment, we studied rankings based on SMRs, empirical Bayes (EB) estimates, and a posterior percentile ranking (PPR) method and indicated circumstances that warrant the more complex procedures in order to obtain a high probability of correctly classifying the regions with high and low relative risks .
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Methods for Data Integration and Applications to Genome-wide Association Studies
  • 批准号:
    10889298
  • 项目类别:
  • 资助金额:
    $29.0万
  • 财政年份:
    2023
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10609504
  • 项目类别:
  • 资助金额:
    $61.31万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10416066
  • 项目类别:
  • 资助金额:
    $32.54万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10263893
  • 项目类别:
  • 资助金额:
    $63.77万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: