课题基金 / 基金详情

Methods for Epidemiology Studies

Methods for Epidemiology Studies
流行病学研究方法
批准号:
8349580
负责人:
Nilanjan Chatterjee
金额:
$324.22万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Nilanjan Chatterjee的其他基金

相似基金

相关文献

中文摘要
翻译
基于下一代基因分型和测序技术,开展了下一代关联研究设计的遗传流行病学调查方法。一个项目调查了使用下一代测序技术进行变异检测和疾病相关检测研究的覆盖深度和样本数量之间的最佳平衡。第二个项目调查了两阶段设计的潜在效用,其中所有受试者可以使用现有全基因组关联研究中常用的标准基因分型平台进行基因分型,而少数受试者可以使用新一代、更密集的基因分型平台进行基因分型,旨在为不太常见的变异提供更高的覆盖率。在大规模全基因组关联研究(GWAS)中,提出了新的拷贝数变异(CNV)的统计方法。目前可用的CNV调用包检测短CNV的功率较低,因此没有足够的功率检测与疾病相关的短CNV。Shi博士一直在尝试整合b等位基因频率,SNP探针之间的连锁不平衡以及受试者之间的相关性,以提高CNV呼叫。在c++中开发了一个软件包SegCNV,用于调用不相关主题中的cnv。在存在异质性的全基因组关联研究中,开发了一种新的方法进行meta分析。结合不同但可能相关的性状的GWAS现在被认为是一个特别有前途的方向,可以发现具有小而常见的多效效应的位点。然而,对于个体变异可能只与性状的一个子集相关或在不同方向上表现出影响的情况,用于元分析或汇总分析的经典方法并不是最佳的。我们提出了一种方法,通过不可知论地探索研究的子集来推广标准的固定效应分析方法,以确定是否存在相同方向或相反方向的真实关联信号。一个有效的近似用于快速评估p值。一个用R编程语言开发的名为ASSET的软件包已经被免费分发给外部使用。许多项目涉及基因-环境相互作用研究方法的发展。一个项目研究了在基因-环境相互作用的病例对照研究中纳入抽样权重的策略。第二个项目开发了一种利用来自同一遗传区域的多个snp研究基因-环境相互作用的新方法。一般统计方法已经开发了一种新的方法,用于一般基于重采样的测试的高效p值评估程序,当评估具有小p值的测试统计量时,所需的计算时间仅为标准基于重采样的程序所需计算时间的1%甚至0.0002%;的概率的经验估计在许多分子流行病学研究中是相当普遍的兴趣,这些研究使用严格的测试程序来避免假阳性。该方法已在流行的R编程语言中以一个用户友好的软件包实现,供其他研究人员自由使用。BB成员开发了一种新的更灵活的逻辑回归模型,允许对一些风险进行纯线性建模。该模型的优点是可以自动计算绝对风险和风险差异,这是临床流行病学和转化医学中的两个关键量。该模型与Jay Lubin和Dale Preston在Epicure中为其他目的开发的一些特殊模型有关。许多具有挑战性的计算问题已经得到解决,用于拟合模型的通用软件已经公开可用。低至100%的高危人群。PNF(p)通过指出必须跟踪多少高危人群来评估覆盖100%病例的可行性。给出了这两个准则与Lorenz曲线及其逆曲线的关系,并给出了PCF和PNF估计的分布理论。开发基于影响函数的新方法,用于单个风险模型的推断,以及比较两个风险模型的pcf和pnf,这两个模型都是在相同的验证数据中进行评估的。提出了一类glmm平均结构的GOF检验方法。我们的检验统计量是由协变量空间的分区定义的单元格中观测值与估计模型下期望值之间差的二次形式。当模型参数用极大似然估计时,我们证明了该检验统计量具有渐近卡方分布,并在分析和模拟中研究了它在局部替代下的功率。对于线性混合模型,我们还研究了用最小二乘法和矩量法估计参数时的整定问题。几个数据示例说明了这些方法。进行了广泛的模拟,比较了2度(FP2)和4度(FP4)的FPs以及使用广义交叉验证(GCV)和限制最大似然(REML)进行平滑参数选择的p样条的两种变体。使用不同的样本量和信噪比,评估了p样条和FPs恢复连续、二元和生存结果与线性、二次和更复杂的非线性函数之间关联的真实“函数形式”的能力。对于更弯曲的函数FP2,目前实现中用于拟合FPs的默认设置,与基于样条的估计器和FP4相比,显示出相当大的偏差和更高的均方误差(MSE),后者在大多数模拟设置中表现同样良好。然而,由于特定的原点选择,FPs容易产生伪影,而基于GCV的p样条曲线有时会显示出摇摆不定的估计,特别是对于小样本量。暴露评估、暴露测量误差和暴露数据缺失比较了亚抽样下两种诊断试验的诊断准确性和一致性。采用两阶段半参数Cox比例风险回归模型,开发了纳入多种疾病特征数据的队列研究中的疾病发病率分析方法,该模型允许通过不同疾病特征的水平来检查协变量影响的异质性。提出了一种处理竞争风险数据中缺失失效原因的估计方程方法。证明了在一般缺失随机假设下估计方程方法的渐近无偏性,提出了一种新的基于影响函数的夹心方差估计量。这些方法通过模拟研究和涉及癌症预防研究(CPS-II)营养队列的真实数据应用来说明。描述性流行病学研究方法确定死亡率最高和最低的地区,并制作相应的彩色编码地图,有助于流行病学家确定有希望进行分析病因学研究的地区。基于带有协变量的两阶段PoissonGamma模型,我们使用已知危险因素(如吸烟率)的信息来调整死亡率,并揭示可能反映先前被掩盖的病因学关联的相对风险的剩余变异。除了协变量调整外,我们还研究了基于smr、经验贝叶斯(EB)估计和后验百分位排名(PPR)方法的排名,并指出了需要更复杂程序的情况,以便获得高概率正确分类相对风险高和低的地区。
英文摘要
Methods for Genetic Epidemiology Investigations have been conducted on designs for next generation association studies based on next generation genotyping and sequencing technologies. One project has investigated the optimal balance between coverage depth and number of samples for studies of variant detection and disease-association testing using next generation sequencing technologies. A second project investigated potential utility of two-stage designs where all subjects may be genotyped using a standard genotyping platforms that have been commonly used for existing genome-wide association studies and a small number of subjects may be genotyped using newer generation, denser, genotyping platforms designed to provide higher coverage for less common variants. New statistical methods for calling copy number variations (CNV) in large scale genome-wide association studies (GWAS) have been developed. The currently available CNV calling packages have low power for detecting short CNV and hence dont have sufficient power for detecting short CNVs associated with diseases. Dr. Shi has been trying to integrate the B-allele frequencies, the linkage disequilibrium among SNP probes and the relatedness among subjects to improve the CNV calling. A software package SegCNV in C++ for calling CNVs in unrelated subjects has been developed. A new method has been developed for conducting meta-analysis of genome-wide association studies in presence of heterogeneity. Combining GWAS of distinct, but possibly related, traits is now considered a particularly promising direction for the discovery of loci with small but common pleiotropic effects. Classical approaches for meta- or pooled- analysis, however, are not optimal for such a setting in which individual variants are likely to be associated with only a subset of the traits or demonstrate effects in different directions. We propose a method that generalizes standard fixed-effects analytic approaches by agnostically exploring subsets of studies for the presence of true association signals, in either the same direction or possibly in opposite directions. An efficient approximation is used for rapid evaluation of p-values. A software package in R programming language called ASSET has been developed is being freely distributed for external use. A number of projects involved development of methods for studies of gene-environment interactions. One project studied strategies for incorporation of sampling weights in case-control studies of gene-environment interaction. A second project developed a novel method for studying gene-environment interaction using multiple SNPs from same genetic region. General statistical methods New method has been developed for highly efficient p-value evaluation procedure for the general resampling-based test that requires as little as 1% of even 0.0002% of the computing time required by standard resampling-based procedures when evaluating a test statistic with a small p-value; an empirical estimate of probabilities of the order of is quite common interest in many molecular epidemiology studies that use rigorous testing procedures to avoid false positives . The method has been implemented in a user friendly package in the popular R programming language for other researchers to use freely. BB members have developed a new more flexible logistic regression model that allows some exposures to be modeled purely linearly. The advantage of this model is that it can automatically calculate absolute risks and risk differences -- the 2 key quantities in clinical epidemiology and translational medicine. The model is related to some of the special models Jay Lubin and Dale Preston have developed for other purposes in Epicure. A number of challenging computational issues have been solved and a general-purpose software for fitting the model has been made publicly available. lows 100q% of the population at highest risk. PNF(p) assess the feasibility of covering 100p% of cases by indicating how much of the population at highest risk must be followed. Showed the relationship of those two criteria to the Lorenz curve and its inverse, and present distribution theory for estimates of PCF and PNF. Develop new methods, based on influence functions, for inference for a single risk model, and also for comparing the PCFs and PNFs of two risk models, both of which were evaluated in the same validation data. Proposed a class of GOF tests for the mean structure of GLMMs. Our test statistic is a quadratic form of the difference between observed values and the values expected under the estimated model in cells defined by a partition of the covariate space. We show that this test statistic has an asymptotic chi-squared distribution when model parameters are estimated by maximum likelihood, and study its power under local alternatives both analytically and in simulations. For the case of linear mixed models, we also study the setting when parameters are estimated by least squares and method of moments. Several data examples illustrate the methods. Conducted extensive simulations to compare FPs of degree 2 (FP2) and degree 4 (FP4) and two variants of P-splines that used generalized cross validation (GCV) and restricted maximum likelihood (REML) for smoothing parameter selection. The ability of P-splines and FPs to recover the true'' functional form of the association between continuous, binary and survival outcomes and exposure for linear, quadratic and more complex, non-linear functions, using different sample sizes and signal to noise ratios was evaluated. For more curved functions FP2, the current default setting in implementations for fitting FPs, showed considerable bias and consistently higher mean squared error (MSE) compared to spline-based estimators and FP4, that performed equally well in most simulation settings. FPs however, are prone to artefacts due to the specific choice of the origin, while P-splines based on GCV reveal sometimes wiggly estimates in particular for small sample sizes. Exposure Assessment, Errors in Exposure Measurements, and Missing Exposure Data Compared the diagnostic accuracy and agreement of two diagnostic tests under subsampling. Developed methods for analysis of disease incidence in cohort studies incorporating data on multiple disease traits using a two-stage semiparametric Cox proportional hazards regression model that allows one to examine the heterogeneity in the effect of the covariates by the levels of the different disease traits. Proposed an estimating-equation approach for handling missing cause of failure in competing-risk data. Proved asymptotic unbiasedness of the estimating-equation method under a general missing-at-random assumption and propose a novel influence-function based sandwich variance estimator. The methods are illustrated using simulation studies and a real data application involving the Cancer Prevention Study (CPS-II) nutrition cohort. Methods for descriptive epidemiologic studies Identifying regions with the highest and lowest mortality rates and producing the corresponding color-coded maps help epidemiologists identify promising areas for analytic etiological studies. Based on a two-stage PoissonGamma model with covariates, we used information on known risk factors, such as smoking prevalence, to adjust mortality rates and reveal residual variation in relative risks that may reflect previously masked etiological associations. In addition to covariate adjustment, we studied rankings based on SMRs, empirical Bayes (EB) estimates, and a posterior percentile ranking (PPR) method and indicated circumstances that warrant the more complex procedures in order to obtain a high probability of correctly classifying the regions with high and low relative risks .
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Methods for Data Integration and Applications to Genome-wide Association Studies
  • 批准号:
    10889298
  • 项目类别:
  • 资助金额:
    $29.0万
  • 财政年份:
    2023
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10609504
  • 项目类别:
  • 资助金额:
    $61.31万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10416066
  • 项目类别:
  • 资助金额:
    $32.54万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10263893
  • 项目类别:
  • 资助金额:
    $63.77万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: