课题基金 / 基金详情

Methods for Epidemiology Studies

Methods for Epidemiology Studies
流行病学研究方法
批准号:
7966676
负责人:
Nilanjan Chatterjee
金额:
$301.69万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:

项目摘要

项目成果

Nilanjan Chatterjee的其他基金

相似基金

相关文献

中文摘要
翻译
遗传流行病学方法病例对照全基因组关联研究提供了大量的遗传信息,可用于研究继发性表型。在病例对照研究中,已经开发出方法来分析遗传标记和继发表型之间的关联,以考虑由于使用病例而导致的潜在偏差,当继发表型和遗传标记之间存在相互作用时,对原发疾病的风险。所提出的方法自适应地结合了病例和对照数据,而如果有强有力的证据表明相互作用,则只对对照进行分析。模拟和渐近理论表明,与仅对对照进行分析相比,自适应加权方法可以减少预指定SNP估计的均方误差,并增加在全基因组研究中发现新关联的能力。与单阶段设计相比,大型两阶段全基因组关联研究(GWAS)已被证明可以减少所需的基因分型,并且功率损失很小,前提是将相当一部分病例和对照组pi(样本)纳入第一阶段。一种新的衡量能力的方法,检测概率(DP),已经被定义为一个给定的疾病相关单核苷酸多态性(SNP)在阶段1的p值的最低等级中具有p值的概率,以及在阶段1选择的SNP中在阶段2具有p值的概率。我们的研究结果表明,应该避免第一阶段较小的多阶段设计(例如,pi(样本)<= 0.25),并且在早期研究中使用小的第一阶段进行额外的基因分型将产生先前未选择的疾病相关snp。越来越多的人认识到,途径分析——对结果与生物学途径中一组snp之间的关联进行联合测试——可能潜在地补充单snp分析,并为复杂疾病的遗传结构提供额外的见解。基于自适应秩截断积(ARTP)统计,提出了一类高度灵活的途径分析方法,该方法可以有效地结合途径中不同snp和基因的关联证据。计算效率的排列算法被开发用于评估所提出的测试统计的统计显著性。这些方法被用于研究尼古丁受体通路与吸烟行为之间的关系。DCEG最近对NHANES-III家庭成员进行了基因分型,开创了真正以美国人群为基础的基因研究领域。然而,基于问卷调查数据的NHANES-III家庭中报告的家庭关系是不完整的,对于家庭成员的实际生物亲缘关系没有定论。已经开发了统计算法,用于使用DNA指纹(Identifiler分析中的str)来推断NHANES-III家庭中的家庭关系。这项研究的结果将对未来的家庭研究和NHANES的GWAS至关重要。<br> <br> <br> <b>环境暴露相对风险模型</b><br>复杂的疾病,如癌症,通常可以根据疾病的各种病理和分子特征划分为亚型。采用两阶段、半参数Cox比例风险回归模型对纳入多种疾病特征数据的队列研究中的疾病发病率进行了分析,该模型允许人们通过不同疾病特征的水平来检查协变量影响的异质性。对于存在缺失疾病特征的推断,我们提出了一种一般化的估计方程方法来处理竞争风险数据中缺失的失败原因。通过模拟研究和涉及癌症预防研究(CPS-II)营养队列的真实数据应用说明了这些方法。<br> <br>对于大多数疾病,单一的生物标志物在实际应用中没有足够的敏感性或特异性。已经开发了一种方法,将几个生物标志物组合成一个复合标记评分,而不假设预测因子分布的模型。使用足够的降维技术,原始标记被替换为低维版本,通过标记的线性变换获得,该标记包含足够的信息,用于预测者对结果的回归。利用线性变换的渐近性质,通过似然比统计量将其组合成一个标量诊断评分。该评分的表现是通过接受者-操作者特征曲线(ROC)下的面积来评估的,ROC是一种流行的对二元疾病结果的单一连续诊断标记的区分能力的汇总测量。还推导了用于评估个体生物标志物对诊断评分贡献的渐近卡方检验。在流行病学研究中,部分问卷设计(PQD)通过将问卷的不同子集分配给不同但重叠的研究参与者,可以减少问卷冗长所带来的成本、时间和其他实际负担。在PQD和其他可以产生非单调缺失数据的研究设置下,已经开发了回归模型的半参数推理方法。这些方法使用来自非霍奇金淋巴瘤病例对照研究的数据进行说明,其中主要化学物质暴露的数据是使用两种不同的仪器在两个不同但重叠的参与者子集上收集的。对于许多疾病,由于可能不存在完美的“黄金标准”,或者获得标准的成本太高,因此很难或不可能做出明确的诊断。提出了一种方法,使用连续的测试结果来估计特定人群中的疾病患病率,并估计可能影响患病率的因素的影响。受乌干达镰状细胞性贫血儿童中人类疱疹病毒8型研究的启发,采用两种酶免疫测定法评估感染状况,拟合了一个双组分多变量混合模型。组件密度使用参数密度建模,参数密度包括数据转换和灵活的转换模型。此外,混合比例,即潜在变量对应于真实未知感染状态的概率,通过逻辑回归建模以纳入协变量。该模型包括多元正态密度的混合,作为一种特殊情况,能够适应数据中不寻常的形状和偏度。在模拟中评估了模型的性能,并将该方法应用于乌干达的研究中。
英文摘要
<b>Methods for Genetic Epidemiology</b><br> Case-control genome-wide association studies provide a vast amount of genetic information that may be used to investigate secondary phenotypes. Methods have been developed to analyze the association between genetic markers and a secondary phenotypes in a case-control studies accounting for potential bias due to use of cases when there is an interaction between the secondary phenotype and the genetic marker on the risk of the primary disease. The proposed method adaptively combines the case and control data, while reducing to the controls only analysis if there is strong evidence of an interaction. Simulations and asymptotic theory indicate that the adaptively weighted method can reduce the mean square error for estimation with a pre-specified SNP and increase the power to discover a new association in a genome-wide study, compared to an analysis of controls only.<br> <br> Large two-stage genome-wide association studies (GWAS) have been shown to reduce required genotyping with little loss of power, compared to a one-stage design, provided a substantial fraction of cases and controls, pi (sample), is included in stage 1. A new measure of power, the detection probability (DP), has been defined as the probability that a given disease-associated single-nucleotide polymorphism (SNP) will have a p-value among the lowest ranks of p-values at stage 1, and, among those SNPs selected at stage 1, at stage 2. Our results suggest that multistage designs with small first stages (e.g., pi (sample) <= 0.25) should be avoided, and that additional genotyping in earlier studies with small first stages will yield previously unselected disease-associated SNPs.<br> <br> It is increasingly recognized that pathway analysesa joint test of association between the outcome and a group of SNPs within a biological pathwaycould potentially complement single-SNP analysis and provide additional insights for the genetic architecture of complex diseases. A class of highly flexible pathway analysis approaches has been proposed based on an adaptive rank truncated product (ARTP) statistic that can effectively combine evidence of associations over different SNPs and genes within a pathway. Computationally-efficient permutation algorithms are developed for evaluating the statistical significance of the proposed test-statistics. The methods are applied to a study of the association between the nicotinic receptor pathway and cigarette smoking behaviors.<br> <br> DCEG has recently conducted genotyping of NHANES-III household members, inaugurating the field of truly US population-based genetic research. However, reported family relationships within NHANES-III households based on questionnaire data are incomplete and inconclusive with regards to actual biological relatedness of family members. Statistical algorithms have been developed for using DNA fingerprints (the STRs in the Identifiler assay) to infer family relationships within NHANES-III households. The findings from this study will be pivotal for future family studies and GWAS within NHANES.<br> <br> <br> <b>Models for Relative Risks of Environmental Exposures</b><br> Complex diseases, like cancer, can often be classified into subtypes using various pathological and molecular traits of the disease. Methods have been developed for analysis of disease incidence in cohort studies incorporating data on multiple disease traits using a two-stage, semi-parametric Cox proportional hazard regression model that allows one to examine the heterogeneity in the effect of the covariates by the levels of the different disease traits. For inference in the presence of missing disease traits, we propose a generalization of an estimating-equation approach for handling missing cause of failure in competing-risk data. The methods are illustrated using simulation study and a real data application involving the Cancer Prevention Study (CPS-II) nutrition cohort.<br> <br> For most diseases, single biomarkers do not have adequate sensitivity or specificity for practical purposes. An approach has been developed to combine several biomarkers into a composite marker score without assuming a model for the distribution of the predictors. Using sufficient dimension reduction techniques, the original markers are replaced with a lower-dimensional version, obtained through linear transformations of markers that contain sufficient information for regression of the predictors on the outcome. The linear transformations are combined using their asymptotic properties into a scalar diagnostic score via the likelihood ratio statistic. The performance of this score is assessed by the area under the receiver-operator characteristics curve (ROC), a popular summary measure of the discriminatory ability of a single continuous diagnostic marker for binary disease outcomes. An asymptotic chi-squared test for assessing individual biomarker contribution to the diagnostic score is also derived.<br> <br> <b>Exposure Assessment, Errors in Exposure Measurements, and Missing Exposure Data</b><br> In epidemiologic studies, partial questionnaire design (PQD) can reduce cost, time and other practical burdens associated with lengthy questionnaire by assigning different subsets of the questionnaire to different, but overlapping, subsets of the study participants. Methods for semi-parametric inference for regression model under PQD and other study settings that can generate non-monotone missing data in covariates have been developed. The methods are illustrated using data from a case-control study of non-Hodgkin's lymphoma where the data on the main chemical exposures of interest are collected using two different instruments on two different, but overlapping, subsets of the participants.<br> <br> For many diseases, it is difficult or impossible to establish a definitive diagnosis because a perfect "gold standard" may not exist or may be too costly to obtain. A method is proposed to use continuous test results to estimate prevalence of disease in a given population and to estimate the effects of factors that may influence prevalence. Motivated by a study of human herpes virus 8 among children with sickle-cell anemia in Uganda, where 2 enzyme immunoassays were used to assess infection status, a 2-component multivariate mixture model is fitted. The component densities are modeled using parametric densities that include data transformation as well as flexible transformed models. In addition, the mixing proportion, the probability of a latent variable corresponding to the true unknown infection status, is modeled via a logistic regression to incorporate covariates. This model includes mixtures of multivariate normal densities as a special case and is able to accommodate unusual shapes and skewness in the data. The model performance is assessed in simulations and results are presented from application of the methods to the Ugandan study.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Methods for Data Integration and Applications to Genome-wide Association Studies
  • 批准号:
    10889298
  • 项目类别:
  • 资助金额:
    $29.0万
  • 财政年份:
    2023
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10609504
  • 项目类别:
  • 资助金额:
    $61.31万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10416066
  • 项目类别:
  • 资助金额:
    $32.54万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
  • 批准号:
    10263893
  • 项目类别:
  • 资助金额:
    $63.77万
  • 财政年份:
    2020
  • 负责人:
    Nilanjan Chatterjee
  • 依托单位:
海外基金