Methods for Epidemiology Studies
Methods for Epidemiology Studies
批准号:
7593206
负责人:
Nilanjan Chatterjee
金额:
$159.94万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AccountingAlgorithmsArsenicBiometryBladderCase-Control StudiesCategoriesCessation of lifeCigaretteClassComplexDataDetectionDiseaseDisease AssociationDisease regressionEnvironmental ExposureEpidemiologic MethodsEpidemiologic StudiesEsophagusExhibitsFollow-Up StudiesGene OrderGenotypeHaplotypesJointsJudgmentLinear RegressionsLinkage DisequilibriumLogistic RegressionsLogisticsMalignant NeoplasmsMalignant neoplasm of lungMeasurementMethodologyMethodsMissionModelingModificationNegative FindingNumbersOccupational EpidemiologyOral cavityOutcomePancreasPaperPatternPopulationPopulation Attributable RisksProbabilityProceduresPropertyPublishingRandomized Clinical TrialsRare DiseasesRateRegression AnalysisRelative (related person)Relative RisksResearch DesignRiskRisk EstimateRisk FactorsSample SizeSamplingScoreSeriesSingle Nucleotide PolymorphismSmokerSmokingSourceStagingStandards of Weights and MeasuresStatistical MethodsStudy modelsSurveysTechniquesTestingTimeTreesVariantbasecancer sitecase controlcohortdaydesigndisorder riskdrinking waterepidemiology studyfollow-upgene interactiongenetic associationgenetic epidemiologygenetic variantgenome wide association studygraphical user interfacehazardhuman dataimprovedmembermortalityprogramsresponsesimulationsizetime usetrend
中文摘要
遗传流行病学方法随着越来越多的基于人群的研究表明遗传变异与疾病风险之间存在关联,有必要改进独立样本随访研究(II期)的设计,以证实在初始阶段(I期)观察到的关联证据。我们建议使用为随机临床试验开发的灵活设计来计算随访研究的样本量。我们应用了一个自举程序来纠正回归均值,也称为赢家诅咒,这是由于选择跟踪具有最强关联的标记而导致的。标准回归模型便于评估主效应和低阶相互作用,但不适用于探索复杂的高阶基因相互作用。基于树的方法对于解纠缠可能的相互作用是一种有吸引力的替代方法,但它在建模可加性主效应方面存在困难。我们提出了一类新的半参数回归模型,称为部分线性树基回归(PLTR)模型,它具有广义线性回归和树模型的优点。我们研究了病例对照全基因组关联研究(CCGWASs)程序的特性,该程序选择卡方趋势检验最大(或相应p值最小)的snp。我们发现,对于罕见疾病,如果SNP基因型在源人群中是独立的,则SNP的关联检测是独立的。这一结果使我们能够开发分析和模拟技术来研究CCGWASs。这些分析表明,需要大量样本才能具有较高的检测概率(真正的疾病SNP出现在卡方值顶部的机会)。统计能力计算为遗传关联研究的设计和解释提供了信息,但很少有程序专门用于不相关受试者的单核苷酸多态性(snp)的病例对照研究。开发了算法和图形用户界面来计算在显性、共显性和隐性模型下的样本量和SNP或单倍型效应的最小可检测风险。程序允许调整由于联动不平衡或多重测试的多重比较。<br><br><b>调查抽样方法和应用</b><br>我们发表了估计各种原因导致的死亡人数(AD)的方法。我们的方法包括首先估计经混杂协变量调整后的人群归因风险(AR),然后将AR乘以特定时期人群中发生的生命死亡率统计数据确定的死亡人数。从队列死亡率随访数据中获得的调整后相对危险度的比例风险回归估计值与风险因素的联合分布相结合,计算调整后危险度。<br><br>我们开发了新的统计方法,用于对聚类数据进行逻辑回归分析,其中在一些协变量类别中几乎没有积极结果。当使用适当的聚类水平方差估计器时,逻辑回归系数通常的渐近Wald和分数假设检验可能会缓慢地收敛到名义水平。我们提出了一种基于模拟的方法来测试逻辑回归系数,该方法在维持名义水平方面优于广义Wald和分数检验以及自举假设检验。所提出的方法在使用十分位风险表检验逻辑回归模型的拟合优度时也很有用。环境暴露相对风险模型为研究吸烟持续时间和吸烟强度的联合影响,我们建立了总包年和每日卷数的3参数线性超额吸烟风险(ERR)模型,利用一项大型肺癌病例对照研究的数据,比较低强度长期暴露的总暴露量与等量高强度短时间暴露的总暴露量。该模型表明,低于每天1520支香烟存在直接暴露率(或暴露率增强)效应,即高强度(持续时间较短)吸烟者的ERR/包年大于低强度(持续时间较长)吸烟者。每天超过20支烟,存在反向暴露率(或降低效价)效应,即高强度吸烟者的ERR/包年比低强度吸烟者小。我们在一系列分析中探索了这种建模方法。将该模型应用于各种癌症研究的数据,包括肺癌、膀胱癌、口腔癌、胰腺癌和食道癌,结果显示,研究中的效力效应一致降低,这在统计上是均匀的,表明在考虑了总包年之后,不同癌症部位的强度模式具有可比性。研究相互作用和效应修正的模型的扩展表明,<i>NAT2</i>状态的吸烟风险变化是由吸烟强度而不是总包年暴露的相互作用引起的。此外,<i>NAT2</i>慢乙酰化者吸烟风险的相对增加随吸烟强度的增加而增加。<br><br><b>暴露评估、暴露测量误差和暴露数据缺失</b><br>我们发表了两篇说明文,讨论了混淆和暴露错误分类的实际影响。在职业流行病学中,这些因素经常被提出来论证观察到的结果是假阳性或假阴性的发现。我们注意到,在职业流行病学中,大量混淆的例子是罕见的。我们还注意到,在非差分测量误差下,由于期望偏差的方向和幅度,由于误分类而产生的假阳性结果是不可能的。我们建议,在对研究设计或数据解释做出判断时,应考虑所有潜在的局限性,更仔细、更现实地考虑发生的可能性、影响的方向和程度。来自世界上饮用水中砷浓度非常高的地区的流行病学数据表明,砷暴露与几种内部癌症风险之间存在很强的关联,并且可以认为这种关联是因果关系。在较低暴露水平下,由于缺乏明确的人体数据,采用高暴露研究的外推法来估计风险。对低暴露人群的研究受到了限制,因为难以估计过去的暴露情况,而且风险的增加相对较小。不同情况下暴露错误分类和小研究规模对风险估计的影响用图表说明
英文摘要
<b>Methods for Genetic Epidemiology</b><br>As more population-based studies suggest associations between genetic variants and disease risk, there is a need to improve the design of follow-up studies (stage II) in independent samples to confirm evidence of association observed at the initial stage (stage I). We proposed to use flexible designs developed for randomized clinical trials in the calculation of sample size for follow-up studies. We applied a bootstrap procedure to correct for regression to the mean, also called winners curse, resulting from choosing to follow up the markers with the strongest associations.<br><br>Standard regression models were convenient for assessing main effects and low-order interactions but not for exploring complex higher-order gene-gene interactions. Tree-based methodology is an attractive alternative for disentangling possible interactions, but it has difficulty in modeling additive main effects. We proposed a new class of semi-parametric regression models, termed partially linear tree-based regression (PLTR) models, which exhibit the advantages of both generalized linear regression and tree models.<br><br>We studied the properties of procedures for case-control genome-wide association studies (CCGWASs) that select the SNPs whose chi-square trend tests are largest (or whose corresponding p-values are smallest). We showed that for rare diseases association tests for SNPs are independent if the SNP genotypes are independent in the source population. This result allowed us to develop analytic and simulation techniques to study CCGWASs. These analyses showed that large samples are needed to have a high detection probability (the chance a true disease SNP appears in the top ranks of chi-square values).<br><br>Statistical power calculations inform the design and interpretation of genetic association studies, but few programs are tailored to case-control studies of single nucleotide polymorphisms (SNPs) in unrelated subjects. Algorithms and graphical user interfaces were developed to calculate sample size and minimum detectable risk for SNP or haplotype effects under dominant, co-dominant, and recessive models. The programs allowed adjustments for multiple comparisons due to linkage disequilibrium or multiple testing.<br><br><b>Survey Sampling Methods and Applications</b><br>We published methods for estimating the attributable number of deaths (AD) from all causes. Our approach involved first estimating population attributable risk (AR) adjusted for confounding covariates, then multiplying the AR by the number of deaths determined from vital mortality statistics that occurred in the population for a specific time period. Proportional hazard regression estimates of adjusted relative hazards obtained from mortality follow-up data from a cohort was combined with a joint distribution of risk factors to compute an adjusted AR.<br><br>We developed new statistical methods for inference from logistic regression analysis with clustered data where there are few positive outcomes in some of the covariate categories. The usual asymptotic Wald and score hypothesis tests for logistic regression coefficients can be slow to converge to nominal levels when appropriate cluster-level variance estimators are used. We presented a simulation-based method for testing logistic regression coefficients which compared favorably to generalized Wald and score tests and a bootstrap hypothesis test in terms of maintaining nominal levels. The proposed methods were also useful when testing goodness-of-fit of logistic regression models using deciles-of-risk tables.<br><br><b>Models for Relative Risks of Environmental Exposures</b><br>To study the joint effects of smoking duration and intensity, we developed a 3-parameter linear excess RR (ERR) model in total pack-years and cigarettes per day to compare total exposure delivered at low intensity for a long period of time with an equal total exposure delivered at high intensity for a short period of time using data from a large case-control study of lung cancer. The model suggested that below 1520 cigarettes per day there was a direct exposure rate (or exposure rate enhancement) effect, i.e., the ERR/pack-year for higher intensity (and shorter duration) smokers was greater than for lower-intensity (and longer duration) smokers. Above 20 cigarettes per day, there was an inverse-exposure-rate (or reduced potency) effect, i.e., the ERR/pack-year for higher intensity smokers was smaller than for lower-intensity smokers. We explored this modeling approach in a series of analyses.<br><br>Application of this model to data from various studies of cancer, including cancers of the lung, bladder, oral cavity, pancreas, and esophagus revealed consistent reduced potency effects across studies, which were statistically homogeneous, indicating that after accounting for total pack-years, intensity patterns were comparable across the diverse cancer sites.<br><br>An extension of the model for studying interactions and effect modification revealed that variations in smoking risk with <i>NAT2</i> status resulted from interactions with smoking intensity and not total pack-years of exposure. In addition, the relative increase in smoking risk in <i>NAT2</i> slow acetylators increased with smoking intensity.<br><br><b>Exposure Assessment, Errors in Exposure Measurements, and Missing Exposure Data</b><br>We published two expository papers discussing the practical impacts of confounding and exposure misclassification. In occupational epidemiology, these factors are routinely raised to argue that an observed result is either a false positive or a false negative finding. We noted that examples of substantial confounding were rare in occupational epidemiology. We also noted that false positive results due to misclassification was unlikely given the expected direction and magnitude of bias expected under non-differential measurement error. We suggested that all potential limitations are considered and that the likelihood of occurrence and the direction and magnitude of effects should be more carefully and realistically considered when making judgments about study design or data interpretation.<br><br>Epidemiologic data from regions of the world with very high arsenic concentrations in drinking water show a strong association between arsenic exposure and risk of several internal cancers, and the association can be considered causal. At lower levels of exposure, in the absence of unambiguous human data, extrapolation from the high exposure studies are used to estimate risk. Studies in lower expose populations have been limited by the challenge of estimating past exposures, and relatively small increases in risk. The effects on risk estimates of exposure misclassification and small study size under various scenarios were graphically illustrated
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Methods for Data Integration and Applications to Genome-wide Association Studies
-
批准号:10889298
-
项目类别:
-
资助金额:$29.0万
-
财政年份:2023
-
负责人:Nilanjan Chatterjee
-
依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
-
批准号:10609504
-
项目类别:
-
资助金额:$61.31万
-
财政年份:2020
-
负责人:Nilanjan Chatterjee
-
依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
-
批准号:10416066
-
项目类别:
-
资助金额:$32.54万
-
财政年份:2020
-
负责人:Nilanjan Chatterjee
-
依托单位:
Multifactoral breast cancer risk prediction accounting for ethnic and tumor diversity
-
批准号:10263893
-
项目类别:
-
资助金额:$63.77万
-
财政年份:2020
-
负责人:Nilanjan Chatterjee
-
依托单位:
Robust Methods for Polygenic Analysis to Inform Disease Etiology and Enhance Risk Prediction
-
批准号:9920753
-
项目类别:
-
资助金额:$54.72万
-
财政年份:2019
-
负责人:Nilanjan Chatterjee
-
依托单位:
Robust Methods for Polygenic Analysis to Inform Disease Etiology and Enhance Risk Prediction
-
批准号:10359748
-
项目类别:
-
资助金额:$57.58万
-
财政年份:2019
-
负责人:Nilanjan Chatterjee
-
依托单位:
Robust Methods for Polygenic Analysis to Inform Disease Etiology and Enhance Risk Prediction
-
批准号:10112944
-
项目类别:
-
资助金额:$55.83万
-
财政年份:2019
-
负责人:Nilanjan Chatterjee
-
依托单位:
Robust Methods for Polygenic Analysis to Inform Disease Etiology and Enhance Risk Prediction
-
批准号:10579942
-
项目类别:
-
资助金额:$57.53万
-
财政年份:2019
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:8565443
-
项目类别:
-
资助金额:$323.27万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:9154202
-
项目类别:
-
资助金额:$321.91万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:7733737
-
项目类别:
-
资助金额:$318.2万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:8349580
-
项目类别:
-
资助金额:$324.22万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:8938250
-
项目类别:
-
资助金额:$341.1万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:7966676
-
项目类别:
-
资助金额:$301.69万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:8763630
-
项目类别:
-
资助金额:$317.95万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
Methods for Epidemiology Studies
-
批准号:8177710
-
项目类别:
-
资助金额:$350.55万
-
财政年份:--
-
负责人:Nilanjan Chatterjee
-
依托单位:
海外基金