A mixed-model approach for powerful testing of genetic associations with cancer risk incorporating tumor characteristics.

A mixed-model approach for powerful testing of genetic associations with cancer risk incorporating tumor characteristics.
复制标题

一种混合模型方法,可结合肿瘤特征对与癌症风险的遗传关联进行强有力的测试。

DOI:
10.1093/biostatistics/kxz065
复制
发表时间:
2021
期刊:
Biostatistics (Oxford, England)
影响因子:
--
通讯作者:
Chatterjee,Nilanjan
Chatterjee,Nilanjan
中科院分区:
--
文献类型:
--
作者:
Zhang,Haoyu;Zhao,Ni;Ahearn,ThomasU;Wheeler,William;García-Closas,Montserrat;Chatterjee,Nilanjan

文献摘要

相似文献

癌症通常根据各种特征(包括组织病理学特征和分子标志物)分为亚型。之前的全基因组关联研究已经报告了基因座与癌症亚型之间的异质关联。然而,目前尚不清楚什么是最佳的建模策略,用于处理相关的肿瘤特征,缺失的数据,并在相关性的基础测试中增加自由度。我们建议使用混合效应两阶段多分类模型评分检验(MTOP)来测试遗传关联。在第一阶段,使用标准的多分类模型来指定由肿瘤特征的交叉分类定义的所有可能的亚型。在第二阶段,使用基于基线亚型的病例-对照比值比和与肿瘤标志物相关的病例-病例参数的更简约模型来指定亚型特异性病例-对照比值比。此外,为了减少自由度,我们使用随机效应模型为其他探索性标志物指定病例-病例参数。我们使用期望最大化算法来解释肿瘤标志物的缺失数据。通过对波兰乳腺癌研究(PBCS)的一系列现实场景和数据的模拟,我们表明MTOP在识别风险位点和肿瘤亚型之间异质性关联方面优于其他方法。所提出的方法已经在一个名为TOP(https://github.com/andrewhaoyu/TOP)的用户友好和高速R统计包中实现。
Cancers are routinely classified into subtypes according to various features, including histopathological characteristics and molecular markers. Previous genome-wide association studies have reported heterogeneous associations between loci and cancer subtypes. However, it is not evident what is the optimal modeling strategy for handling correlated tumor features, missing data, and increased degrees-of-freedom in the underlying tests of associations. We propose to test for genetic associations using a mixed-effect two-stage polytomous model score test (MTOP). In the first stage, a standard polytomous model is used to specify all possible subtypes defined by the cross-classification of the tumor characteristics. In the second stage, the subtype-specific case–control odds ratios are specified using a more parsimonious model based on the case–control odds ratio for a baseline subtype, and the case–case parameters associated with tumor markers. Further, to reduce the degrees-of-freedom, we specify case–case parameters for additional exploratory markers using a random-effect model. We use the Expectation–Maximization algorithm to account for missing data on tumor markers. Through simulations across a range of realistic scenarios and data from the Polish Breast Cancer Study (PBCS), we show MTOP outperforms alternative methods for identifying heterogeneous associations between risk loci and tumor subtypes. The proposed methods have been implemented in a user-friendly and high-speed R statistical package called TOP (https://github.com/andrewhaoyu/TOP)