Breast cancer prediction using genome wide single nucleotide polymorphism data.

Breast cancer prediction using genome wide single nucleotide polymorphism data.
复制标题

DOI:
10.1186/1471-2105-14-s13-s3
复制
发表时间:
2013
期刊:
影响因子:
3
通讯作者:
Damaraju S
Damaraju S
中科院分区:
生物学4区
文献类型:
--
作者:
Hajiloo M;Damavandi B;Hooshsadat M;Sangi F;Mackey JR;Cass CE;Greiner R;Damaraju S

文献摘要

被引文献

相似文献

本文介绍并应用了一项全基因组预测研究,以学习一种模型,该模型可以根据新受试者的SNP特征来预测该受试者是否会患上乳腺癌。我们首先使用Affymetrix Human SNP 6.0阵列对696名女性受试者(348名乳腺癌患者和348名明显健康的对照组)进行了基因分型,她们主要来自加拿大艾伯塔省的高加索血统。然后,我们应用EIGENSTRAT人口分层修正法剔除了73名不属于高加索人群的受试者。然后,我们筛选出有任何缺失呼叫的SNP,其基因频率偏离Hardy-Weinberg平衡,或其次要等位基因频率低于5%。最后,我们将MeanDiff特征选择方法和KNN学习方法相结合应用于这个过滤后的数据集,生成了乳腺癌预测模型。该分类器的LOOCV准确率为59.55%。随机排列测试表明,这一结果明显好于基线51.52%的准确率。敏感度分析表明,该分类器对MeanDiff选择的SNPs数目具有较强的鲁棒性。对CGEMS乳腺癌数据集的外部验证表明,MeanDiff和KNN的组合导致LOOCV准确率为60.25%,明显好于其基线50.06%。然后我们考虑了十几种不同的特征选择和学习方法的组合,但发现这些组合都不能产生比我们的模型更好的预测模型。我们还考虑了各种生物特征选择方法,如选择最近全基因组关联研究中报告的与乳腺癌相关的SNPs,选择与KEGG癌症通路相关的基因中的SNPs,或在F-SNP数据库中选择与乳腺癌相关的SNPs来生成预测模型,但再次发现这些模型都没有达到比基线更好的准确性。我们期望通过招募更多的研究对象,提供更准确的表型标记(以适应乳腺癌的异质性),测量其他基因组变化,如点突变和拷贝数变异,以及纳入关于研究对象的非遗传信息,如环境和生活方式因素,来生产更准确的乳腺癌预测模型。
This paper introduces and applies a genome wide predictive study to learn a model that predicts whether a new subject will develop breast cancer or not, based on her SNP profile. We first genotyped 696 female subjects (348 breast cancer cases and 348 apparently healthy controls), predominantly of Caucasian origin from Alberta, Canada using Affymetrix Human SNP 6.0 arrays. Then, we applied EIGENSTRAT population stratification correction method to remove 73 subjects not belonging to the Caucasian population. Then, we filtered any SNP that had any missing calls, whose genotype frequency was deviated from Hardy-Weinberg equilibrium, or whose minor allele frequency was less than 5%. Finally, we applied a combination of MeanDiff feature selection method and KNN learning method to this filtered dataset to produce a breast cancer prediction model. LOOCV accuracy of this classifier is 59.55%. Random permutation tests show that this result is significantly better than the baseline accuracy of 51.52%. Sensitivity analysis shows that the classifier is fairly robust to the number of MeanDiff-selected SNPs. External validation on the CGEMS breast cancer dataset, the only other publicly available breast cancer dataset, shows that this combination of MeanDiff and KNN leads to a LOOCV accuracy of 60.25%, which is significantly better than its baseline of 50.06%. We then considered a dozen different combinations of feature selection and learning method, but found that none of these combinations produces a better predictive model than our model. We also considered various biological feature selection methods like selecting SNPs reported in recent genome wide association studies to be associated with breast cancer, selecting SNPs in genes associated with KEGG cancer pathways, or selecting SNPs associated with breast cancer in the F-SNP database to produce predictive models, but again found that none of these models achieved accuracy better than baseline. We anticipate producing more accurate breast cancer prediction models by recruiting more study subjects, providing more accurate labelling of phenotypes (to accommodate the heterogeneity of breast cancer), measuring other genomic alterations such as point mutations and copy number variations, and incorporating non-genetic information about subjects such as environmental and lifestyle factors.