Machine Learning Strategies for Improved Phenotype Prediction in Underrepresented Populations.

Machine Learning Strategies for Improved Phenotype Prediction in Underrepresented Populations.
复制标题

DOI:
10.1101/2023.10.12.561949
复制
发表时间:
2023-10-17
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Ioannidis AG
Ioannidis AG
中科院分区:
其他
文献类型:
--
作者:
Bonet D;Levin M;Montserrat DM;Ioannidis AG

文献摘要

相似文献

精确医学模型通常对欧洲血统的人群表现更好,因为这一群体在基因组数据集和构建模型的大规模生物库中有过多的代表性。因此,预测模型可能会歪曲或为代表性不足的人群提供不太准确的治疗建议,从而造成健康差距。本研究介绍了一种适应性强的机器学习工具包,该工具包集成了多种现有方法和新技术,以提高基因组数据集中代表性不足人群的预测准确性。通过利用机器学习技术,包括梯度增强和自动化方法,再加上新的种群条件重采样技术,我们的方法显着改善了来自不同种群的单核苷酸多态性(SNP)数据的表型预测。我们使用英国生物银行来评估我们的方法,该银行主要由具有欧洲血统的英国人组成,以及具有亚洲和非洲血统的少数群体代表。性能指标显示,对代表性不足的群体的表型预测有了实质性的改进,预测精度与多数群体相当。这种方法代表了在当前数据集多样性挑战中提高预测精度的重要一步。通过整合量身定制的管道,我们的方法促进了统计遗传学方法更公平的有效性和实用性,为更具包容性的模型和结果铺平了道路。
Precision medicine models often perform better for populations of European ancestry due to the over-representation of this group in the genomic datasets and large-scale biobanks from which the models are constructed. As a result, prediction models may misrepresent or provide less accurate treatment recommendations for underrepresented populations, contributing to health disparities. This study introduces an adaptable machine learning toolkit that integrates multiple existing methodologies and novel techniques to enhance the prediction accuracy for underrepresented populations in genomic datasets. By leveraging machine learning techniques, including gradient boosting and automated methods, coupled with novel population-conditional re-sampling techniques, our method significantly improves the phenotypic prediction from single nucleotide polymorphism (SNP) data for diverse populations. We evaluate our approach using the UK Biobank, which is composed primarily of British individuals with European ancestry, and a minority representation of groups with Asian and African ancestry. Performance metrics demonstrate substantial improvements in phenotype prediction for underrepresented groups, achieving prediction accuracy comparable to that of the majority group. This approach represents a significant step towards improving prediction accuracy amidst current dataset diversity challenges. By integrating a tailored pipeline, our approach fosters more equitable validity and utility of statistical genetics methods, paving the way for more inclusive models and outcomes.