Multi-modality machine learning predicting Parkinson's disease.
Multi-modality machine learning predicting Parkinson's disease.
复制标题
多模式机器学习预测帕金森氏病。
DOI:
10.1038/s41531-022-00288-w
复制
发表时间:
2022-04-01
期刊:
影响因子:
--
通讯作者:
Nalls MA
中科院分区:
文献类型:
--
作者:
Makarious MB;Leonard HL;Vitale D;Iwaki H;Sargent L;Dadu A;Violich I;Hutchins E;Saffo D;Bandres-Ciga S;Kim JJ;Song Y;Maleknia M;Bookman M;Nojopranoto W;Campbell RH;Hashemi SH;Botia JA;Carter JF;Craig DW;Van Keuren-Jensen K;Morris HR;Hardy JA;Blauwendraat C;Singleton AB;Faghri F;Nalls MA
Personalized medicine promises individualized disease prediction and treatment. The convergence of machine learning (ML) and available multimodal data is key moving forward. We build upon previous work to deliver multimodal predictions of Parkinson’s disease (PD) risk and systematically develop a model using GenoML, an automated ML package, to make improved multi-omic predictions of PD, validated in an external cohort. We investigated top features, constructed hypothesis-free disease-relevant networks, and investigated drug–gene interactions. We performed automated ML on multimodal data from the Parkinson’s progression marker initiative (PPMI). After selecting the best performing algorithm, all PPMI data was used to tune the selected model. The model was validated in the Parkinson’s Disease Biomarker Program (PDBP) dataset. Our initial model showed an area under the curve (AUC) of 89.72% for the diagnosis of PD. The tuned model was then tested for validation on external data (PDBP, AUC 85.03%). Optimizing thresholds for classification increased the diagnosis prediction accuracy and other metrics. Finally, networks were built to identify gene communities specific to PD. Combining data modalities outperforms the single biomarker paradigm. UPSIT and PRS contributed most to the predictive power of the model, but the accuracy of these are supplemented by many smaller effect transcripts and risk SNPs. Our model is best suited to identifying large groups of individuals to monitor within a health registry or biobank to prioritize for further testing. This approach allows complex predictive models to be reproducible and accessible to the community, with the package, code, and results publicly available.
登录
查看更多内容
影响因子:
64.8
作者:
Green ED;Gunter C;Biesecker LG;Di Francesco V;Easter CL;Feingold EA;Felsenfeld AL;Kaufman DJ;Ostrander EA;Pavan WJ;Phillippy AM;Wise AL;Dayal JG;Kish BJ;Mandich A;Wellington CR;Wetterstrand KA;Bates SA;Leja D;Vasquez S;Gahl WA;Graham BJ;Kastner DL;Liu P;Rodriguez LL;Solomon BD;Bonham VL;Brody LC;Hutter CM;Manolio TA
通讯作者:
Manolio TA
影响因子:
4.6
作者:
Haynes, Winston A.;Tomczak, Aurelie;Khatri, Purvesh
通讯作者:
Khatri, Purvesh
影响因子:
3.5
作者:
Hill-Burns EM;Ross OA;Wissemann WT;Soto-Ortolaza AI;Zareparsi S;Siuda J;Lynch T;Wszolek ZK;Silburn PA;Mellick GD;Ritz B;Scherzer CR;Zabetian CP;Factor SA;Breheny PJ;Payami H
通讯作者:
Payami H
影响因子:
8.8
作者:
Fernandes, Hugo J. R.;Patikas, Nikolaos;Metzakopian, Emmanouil
通讯作者:
Metzakopian, Emmanouil
影响因子:
7.5
作者:
Geurts, P;Ernst, D;Wehenkel, L
通讯作者:
Wehenkel, L