Systematic auditing is essential to debiasing machine learning in biology.

Systematic auditing is essential to debiasing machine learning in biology.
复制标题

DOI:
10.1038/s42003-021-01674-5
复制
发表时间:
2021-02-10
影响因子:
5.9
通讯作者:
Lage K
Lage K
中科院分区:
生物学2区
文献类型:
--
作者:
Eid FE;Elmarakeby HA;Chan YA;Fornelos N;ElHefnawi M;Van Allen EM;Heath LS;Lage K

文献摘要

参考文献

被引文献

相似文献

用于训练机器学习 (ML) 模型的数据偏差可能会夸大其预测性能,并混淆我们对其学习方式和学习内容的理解。尽管偏差在生物数据中很常见,但在生命科学中应用机器学习时,对机器学习模型进行系统审核以识别和消除这些偏差并不常见。在这里,我们设计了一种系统的、有原则的、通用的方法来审计生命科学中的机器学习模型。我们使用此审计框架来检查三个具有治疗意义的机器学习应用程序中的偏差,并识别阻碍机器学习过程并导致新数据集上的模型性能大幅降低的未识别偏差。最终,我们表明,当数据中没有足够的信号可供学习时,机器学习模型往往主要从数据偏差中学习。我们提供详细的协议、指南和代码示例,以便能够根据其他生物医学应用定制审核框架。法特玛-埃尔扎拉开斋节等。说明了一种识别可能夸大生物机器学习模型性能的偏差的原则方法。当应用于三个生物医学预测问题时,他们识别出以前未识别的偏差,并最终表明,当数据中可学习信号不足时,模型可能主要从数据偏差中学习。
Biases in data used to train machine learning (ML) models can inflate their prediction performance and confound our understanding of how and what they learn. Although biases are common in biological data, systematic auditing of ML models to identify and eliminate these biases is not a common practice when applying ML in the life sciences. Here we devise a systematic, principled, and general approach to audit ML models in the life sciences. We use this auditing framework to examine biases in three ML applications of therapeutic interest and identify unrecognized biases that hinder the ML process and result in substantially reduced model performance on new datasets. Ultimately, we show that ML models tend to learn primarily from data biases when there is insufficient signal in the data to learn from. We provide detailed protocols, guidelines, and examples of code to enable tailoring of the auditing framework to other biomedical applications. Fatma-Elzahraa Eid et al. illustrate a principled approach for identifying biases that can inflate the performance of biological machine learning models. When applied to three biomedical prediction problems, they identify previously unrecognized biases and ultimately show that models are likely to learn primarily from data biases when there is insufficient learnable signal in the data.
DOI: 10.4049/jimmunol.1700893
发表时间: 2017-11-01
期刊: Journal of immunology (Baltimore, Md. : 1950)
影响因子: --
作者:
Jurtz V;Paul S;Andreatta M;Marcatili P;Peters B;Nielsen M
通讯作者: Nielsen M
DOI: 10.1093/bioinformatics/btr514
发表时间: 2011-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Park, Yungki;Marcotte, Edward M.
通讯作者: Marcotte, Edward M.
DOI: 10.1038/nmeth.4627
发表时间: 2018-04
期刊: Nature methods
影响因子: 48
作者:
Ma J;Yu MK;Fong S;Ono K;Sage E;Demchak B;Sharan R;Ideker T
通讯作者: Ideker T
DOI: 10.1186/1471-2105-10-394
发表时间: 2009-11-30
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim Y;Sidney J;Pinilla C;Sette A;Peters B
通讯作者: Peters B
DOI: 10.1016/j.cell.2014.10.050
发表时间: 2014-11-20
期刊: Cell
影响因子: 64.5
作者:
Rolland T;Taşan M;Charloteaux B;Pevzner SJ;Zhong Q;Sahni N;Yi S;Lemmens I;Fontanillo C;Mosca R;Kamburov A;Ghiassian SD;Yang X;Ghamsari L;Balcha D;Begg BE;Braun P;Brehme M;Broly MP;Carvunis AR;Convery-Zupan D;Corominas R;Coulombe-Huntington J;Dann E;Dreze M;Dricot A;Fan C;Franzosa E;Gebreab F;Gutierrez BJ;Hardy MF;Jin M;Kang S;Kiros R;Lin GN;Luck K;MacWilliams A;Menche J;Murray RR;Palagi A;Poulin MM;Rambout X;Rasla J;Reichert P;Romero V;Ruyssinck E;Sahalie JM;Scholz A;Shah AA;Sharma A;Shen Y;Spirohn K;Tam S;Tejeda AO;Wanamaker SA;Twizere JC;Vega K;Walsh J;Cusick ME;Xia Y;Barabási AL;Iakoucheva LM;Aloy P;De Las Rivas J;Tavernier J;Calderwood MA;Hill DE;Hao T;Roth FP;Vidal M
通讯作者: Vidal M