Transformational machine learning: Learning how to learn from many related scientific problems.

Transformational machine learning: Learning how to learn from many related scientific problems.
复制标题

DOI:
10.1073/pnas.2108013118
复制
发表时间:
2021-12-07
影响因子:
11.1
通讯作者:
King RD
King RD
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Olier I;Orhobor OI;Dash T;Davis AM;Soldatova LN;Vanschoren J;King RD

文献摘要

参考文献

被引文献

相似文献

机器学习(ML)是人工智能(AI)的分支,它开发从经验中学习的计算系统。在监督ML中,ML系统从标记的示例中泛化,以学习可以预测未见过示例的标签的模型。示例通常使用直接描述示例的特征来表示。例如,在药物设计中,ML使用描述分子形状等的特征。在存在多个相关ML问题的情况下,可以使用不同类型的功能:通过在其他问题上学习的ML模型对示例进行预测。我们称之为转型ML。我们表明,这将导致更好的预测和更好的理解时,应用于科学问题。几乎所有的机器学习(ML)都是基于使用内在特征来表示示例。当存在多个相关的机器学习问题(任务)时,可以通过首先在其他任务上训练机器学习模型并让它们各自对新任务的每个示例进行预测,将这些特征转换为外部特征,从而产生一种新的表示。我们称之为转换ML(transformational ML)。TML与迁移学习、多任务学习和堆叠非常密切相关并具有协同作用。TML适用于改进任何非线性ML方法。我们使用最重要的非线性ML类别测试了TML:随机森林,梯度提升机,支持向量机,k-最近邻和神经网络。为了确保评估的通用性和鲁棒性,我们使用了来自三个科学领域的数千个ML问题:药物设计,预测基因表达和ML算法选择。我们发现,TML显著提高了所有领域中所有ML方法的预测性能(平均提高4%至50%),并且TML特征通常优于内在特征。TML的使用还通过可解释的ML增强了科学理解。在药物设计中,我们发现TML提供了对药物靶点特异性、药物之间关系以及靶蛋白之间关系的洞察。TML导致了一种基于生态系统的ML方法,其中新任务,示例,预测等协同交互以提高性能。为了对这个生态系统做出贡献,我们所有的数据、代码和50,000 ML模型都使用元数据进行了全面注释,并使用可查找性、可访问性、互操作性和可重用性原则(100 GB)进行了链接和公开发布。
Machine learning (ML) is the branch of artificial intelligence (AI) that develops computational systems that learn from experience. In supervised ML, the ML system generalizes from labelled examples to learn a model that can predict the labels of unseen examples. Examples are generally represented using features that directly describe the examples. For instance, in drug design, ML uses features that describe molecular shape and so on. In cases where there are multiple related ML problems, it is possible to use a different type of feature: predictions made about the examples by ML models learned on other problems. We call this transformational ML. We show that this results in better predictions and improved understanding when applied to scientific problems. Almost all machine learning (ML) is based on representing examples using intrinsic features. When there are multiple related ML problems (tasks), it is possible to transform these features into extrinsic features by first training ML models on other tasks and letting them each make predictions for each example of the new task, yielding a novel representation. We call this transformational ML (TML). TML is very closely related to, and synergistic with, transfer learning, multitask learning, and stacking. TML is applicable to improving any nonlinear ML method. We tested TML using the most important classes of nonlinear ML: random forests, gradient boosting machines, support vector machines, k-nearest neighbors, and neural networks. To ensure the generality and robustness of the evaluation, we utilized thousands of ML problems from three scientific domains: drug design, predicting gene expression, and ML algorithm selection. We found that TML significantly improved the predictive performance of all the ML methods in all the domains (4 to 50% average improvements) and that TML features generally outperformed intrinsic features. Use of TML also enhances scientific understanding through explainable ML. In drug design, we found that TML provided insight into drug target specificity, the relationships between drugs, and the relationships between target proteins. TML leads to an ecosystem-based approach to ML, where new tasks, examples, predictions, and so on synergistically interact to improve performance. To contribute to this ecosystem, all our data, code, and our ∼50,000 ML models have been fully annotated with metadata, linked, and openly published using Findability, Accessibility, Interoperability, and Reusability principles (∼100 Gbytes).
DOI: 10.1093/bioinformatics/bty287
发表时间: 2018-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Öztürk H;Ozkirimli E;Özgür A
通讯作者: Özgür A
DOI: 10.3389/fnhum.2017.00334
发表时间: 2017
影响因子: 2.9
作者:
Lin YP;Jung TP
通讯作者: Jung TP
DOI: 10.1186/s13321-020-00444-5
发表时间: 2020-06-05
影响因子: 8.6
作者:
Cortes-Ciriano, Isidro;Skuta, Ctibor;Svozil, Daniel
通讯作者: Svozil, Daniel
DOI: 10.1136/bmj.1.6070.1191
发表时间: 1977-01-01
影响因子: --
作者:
GHOSE, K;COPPEN, A;CARROLL, D
通讯作者: CARROLL, D
DOI: 10.1145/240455.240472
发表时间: 1996-11-01
影响因子: 22.7
作者:
Imielinski, T;Mannila, H
通讯作者: Mannila, H