Hands-on training about overfitting.

Hands-on training about overfitting.
复制标题

DOI:
10.1371/journal.pcbi.1008671
复制
发表时间:
2021-03
影响因子:
4.3
通讯作者:
Zupan B
Zupan B
中科院分区:
生物学2区
文献类型:
--
作者:
Demšar J;Zupan B

文献摘要

参考文献

被引文献

相似文献

过拟合是机器学习开发模型的关键问题之一。随着机器学习成为计算生物学中的一项重要技术,我们必须在向学生和从业者介绍这项技术的所有课程中包括关于过拟合的培训。我们在这里提出了一个适合入门级课程的过拟合实践培训,可以单独进行或嵌入任何数据科学课程。我们使用基于工作流的机器学习管道设计,基于实验的教学,以及专注于概念而不是基础数学的实践方法。我们在这里详细介绍了我们在培训中使用的数据分析工作流程,并从教学目标的角度激励他们。我们提出的方法依赖于橙子,这是一个开源的数据科学工具箱,它结合了数据可视化和机器学习,并为机器学习和探索性数据分析的教育量身定制。每个老师都在努力寻找一个顿悟的时刻,一个学生获得了她永远记得的基本见解的突然启示。在过去的几年里,本文的作者一直在调整他们的机器学习课程,以包括可能导致学生发现这些发现的材料。我们的目标是将机器学习暴露给计算机科学家,不仅是计算机科学家,还有分子生物学家和生物医学的学生,也就是生物信息学计算方法的最终用户。在本文中,我们设计了一门课程,旨在教授过拟合,这是机器学习中的关键概念之一,需要在数据科学应用中理解,掌握和避免。我们提出了一种实践方法,该方法使用基于开源工作流的数据科学工具箱,该工具箱将数据可视化和机器学习相结合。在提出的过拟合训练中,我们首先欺骗学生,然后暴露问题,最后挑战他们找到解决方案。在本文中,我们提出了过拟合和相关的数据分析工作流程的三个教训,并通过将它们与教师传达的概念联系起来来激励使用引入的计算方法。
Overfitting is one of the critical problems in developing models by machine learning. With machine learning becoming an essential technology in computational biology, we must include training about overfitting in all courses that introduce this technology to students and practitioners. We here propose a hands-on training for overfitting that is suitable for introductory level courses and can be carried out on its own or embedded within any data science course. We use workflow-based design of machine learning pipelines, experimentation-based teaching, and hands-on approach that focuses on concepts rather than underlying mathematics. We here detail the data analysis workflows we use in training and motivate them from the viewpoint of teaching goals. Our proposed approach relies on Orange, an open-source data science toolbox that combines data visualization and machine learning, and that is tailored for education in machine learning and explorative data analysis. Every teacher strives for an a-ha moment, a sudden revelation by the student who gained a fundamental insight she will always remember. In the past years, authors of this paper have been tailoring their courses in machine learning to include material that could lead students to such discoveries. We aim to expose machine learning to practitioners–not only computer scientists but also molecular biologists and students of biomedicine, that is, the end-users of bioinformatics’ computational approaches. In this article, we lay out a course that aims to teach about overfitting, one of the key concepts in machine learning that needs to be understood, mastered, and avoided in data science applications. We propose a hands-on approach that uses an open-source workflow-based data science toolbox that combines data visualization and machine learning. In the proposed training about overfitting, we first deceive the students, then expose the problem, and finally challenge them to find the solution. In the paper, we present three lessons in overfitting and associated data analysis workflows and motivate the use of introduced computation methods by relating them to concepts conveyed by instructors.
DOI: 10.1073/pnas.97.1.262
发表时间: 2000-01-04
影响因子: 11.1
作者:
Brown, MPS;Grundy, WN;Haussler, D
通讯作者: Haussler, D
DOI: 10.1093/bioinformatics/bth474
发表时间: 2005-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Curk, T;Demsar, J;Zupan, B
通讯作者: Zupan, B
DOI: 10.1186/s13040-017-0155-3
发表时间: 2017
期刊: BioData mining
影响因子: 4.5
作者:
Chicco D
通讯作者: Chicco D
DOI: 10.1200/jco.2005.03.156
发表时间: 2005-02-20
影响因子: 45.3
作者:
Chang, JC;Wooten, EC;O'Connell, P
通讯作者: O'Connell, P
DOI: 10.1016/j.inffus.2018.09.012
发表时间: 2019-10
期刊: An international journal on information fusion
影响因子: --
作者:
Zitnik M;Nguyen F;Wang B;Leskovec J;Goldenberg A;Hoffman MM
通讯作者: Hoffman MM