Doctoral Dissertation Research: Evaluating the Promise and Pitfalls of Benchmarking in Machine Learning Research
Doctoral Dissertation Research: Evaluating the Promise and Pitfalls of Benchmarking in Machine Learning Research
批准号:
2124685
负责人:
Jacob Foster
金额:
$2.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2023-07-31
中文摘要
该奖项全部或部分由《2021年美国救援计划法案》(公法117-2)资助。机器学习(ML)在科学和商业上的成功促使政府和企业赞助商在机器学习研究上投入数十亿美元。尽管投入巨大,但关于机器学习领域如何衡量进展的定量研究有限:一个被称为“基准测试”的过程。基准测试是在相同的基准数据集上训练算法后,在定量指标上比较算法的行为。基准测试围绕常见任务组织ML研究人员。在一个重要的基准上实现“最先进”的表现可以激发新的研究轨迹和职业发展:想想2012年“AlexNet”在一个突出的计算机视觉任务中的成功,它帮助引发了当前对深度学习的兴趣。然而,基准测试的实践已经引起了批评,这种几乎无处不在的研究文化并没有推动该领域走向对社会有益的结果,并导致对学术数据集性能最大化的方法的过度投资,但在现实世界中使用时,这些方法在环境上是不可持续的,或者会伤害公众。本论文研究将全面分析基准实践的优势和劣势,涉及几个公共目标:加速科学创新,增加领域内的公平,促进伦理研究(即,朝向有利于社会和避免危害的研究)。通过混合社会学分析,从数千篇论文中提取和分析基准数据的计算方法,以及深入的定性访谈,本研究将产生对ML研究中基准文化的理解,将广度和定量严谨性与深度和解释性的细微差别相结合。该项目对政府和企业资助者、研究人员以及更广泛的社会具有重要意义。本论文由三个子课题组成。第一个子项目探讨了基准文化阻碍创新的证据,因为它倾向于在多个任务中使用相同的数据集,并鼓励研究人员在新兴基准上投资不足,而在成熟基准上投资过多。第二个子项目探讨采用基准和奖励最先进业绩的模式如何与地位和资源相互作用,从而在该领域造成不平等。它检验了一个假设,即高地位的研究人员和机构通过引入基准,在设定该领域的研究议程方面拥有不成比例的权力,同时在最先进的成就上获得不成比例的引用。这两种现象都有可能产生“马太效应”,使代表性不足和资源不足的研究人员/机构处于不利地位。这些子项目使用网络科学、自然语言处理和手动编码来创建大型基准数据集,并跨多个ML任务社区在这些基准上取得进展。第三个子项目包括对跨职业阶段和专业知识的ML研究人员进行定性访谈,以获得对基准文化的第一手看法,并评估改革以改善研究伦理和社会成果。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).The scientific and commercial success of machine learning (ML) has spurred government and corporate sponsors to invest billions of dollars in machine learning research. Despite this massive investment, there is limited quantitative research on how the ML field measures progress: a process called “benchmarking.” Benchmarking is the act of comparing algorithms on a quantitative metric after training them on the same benchmark dataset. Benchmarks organize ML researchers around common tasks. Achieving “state of the art” performance on an important benchmark can spark new research trajectories and advance careers: consider the 2012 success of “AlexNet” in a prominent computer vision task, which helped to launch current interest in deep learning. However, the practice of benchmarking has already engendered criticism that this near-ubiquitous research culture does not push the field towards socially beneficial outcomes, and leads to overinvestment in methods that maximize performance on academic datasets but are environmentally unsustainable or harm the public when used in the real world. This dissertation research will provide a comprehensive analysis of the strengths and weaknesses of benchmarking practices with respect to several public aims: accelerating innovation in science, increasing equity within the field, and promoting ethical research (i.e., an orientation toward research that benefits society and avoids harms). By blending sociological analysis, computational methods for extracting and analyzing benchmarking data from thousands of papers, and in-depth qualitative interviews, this research will produce an understanding of benchmarking culture in ML research that combines breadth and quantitative rigor with depth and interpretive nuance. This project has significant implications for government and corporate funders, researchers, and society more broadly. The dissertation consists of three subprojects. The first subproject explores evidence that benchmarking culture has stymied innovation by favoring utilization of the same datasets across multiple tasks and by incentivizing researchers to underinvest on nascent benchmarks and overinvest on mature ones. The second subproject explores how patterns in the adoption of benchmarks and rewards for state-of-the-art performance interact with status and resources to create inequities in the field. It tests the hypothesis that high-status researchers and institutions have disproportionate power to set the field’s research agenda by introducing benchmarks, while garnering disproportionate citations for state-of-the-art achievements. Both of these phenomena have the potential to create a “Matthew Effect” that disadvantages under-represented and under-resourced researchers/institutions. These subprojects use network science, natural language processing, and manual coding to create a large dataset of benchmarks and progress on those benchmarks across multiple ML task communities. The third subproject consists of qualitative interviews with ML researchers across career stages and expertise to gain first-hand perspectives on benchmarking culture and assess reforms to improve research ethics and societal outcomes.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
--
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
作者:
[Bernard Koch;Emily L. Denton;A. Hanna;J. Foster]
通讯作者:
Bernard Koch;Emily L. Denton;A. Hanna;J. Foster
海外基金