EAGER: Developing data and evaluation methods to assess the generality and robustness of AI systems for abstraction and analogy-making
EAGER: Developing data and evaluation methods to assess the generality and robustness of AI systems for abstraction and analogy-making
批准号:
2139983
负责人:
Melanie Mitchell
金额:
$19.97万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2024-02-29
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The ability of humans to make conceptual abstractions and analogies is at the root of many of our most important cognitive capabilities, such as learning new concepts from small numbers of examples, flexibly adapting our prior knowledge and experience to new situations, and communicating our knowledge to others. While AI has made dramatic progress over the last decade in areas such as vision, natural language processing, and robotics, current AI systems almost entirely lack the ability to form humanlike abstractions and analogies. The lack of such abilities is in part responsible for the lack of robustness in current AI systems, as well as their difficulties with extrapolating what they have learned to diverse situations. While there have been many efforts in past AI research on this topic, each individual effort has generally focused on a specific problem domain, without careful evaluation of the AI system’s robustness within its domain or its generality across different domains. In this project we will promote progress in AI by creating a web-based platform that offers a diverse set of abstraction and analogy-making challenges for the research community as well as new evaluation methods that test for generality and robustness within and across different challenge domains. We will use our platform to evaluate selected existing AI approaches and to measure human performance on our challenges in order to compare with AI systems’ performance. Our work will contribute to the AI research community by spurring new approaches and evaluation methods for abstraction and analogy-making in machines, and will contribute more broadly via the development of methods for robust and generalizable AI systems. Our specific research plan is to (1) curate an initial suite of idealized challenge domains inspired by Hofstadter’s letter-string analogies, Raven’s progressive matrices, Bongard problems, and Chollet’s Abstraction and Reasoning Corpus; (2) develop evaluation methods along dimensions such as robustness to variations on a particular concept, generality across domains, and scalability to more complex instances of a problem; (3) evaluate selected AI methods for abstraction and analogy using our evaluation methods; and (4) measure human benchmarks on our challenge suite using paid participants on the Amazon Mechanical Turk platform. At the end of the project period, we will have demonstrated the utility and promise of our challenge problems and evaluations, and will have gained insight into their limitations. This work will set the stage for future efforts on expanding our challenge suite, improving our evaluation metrics, and developing and evaluating novel AI approaches to abstraction and analogy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2206.14187
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
作者:
[Victor Vikram Odouard;M. Mitchell]
通讯作者:
Victor Vikram Odouard;M. Mitchell
How do we know how smart AI systems are?
我们如何知道人工智能系统有多智能?
DOI:
10.1126/science.adj5957
发表时间:
2023
期刊:
Science
影响因子:
56.9
作者:
[Mitchell, Melanie]
通讯作者:
Mitchell, Melanie
Rethink reporting of evaluation results in AI
重新思考人工智能评估结果的报告
DOI:
10.1126/science.adf6369
发表时间:
2023
期刊:
Science
影响因子:
56.9
作者:
[Burnell, Ryan, Schellaert, Wout, Burden, John, Ullman, Tomer D., Martinez-Plumed, Fernando, Tenenbaum, Joshua B., Rutar, Danaja, Cheke, Lucy G., Sohl-Dickstein, Jascha, Mitchell, Melanie]
通讯作者:
Mitchell, Melanie
The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain
ConceptARC 基准:评估 ARC 领域的理解和泛化
DOI:
--
发表时间:
2023
期刊:
Transactions on machine learning research
影响因子:
--
作者:
[Moskvichev, Arsenii Kirillovich, Odouard, Victor Vikram, Mitchell, Melanie]
通讯作者:
Mitchell, Melanie
AI Institute: Planning: Foundations of Intelligence in Natural and Artificial Systems
-
批准号:2020103
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2020
-
负责人:Melanie Mitchell
-
依托单位:
Workshop on Artificial Intelligence and the "Barrier of Meaning"
-
批准号:1832717
-
项目类别:Standard Grant
-
资助金额:$2.01万
-
财政年份:2018
-
负责人:Melanie Mitchell
-
依托单位:
RI: Small: Visual Situation Recognition: An Integration of Deep Networks and Analogy-Making
-
批准号:1423651
-
项目类别:Standard Grant
-
资助金额:$44.98万
-
财政年份:2014
-
负责人:Melanie Mitchell
-
依托单位:
RI: Small: Collaborative Research: A Scalable Architecture for Image Interpretation
-
批准号:1018967
-
项目类别:Standard Grant
-
资助金额:$34.13万
-
财政年份:2010
-
负责人:Melanie Mitchell
-
依托单位:
Evolving Cellular Automata: With Genetic Algorithms
-
批准号:9705830
-
项目类别:Continuing Grant
-
资助金额:$29.78万
-
财政年份:1998
-
负责人:Melanie Mitchell
-
依托单位:
Postdoc: Automatic Programming of Decentralized Parallel Architectures
-
批准号:9503162
-
项目类别:Standard Grant
-
资助金额:$2.56万
-
财政年份:1995
-
负责人:Melanie Mitchell
-
依托单位:
海外基金