CAREER: Active Learning through Rich and Transparent Interactions
CAREER: Active Learning through Rich and Transparent Interactions
批准号:
1350337
负责人:
Mustafa Bilgic
金额:
$54.99万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-05-01 至 2020-12-31
中文摘要
机器学习模型是在由人类注释(标记)的数据上进行训练的。训练后的模型的准确度通常会随着带注释的数据样本的数量而提高。然而,注释需要时间、金钱和精力。主动学习的目的是通过确定哪些例子最具信息量,并将人类标签者引向它们,从而将成本降至最低。主动学习的改进将降低与数据注释相关的成本,并导致智能系统的更快实现,这些应用包括机器人、语音技术、错误和异常检测(例如,在医学、金融欺诈和基础设施的基于条件的维护中)、定向广告、人机界面和生物信息学。在传统的主动学习方法中,算法受限于它们可以获取的信息类型,并且它们通常不向用户提供任何理由来解释为什么选择特定样本进行注释。这个职业项目开发了一种名为“丰富而透明的主动学习”的新范式。这一新的范例在算法和用户之间开辟了一条沟通渠道,他们可以借此交换一系列丰富的问题、答案和解释。通过使用来自用户的丰富反馈,算法将能够更经济地学习目标概念,减少建立准确预测模型所需的资源。通过解释他们的推理,这些算法将实现透明度,建立信任,并开放自己接受审查。为此,该项目开发了一些方法,允许算法使用丰富的查询集进行资源高效的模型训练,并生成信息丰富但不会让用户感到不知所措的解释。开发的方法建立在期望损失最小化、信息论和人机交互原理的基础上。使用公开可用的数据集和作为该项目的一部分进行的用户研究对方法进行评估。该项目开发了两个高影响力的现实世界问题的案例研究:检测欺诈性的医疗保健索赔,以及识别有疾病风险的患者。丰富和透明的主动学习范式提供了独特的教育机会。与以黑盒形式运作的标准机器学习算法不同,交互式和透明的机器学习预计将提高学生对数据科学的兴趣和动机。两名博士和几名本科生和高中生正在接受该奖项的培训。一门关于交互式机器学习的新研究生课程正在开发中。最后,PI通过与芝加哥一所公立高中合作,确保与代表不足的群体进行有效的接触,该高中的学生人口中有90%是少数族裔。
英文摘要
Machine learning models are trained on data that are annotated (labeled) by humans. The accuracy of the trained models generally improves with the number of annotated data examples. Yet, annotating takes time, money, and effort. Active learning aims to minimize the costs by determining which exemples are most informative and directing the human labeler to them. Improvements in active learning will lower the costs associated with data annotation and lead to faster implementations of intelligent systems for a range of applications including robotics, speech technology, error and anomaly detection (for example in medicine, financial fraud, and condition-based maintenance of infrastructure), targeted advertising, human-computer interfaces, and bioinformatics.In traditional active learning approaches, algorithms are limited in the types of information they can acquire, and they often do not provide any rationale to the user as to why a particular exemplar is chosen for annotation. This CAREER project develops a new paradigm dubbed "rich and transparent active learning." This new paradigm opens a communication channel between algorithms and users whereby they can exchange a rich set of queries, answers, and explanations. By using rich feedback from users the algorithms will be able to learn the target concept more economically, reducing the resources required to build an accurate predictive model. By explaining their reasoning, these algorithms will achieve transparency, build trust, and open themselves to scrutiny. Towards that end, the project develops methods that allow algorithms to use a rich set of queries for resource-efficient model training, and generate explanations that are informative but not overwhelming for the users. The methods developed build on expected loss minimization, information theory, and principles from human-computer interaction. Approaches are evaluated using publicly available datasets and user studies carried out as part of the project. The project develops case studies on two high-impact real-world problems: detecting fraudulent health-care claims, and identifying patients at risk of disease.The rich and transparent active learning paradigm provides unique educational opportunities. In contrast to standard machine learning algorithms, operated as black boxes, interactive and transparent machine learning is expected to raise students' interest and motivation for data science. Two PhD and several undergraduate and high school students are being trained under this award. A new graduate course on interactive machine learning is being developed. Finally the PI ensures effective outreach to under-represented groups by partnering with a Chicago public high school whose student population includes 90% minorities.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1145/3442381.3450113
发表时间:
2021-04
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
作者:
[Ping Liu;K. Shivaram;A. Culotta;Matthew A. Shapiro;M. Bilgic]
通讯作者:
Ping Liu;K. Shivaram;A. Culotta;Matthew A. Shapiro;M. Bilgic
EAGER:AI-DCL: Understanding the Relationship between Algorithmic Transparency and Filter Bubbles in Online Media
-
批准号:1927407
-
项目类别:Standard Grant
-
资助金额:$29.99万
-
财政年份:2019
-
负责人:Mustafa Bilgic
-
依托单位:
国内基金
海外基金
光-电驱动下的AIE-active手性高分子CPL液晶器件研究
-
批准号:92156014
-
项目类别:重大研究计划
-
资助金额:70.0万元
-
批准年份:2021
-
负责人:成义祥
-
依托单位:
光-电驱动下的AIE-active手性高分子CPL液晶器件研究
-
批准号:--
-
项目类别:--
-
资助金额:70万元
-
批准年份:2021
-
负责人:成义祥
-
依托单位: