Kaleidoscope: Semantically-grounded, context-specific ML model evaluation

Kaleidoscope: Semantically-grounded, context-specific ML model evaluation
复制标题

DOI:
10.1145/3544548.3581482
复制
发表时间:
2023-04
期刊:
Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
影响因子:
--
通讯作者:
Harini Suresh;Divya Shanmugam;Tiffany Chen;Annie G Bryan;A. D'Amour;John Guttag;Arvindmani Satyanarayan
Harini Suresh;Divya Shanmugam;Tiffany Chen;Annie G Bryan;A. D'Amour;John Guttag;Arvindmani Satyanarayan
中科院分区:
其他
文献类型:
--
作者:
Harini Suresh;Divya Shanmugam;Tiffany Chen;Annie G Bryan;A. D'Amour;John Guttag;Arvindmani Satyanarayan

文献摘要

相似文献

所需的模型行为通常因环境(例如,不同的地理位置、社区或机构)而异,但几乎没有基础设施来促进对部署决策和建立信任至关重要的特定于环境的评估。在这里,我们介绍了万花筒,这是一个根据用户驱动的、领域相关的概念来评估模型的系统。万花筒的迭代工作流程可以将几个例子概括为一个更大、更多样化的集合,代表一个重要的概念。这些示例集可用于以语义有意义的方式测试模型输出或模型行为的转变。例如,我们可以构建一个“排外评论”集,并测试它的例子更有可能被内容审核模型标记,而不是“民间讨论”集。为了评估万花筒,我们将其与基于模板和基于DSL的分组方法进行了比较,并对13名Reddit用户进行了可用性研究,测试了一个内容审核模型。我们发现,万花筒便于在不同的、概念上有意义的示例集上进行迭代的探索性假设测试。
Desired model behavior often differs across contexts (e.g., different geographies, communities, or institutions), but there is little infrastructure to facilitate context-specific evaluations key to deployment decisions and building trust. Here, we present Kaleidoscope, a system for evaluating models in terms of user-driven, domain-relevant concepts. Kaleidoscope’s iterative workflow enables generalizing from a few examples into a larger, diverse set representing an important concept. These example sets can be used to test model outputs or shifts in model behavior in semantically-meaningful ways. For instance, we might construct a “xenophobic comments” set and test that its examples are more likely to be flagged by a content moderation model than a “civil discussion” set. To evaluate Kaleidoscope, we compare it against template- and DSL-based grouping methods, and conduct a usability study with 13 Reddit users testing a content moderation model. We find that Kaleidoscope facilitates iterative, exploratory hypothesis testing across diverse, conceptually-meaningful example sets.