Computer Vision and Conflicting Values: Describing People with Automated Alt Text.

Computer Vision and Conflicting Values: Describing People with Automated Alt Text.
复制标题

计算机视觉和冲突的价值观:用自动替代文本描述人。

DOI:
10.1145/3461702.3462620
复制
发表时间:
2021
期刊:
and Society (AIES
影响因子:
--
通讯作者:
Hanley, M.
Hanley, M.
中科院分区:
--
文献类型:
--
作者:
Hanley, M.

文献摘要

相似文献

学者们最近引起了人们对使用计算机视觉自动生成图像中人物描述所带来的一系列有争议的问题的关注。尽管存在这些担忧,自动图像描述已成为确保盲人和低视力人群公平获取信息的重要工具。在本文中,我们调查了采用计算机视觉来生成替代文本的公司所面临的道德困境:为盲人和低视力人士提供图像的文本描述。我们使用 Facebook 的自动替代文本工具作为我们的主要案例研究。首先,我们分析 Facebook 针对身份类别(例如种族、性别、年龄等)采取的政策,以及公司关于是否在替代文本中呈现这些术语的决定。然后,我们描述了博物馆社区中实践的替代方法和手动方法,重点关注博物馆如何确定在文化文物的替代文本描述中包含哪些内容。我们比较这些政策,利用显着的对比点来开发一个分析框架,描述这些政策选择背后的特定担忧。最后,我们考虑了两种似乎可以回避其中一些问题的策略,发现没有简单的方法可以避免使用计算机视觉自动化替代文本所带来的规范困境。
Scholars have recently drawn attention to a range of controversial issues posed by the use of computer vision for automatically generating descriptions of people in images. Despite these concerns, automated image description has become an important tool to ensure equitable access to information for blind and low vision people. In this paper, we investigate the ethical dilemmas faced by companies that have adopted the use of computer vision for producing alt text: textual descriptions of images for blind and low vision people. We use Facebook's automatic alt text tool as our primary case study. First, we analyze the policies that Facebook has adopted with respect to identity categories, such as race, gender, age, etc., and the company's decisions about whether to present these terms in alt text. We then describe an alternative---and manual---approach practiced in the museum community, focusing on how museums determine what to include in alt text descriptions of cultural artifacts. We compare these policies, using notable points of contrast to develop an analytic framework that characterizes the particular apprehensions behind these policy choices. We conclude by considering two strategies that seem to sidestep some of these concerns, finding that there are no easy ways to avoid the normative dilemmas posed by the use of computer vision to automate alt text.