A Comparative Analysis of Emotion-Detecting AI Systems with Respect to Algorithm Performance and Dataset Diversity

A Comparative Analysis of Emotion-Detecting AI Systems with Respect to Algorithm Performance and Dataset Diversity
复制标题

DOI:
10.1145/3306618.3314284
复制
发表时间:
2019-01
期刊:
Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
De'Aira G. Bryant;A. Howard
De'Aira G. Bryant;A. Howard
中科院分区:
其他
文献类型:
--
作者:
De'Aira G. Bryant;A. Howard

文献摘要

被引文献

相似文献

在最近的新闻中,组织一直在考虑将面部和情感识别用于涉及青年的应用,例如解决学校的监控和安全问题。然而,面部情绪识别研究的大部分努力都集中在成年人身上。儿童,特别是在他们的早年,已经被证明表达情感与成人完全不同。因此,在将此类算法部署到影响青年福祉和环境的环境中之前,应仔细检查其准确性是否适合这一目标人口。在这项工作中,我们利用几个数据集,其中包含与他们的情绪状态相关联的儿童的面部表情,以评估八个不同的商业情绪分类系统。我们将相应数据集提供的地面真值标签与分类系统给出的具有最高置信度的标签进行比较,并根据匹配分数(TPR)、阳性预测值和计算失败率评估结果。总体结果表明,与先前使用成人数据集和初始人类评级的工作相比,情感识别系统在儿童表情数据集上显示出低于标准的性能。然后,我们确定了与儿童情绪自动识别相关的限制,并通过数据多样化,数据集问责制和算法监管来提高识别准确性。
In recent news, organizations have been considering the use of facial and emotion recognition for applications involving youth such as tackling surveillance and security in schools. However, the majority of efforts on facial emotion recognition research have focused on adults. Children, particularly in their early years, have been shown to express emotions quite differently than adults. Thus, before such algorithms are deployed in environments that impact the wellbeing and circumstance of youth, a careful examination should be made on their accuracy with respect to appropriateness for this target demographic. In this work, we utilize several datasets that contain facial expressions of children linked to their emotional state to evaluate eight different commercial emotion classification systems. We compare the ground truth labels provided by the respective datasets to the labels given with the highest confidence by the classification systems and assess the results in terms of matching score (TPR), positive predictive value, and failure to compute rate. Overall results show that the emotion recognition systems displayed subpar performance on the datasets of children's expressions compared to prior work with adult datasets and initial human ratings. We then identify limitations associated with automated recognition of emotions in children and provide suggestions on directions with enhancing recognition accuracy through data diversification, dataset accountability, and algorithmic regulation.