Visual Cues for Disrespectful Conversation Analysis

Visual Cues for Disrespectful Conversation Analysis
复制标题

不尊重谈话分析的视觉线索

DOI:
10.1109/acii.2019.8925440
复制
发表时间:
2019
期刊:
2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII)
影响因子:
--
通讯作者:
Ehsan Hoque
Ehsan Hoque
中科院分区:
--
文献类型:
--
作者:
Samiha Samrose;Wenyi Chu;C. He;Yuebai Gao;Syeda Sarah Shahrin;Zhen Bai;Ehsan Hoque

文献摘要

被引文献

相似文献

有毒、辱骂或不敬的行为分析是一个很重要的问题,以前主要从语言的角度来解决。本文提出了一种新的包含不尊重和非不尊重标签的视频数据集,并利用视觉线索对这种行为进行了分析。该数据集来自YouTube新闻节目中的两方对话视频,在视频中,主持人和嘉宾通过电话会议进行互动。每段视频都由三名训练有素的评分员进行评分,以识别通过面部和手势、声音和语言表达的不敬。通过分解混杂因素,我们生成了相应的两两不尊重样本。为了更好地展示视觉线索在不敬互动中的影响,我们提供了222个标记片段(时长=974.41(S),平均时长=4.39(S))。我们提取并分析了无礼行为中常见的面部动作单位(AVs)。结果显示,经Bonferroni矫正后,内眼窝隆起(AV01)、唇角抑制(AV15)和下巴隆起(AV17)均有统计学意义。对于预测,我们使用Logistic回归和线性支持向量机分别构建了两个分类器,准确率分别为62.61%和61.48%。为了深入分析人脸和手势的整体特征,我们使用主题提取进行了定性分析。我们的定性分析提供了有关利用同步和异步功能的进一步见解,以及将文本和音频数据与视觉提示相结合以更好地检测不尊重行为。
Toxic, abusive, or disrespectful behavior analysis is a non-trivial problem previously addressed mostly from the language perspective. In this paper, we present a novel video dataset containing disrespect and non-disrespect labels, and introduce such behavior analysis by using visual cues. The dataset is collected from YouTube news show videos of two-party conversations, in which a host and a guest interact through teleconferencing. Each video is labeled by three trained raters to identify disrespect expressed through face and gesture, voice, and language. By resolving confounding factors, we generate the corresponding pairwise samples of non-disrespect. To particularly show the influence of visual cues in disrespectful interactions, we present 222 labeled clips (duration=974.41(s), mean duration=4.39(s)). We extract and analyze the facial action units (AVs) prevalent in disrespectful behavior. Our result shows statistically significant differences after Bonferroni correction for Inner Brow raise (AV01), Lip Corner Depressor (AV15), and Chin Raiser (AV17). For prediction, we build two classifiers using logistic regression and linear Support Vector Machine with 62.61 % and 61.48 % accuracy, respectively. For an in-depth analysis of overall face and gesture features, we conduct a qualitative analysis using theme extraction. Our qualitative analysis provides further insights on leveraging synchronous and asynchronous features, along with combining text and audio data with visual cues to better detect disrespect.