Evaluating Speech, Face, Emotion and Body Movement Time-series Features for Automated Multimodal Presentation Scoring

Evaluating Speech, Face, Emotion and Body Movement Time-series Features for Automated Multimodal Presentation Scoring
复制标题

评估语音、面部、情绪和身体运动时间序列特征以实现自动多模式演示评分

DOI:
--
复制
发表时间:
2015
期刊:
International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
D. Suendermann
D. Suendermann
中科院分区:
--
文献类型:
--
作者:
Vikram Ramanarayanan;C. W. Leong;L. Chen;G. Feng;D. Suendermann

文献摘要

被引文献

相似文献

我们分析了如何融合功能,从不同的多模态数据流,如语音,面部,身体运动和情感轨迹可以应用到多模态演示文稿的评分。我们从这些数据流中计算时间聚合和基于时间序列的特征-前者是在整个时间序列中计算的统计泛函和其他累积特征,而后者,被称为同现直方图,捕获不同的原型身体姿势或面部配置如何在多模态,多变量时间序列的演变过程中在不同的时间滞后内共同出现。我们研究了这些功能的相对效用,沿着策划的语音流功能在预测演示熟练度的多个方面的人类评分。我们发现,不同的方式是有用的,在预测不同的方面,甚至优于一个天真的人类评分员之间的协议基线的一个子集的分析方面。
We analyze how fusing features obtained from different multimodal data streams such as speech, face, body movement and emotion tracks can be applied to the scoring of multimodal presentations. We compute both time-aggregated and time-series based features from these data streams--the former being statistical functionals and other cumulative features computed over the entire time series, while the latter, dubbed histograms of cooccurrences, capture how different prototypical body posture or facial configurations co-occur within different time-lags of each other over the evolution of the multimodal, multivariate time series. We examine the relative utility of these features, along with curated speech stream features in predicting human-rated scores of multiple aspects of presentation proficiency. We find that different modalities are useful in predicting different aspects, even outperforming a naive human inter-rater agreement baseline for a subset of the aspects analyzed.