Multimodal modeling of collaborative problem-solving facets in triads

Multimodal modeling of collaborative problem-solving facets in triads
复制标题

DOI:
10.1007/s11257-021-09290-y
复制
发表时间:
2021-02
影响因子:
3.6
通讯作者:
Angela E. B. Stewart;Z. Keirn;S. D’Mello
Angela E. B. Stewart;Z. Keirn;S. D’Mello
中科院分区:
计算机科学3区
文献类型:
--
作者:
Angela E. B. Stewart;Z. Keirn;S. D’Mello

文献摘要

被引文献

相似文献

协作解决问题(CPS)是无处不在的日常生活中,包括工作,家庭,休闲活动等,随着越来越多的远程协作发生,下一代协作界面可以增强CPS的过程和结果与动态干预或生成反馈行动后的评论。CPS过程的自动建模(这里称为facets)是实现这一目标的先驱。因此,我们建立自动化检测器的三个关键CPS方面的建设共享知识,谈判和协调,并保持团队的功能,从一个有效的CPS框架。我们使用了32个黑社会的数据,他们通过商业视频会议软件进行合作,以解决视觉编程任务中具有挑战性的问题。我们使用自动语音识别生成了11,163个话语的转录本,然后由受过训练的人进行编码,以获得CPS三个方面的证据。我们使用标准和深度顺序学习分类器,以独立于团队的方式从语言,任务上下文,面部表情和声学韵律特征中对人类编码的方面进行建模。我们发现,依赖于非语言信号的模型产生了高于机会的准确性(受试者工作特征曲线下的面积,AUROC),范围从0.53到0.83,当包括语言信息时,模型准确性增加(AUROCS从0.72到0.86)。深度顺序学习方法与标准分类器相比没有优势。总体而言,使用语言和任务上下文特征的随机森林分类器表现最好,分别在构建共享知识、谈判/协调和维护团队功能方面获得了0.86、0.78和0.79的AUROC分数。我们讨论了我们的工作应用到实时系统,评估CPS和干预,以改善CPS的结果。
Collaborative problem-solving (CPS) is ubiquitous in everyday life, including work, family, leisure activities, etc. With collaborations increasingly occurring remotely, next-generation collaborative interfaces could enhance CPS processes and outcomes with dynamic interventions or by generating feedback for after-action reviews. Automatic modeling of CPS processes (called facets here) is a precursor to this goal. Accordingly, we build automated detectors of three critical CPS facets—construction of shared knowledge, negotiation and coordination, and maintaining team function—derived from a validated CPS framework. We used data of 32 triads who collaborated via a commercial videoconferencing software, to solve challenging problems in a visual programming task. We generated transcripts of 11,163 utterances using automatic speech recognition, which were then coded by trained humans for evidence of the three CPS facets. We used both standard and deep sequential learning classifiers to model the human-coded facets from linguistic, task context, facial expressions, and acoustic–prosodic features in a team-independent fashion. We found that models relying on nonverbal signals yielded above-chance accuracies (area under the receiver operating characteristic curve, AUROC) ranging from .53 to .83, with increases in model accuracy when language information was included (AUROCS from .72 to .86). There were no advantages of deep sequential learning methods over standard classifiers. Overall, Random Forest classifiers using language and task context features performed best, achieving AUROC scores of .86, .78, and .79 for construction of shared knowledge, negotiation/coordination, and maintaining team function, respectively. We discuss application of our work to real-time systems that assess CPS and intervene to improve CPS outcomes.