Modeling Word Importance in Conversational Transcripts: Toward improved live captioning for Deaf and hard of hearing viewers

Modeling Word Importance in Conversational Transcripts: Toward improved live captioning for Deaf and hard of hearing viewers
复制标题

DOI:
10.1145/3587281.3587290
复制
发表时间:
2023-04
期刊:
Proceedings of the 20th International Web for All Conference
影响因子:
--
通讯作者:
Akhter Al Amin;Saad Hassan;Matt Huenerfauth;Cecilia Ovesdotter Alm
Akhter Al Amin;Saad Hassan;Matt Huenerfauth;Cecilia Ovesdotter Alm
中科院分区:
其他
文献类型:
--
作者:
Akhter Al Amin;Saad Hassan;Matt Huenerfauth;Cecilia Ovesdotter Alm

文献摘要

相似文献

尽管自动语音识别(ASR)系统最近有所改进,但它们在实时对话环境中的准确性并不完美。对字幕转录中每个单词的重要性进行分类能够使评估指标最能反映聋人和重听(DHH)读者对字幕质量的判断。先前的研究已提出使用词嵌入,例如word2vec或BERT嵌入,来对对话转录中的单词重要性进行建模。近期的研究还发布了一个人工标注的单词重要性数据集。我们对该数据集进行了词元层面的分析,并探究了词性(POS)分布。然后,我们用词性标签扩充了数据集,并通过使用掩码生成5%的额外文本减少了类别不平衡。最后,我们研究了各种有监督模型如何学习单词的重要性。在我们扩充后的数据集上训练的最佳性能模型比之前的模型表现更好。我们的研究结果可以为从聋人和重听用户的角度衡量实时字幕质量的指标设计提供参考。
Despite the recent improvements in automatic speech recognition (ASR) systems, their accuracy is imperfect in live conversational settings. Classifying the importance of each word in a caption transcription can enable evaluation metrics that best reflect Deaf and Hard of Hearing (DHH) readers’ judgment of the caption quality. Prior work has proposed using word embeddings, e.g., word2vec or BERT embeddings, to model word importance in conversational transcripts. Recent work also disseminated a human-annotated word importance dataset. We conducted a word-token level analysis on this dataset and explored Part-of-Speech (POS) distribution. We then augmented the dataset with POS tags and reduced the class imbalance by generating 5% additional text using masking. Finally, we investigated how various supervised models learn the importance of words. The best performing model trained on our augmented dataset performed better than prior models. Our findings can inform the design of a metric for measuring live caption quality from DHH users’ perspectives.