Modeling Word Importance in Conversational Transcripts: Toward improved live captioning for Deaf and hard of hearing viewers
Modeling Word Importance in Conversational Transcripts: Toward improved live captioning for Deaf and hard of hearing viewers
复制标题
DOI:
10.1145/3587281.3587290
复制
发表时间:
2023-04
期刊:
影响因子:
--
通讯作者:
Akhter Al Amin;Saad Hassan;Matt Huenerfauth;Cecilia Ovesdotter Alm
中科院分区:
文献类型:
--
作者:
Akhter Al Amin;Saad Hassan;Matt Huenerfauth;Cecilia Ovesdotter Alm
Despite the recent improvements in automatic speech recognition (ASR) systems, their accuracy is imperfect in live conversational settings. Classifying the importance of each word in a caption transcription can enable evaluation metrics that best reflect Deaf and Hard of Hearing (DHH) readers’ judgment of the caption quality. Prior work has proposed using word embeddings, e.g., word2vec or BERT embeddings, to model word importance in conversational transcripts. Recent work also disseminated a human-annotated word importance dataset. We conducted a word-token level analysis on this dataset and explored Part-of-Speech (POS) distribution. We then augmented the dataset with POS tags and reduced the class imbalance by generating 5% additional text using masking. Finally, we investigated how various supervised models learn the importance of words. The best performing model trained on our augmented dataset performed better than prior models. Our findings can inform the design of a metric for measuring live caption quality from DHH users’ perspectives.