BehancePR: A Punctuation Restoration Dataset for Livestreaming Video Transcript

BehancePR: A Punctuation Restoration Dataset for Livestreaming Video Transcript
复制标题

DOI:
10.18653/v1/2022.findings-naacl.149
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Viet Dac Lai;Amir Pouran Ben Veyseh;Franck Dernoncourt;Thien Huu Nguyen
Viet Dac Lai;Amir Pouran Ben Veyseh;Franck Dernoncourt;Thien Huu Nguyen
中科院分区:
其他
文献类型:
--
作者:
Viet Dac Lai;Amir Pouran Ben Veyseh;Franck Dernoncourt;Thien Huu Nguyen

文献摘要

被引文献

相似文献

鉴于直播视频数量的不断增加,直播视频文本的自动语音识别和后处理对于高效的数据管理和知识挖掘至关重要。这个过程中的一个关键步骤是标点符号恢复,从视频转录中恢复基本的文本结构,如短语和句子边界。这项工作提出了一个新的人类注释语料库,称为BehancePR,用于在直播视频转录中进行标点符号恢复。我们在BehancePR上的实验证明了这一领域标点符号恢复的挑战。此外,我们还发现,流行的自然语言处理工具包,如斯坦福大学Stanza,Spacy和Trankit在检测直播视频的非标点文字稿上的句子边界时表现不佳。该数据集可在http://github.com/ nlp-uoclave/behancepr上公开访问。
Given the increasing number of livestreaming videos, automatic speech recognition and post-processing for livestreaming video transcripts are crucial for efficient data manage-ment as well as knowledge mining. A key step in this process is punctuation restoration which restores fundamental text structures such as phrase and sentence boundaries from the video transcripts. This work presents a new human-annotated corpus, called BehancePR, for punctuation restoration in livestreaming video transcripts. Our experiments on BehancePR demonstrate the challenges of punctuation restoration for this domain. Furthermore, we show that popular natural language processing toolkits like Stanford Stanza, Spacy, and Trankit underperform on detecting sentence boundary on non-punctuated transcripts of livestreaming videos. The dataset is publicly accessible at http://github.com/ nlp-uoregon/behancepr .