Comparing Named Entity Recognition on Transcriptions and Written Texts

Comparing Named Entity Recognition on Transcriptions and Written Texts
复制标题

比较转录和书面文本的命名实体识别

DOI:
--
复制
发表时间:
2015
期刊:
Italian Natural Language Processing within the PARLI Project
影响因子:
--
通讯作者:
Roberto Zanoli
Roberto Zanoli
中科院分区:
--
文献类型:
--
作者:
Firoj Alam;B. Magnini;Roberto Zanoli

文献摘要

被引文献

相似文献

识别文本中的命名实体(例如,人名、地点和组织名称)的能力已经被证明是包括信息检索和信息提取在内的几个自然语言处理领域的一项重要任务。然而,尽管在书面文本的命名实体识别方面取得了很大的努力和成果,但从口头文档的自动转录中识别命名实体的问题仍然远未解决。事实上,自动语音识别(ASR)的输出经常包含转录错误;此外,许多命名实体是词汇表外的单词,这使得它们对ASR不可用。本文对从书面文本和转录文本中提取命名实体进行了比较分析。对于抄本,我们使用的是口头广播新闻,而对于书面文本,我们使用的是同一领域的报纸的抄本和广播新闻的人工抄写。使用在Evalita 2007上展示的最佳命名实体识别系统进行了多次实验。
The ability to recognize named entities (e.g., person, location and organization names) in texts has been proved as an important task for several natural language processing areas, including Information Retrieval and Information Extraction. However, despite the efforts and the achievements obtained in Named Entity Recognition from written texts, the problem of recognizing named entities from automatic transcriptions of spoken documents is still far from being solved. In fact, the output of Automatic Speech Recognition (ASR) often contains transcription errors; in addition, many named entities are out-of-vocabulary words, which makes them not available to the ASR. This paper presents a comparative analysis of extracting named entities both from written texts and from transcriptions. As for transcriptions, we have used spoken broadcast news, while for written texts we have used both newspapers of the same domain of the transcriptions and the manual transcriptions of the broadcast news. The comparison was carried on a number of experiments using the best Named Entity Recognition system presented at Evalita 2007.