Introduction to the Special Issue on Recent Advances in Asian Language Spoken Document Retrieval

Introduction to the Special Issue on Recent Advances in Asian Language Spoken Document Retrieval
复制标题

DOI:
10.1145/1482343.1482344
复制
发表时间:
2009-03
期刊:
ACM Trans. Asian Lang. Inf. Process.
影响因子:
--
通讯作者:
Chung-Hsien Wu;Haizhou Li
Chung-Hsien Wu;Haizhou Li
中科院分区:
其他
文献类型:
--
作者:
Chung-Hsien Wu;Haizhou Li

文献摘要

被引文献

相似文献

美国计算机协会亚洲语言信息处理汇刊关于亚洲语言口语文件检索的最新进展特刊致力于亚洲语言口语文件处理这一新兴领域的最新进展。由于广播、电视、讲座和电话录音等各种音频媒体来源的口语内容迅速增加,导致对这些相对非结构化的口语文档进行更有效的自动索引和检索的需求不断增加。然而,就像文本文档一样,口语文档可以通过主题、主题和语义概念等属性来描述。因此,大量的口头文件应该像文本文件一样可供我们使用。然而,与处理具有更好结构的文本文档(例如,具有标题、标题和段落)不同,检索和浏览语音文档严重依赖于语音识别引擎的性能,而语音识别引擎还远远不够完美。在这期特刊中,我们鼓励提交关于口头文件检索、口头文件摘要、口头文件翻译和其他相关研究领域的新技术的报告。特别令人感兴趣的是直接解决涉及亚洲口语问题的研究。语音文档检索(SDR)本质上是根据用户的请求从大量的语音文档中检索摘录的任务。SDR是语音搜索、语音监控、语音数据挖掘和呼叫中心自动化等许多应用中的关键元素。作为一门涉及自动语音识别、自然语言处理和信息检索的跨学科研究,SDR领域受益于语音和语言处理的进步以及大型语音数据库的可用性。这些庞大的音频档案中有价值的内容需要以一种有效而实用的方式来获取。这种需求反过来又产生了口头文件摘要领域,它寻求从口头文件中提取重要信息,同时删除冗余的、不正确的信息,以产生用户友好的、可略读的口头文件摘要。信息提取和口语摘要生成等研究问题是我们在口语文档摘要领域研究的重点。
This special issue on Recent Advances in Asian Language Spoken Document Retrieval of the ACM Transactions on Asian Language Information Processing is devoted to recent advances in the burgeoning field of spoken document processing of Asian languages. Rapidly increasing spoken language content from the various audio media sources such as radio, television, lectures, and telephony recordings has led to an increasing demand for more effective automatic indexing and retrieval of these relatively unstructured spoken documents. Yet, just like text documents, spoken documents can be described by attributes such as subjects, topics, and semantic concepts. As such, the vast amount of spoken documents available should be as accessible to us as are text documents. However, unlike handling text documents which are better structured, for example, with titles, headings, and paragraphs, retrieving and browsing spoken documents relies heavily on the performance of speech recognition engines, which are still far from perfect. For this special issue, we encouraged submissions that report on novel techniques for tasks such as spoken document retrieval, spoken document summarization, spoken document translation, and other related areas of research. Of particular interest are studies that directly address problems involving spoken Asian languages. Spoken Document Retrieval (SDR) is essentially the task of retrieving excerpts from a large collection of spoken documents based on a user’s request. SDR is the key element in many applications such as voice search, voice surveillance, voice data mining, and call center automation. As an interdisciplinary research involving automatic speech recognition, natural language processing, and information retrieval, the area of SDR has benefited much from the advances in speech and language processing as well as from the availability of large spoken databases. The valuable content in these large audio archives needs to be accessed in an effective yet practical way. This need in turn gives rise to the area of spoken document summarization, which seeks to distill salient information while removing redundant, incorrect information from spoken documents to produce user-friendly, skim-able spoken document summaries. Research problems, such as information extraction and spoken summary generation, are our main focus in the area of spoken document summarization.