PodCastle: Recent Advances of a Spoken Document Retrieval Service Improved by Anonymous User Contributions

PodCastle: Recent Advances of a Spoken Document Retrieval Service Improved by Anonymous User Contributions
复制标题

PodCastle:通过匿名用户贡献改进的语音文档检索服务的最新进展

DOI:
--
复制
发表时间:
2011
期刊:
Interspeech
影响因子:
--
通讯作者:
J. Ogata
J. Ogata
中科院分区:
--
文献类型:
--
作者:
Masataka Goto;J. Ogata

文献摘要

被引文献

相似文献

的困难和问题·无法避免对各种类型的语音数据犯识别错误语音语料库无法预先准备·难以支持新单词/短语(专有名称和流行语)播客通常包含词汇表外的单词·难以推出高精度的语音文档检索服务用户可能会对ASR结果感到失望基于ASR和众包的语音检索Web服务·收集和放大匿名用户的自愿贡献自动从网络学习·自动收集新单词/短语、它们的发音和用法示例新闻文章(雅虎!新闻)和网络词典·向ASR词典添加新词(24万字)用户可以查找并更正ASR错误·独创的高效纠错界面[绪方和后藤,InterSpeech 2005]·通过正确的索引提高检索性能·通过自动学习(自适应/训练)搜索功能提高识别性能·自动语音识别结果的全文搜索·包含查询的语音数据列表与文本摘录一起显示·摘录可以单独回放阅读功能·查看语音数据的转录文本·每个单词根据自动语音识别的可靠性程度进行着色·全文可由外部搜索引擎索引和访问(例如,GOOGLE)注释功能(纠错)·添加注释以纠正自动语音识别错误·从列表中选择正确的候选者列表是通过使用一个浓缩了巨大内部词图的混淆网络生成的·键入正确的文本·更正的错误可用于提高检索和识别性能搜索播客本文描述了一个公共Web服务PodCastle,它提供基于自动语音识别的日语播客全文搜索。这是我们的研究方法-语音识别研究2.0的一个实例,该方法旨在为用户提供基于Web 2.0的Web服务,使他们能够体验最先进的语音每
 Difficulties and Problems • Cannot avoid making recognition errors for various types of speech data Speech corpus cannot be prepared in advance • Difficult to support new words/phrases (proper names and buzzwords) Podcasts often include out-of-vocabulary words • Difficult to launch a spoken document retrieval service with high accuracy Users might be disappointed by ASR results  Speech retrieval web service based on ASR and crowdsourcing • Collect and amplify voluntary contributions by anonymous users  Automatic learning from the web • Automatically collect new words/phrases, their pronunciation, and usage examples News articles (Yahoo! news) and web dictionaries • Add new words to ASR dictionary (0.24M words)  Users can find and correct ASR errors • Original efficient error correction interface [Ogata & Goto, Interspeech 2005] • Improve retrieval performances by correct indices • Improve recognition performances by automatic learning (adaptation/training)  Searching function • Full-text search of ASR results • List of speech data containing a query is displayed together with text excerpts • Excerpts can be played back individually  Reading function • View the transcribed text of speech data • Each word is colored according to the degree of ASR reliability • Full text can be indexed and accessed by external search engines (e.g., Google)  Annotating function (error correction) • Add "annotations" to correct ASR errors • Select the correct candidate from the list The list is generated by using a confusion network that condenses a huge internal word graph • Type in the correct text • Corrected errors can be used for improving retrieval and recognition performances Search podcast In this paper, we describe a public web service, "PodCastle", that provides full-text searching of Japanese podcasts on the basis of automatic speech recognition. This is an instance of our research approach, "Speech Recognition Research 2.0", which is aimed at providing users with a web service based on Web 2.0 so that they can experience state-of-the-art speech per-