PodCastle: Recent Advances of a Spoken Document Retrieval Service Improved by Anonymous User Contributions
PodCastle: Recent Advances of a Spoken Document Retrieval Service Improved by Anonymous User Contributions
复制标题
PodCastle:通过匿名用户贡献改进的语音文档检索服务的最新进展
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
J. Ogata
中科院分区:
文献类型:
--
作者:
Masataka Goto;J. Ogata
Difficulties and Problems • Cannot avoid making recognition errors for various types of speech data Speech corpus cannot be prepared in advance • Difficult to support new words/phrases (proper names and buzzwords) Podcasts often include out-of-vocabulary words • Difficult to launch a spoken document retrieval service with high accuracy Users might be disappointed by ASR results Speech retrieval web service based on ASR and crowdsourcing • Collect and amplify voluntary contributions by anonymous users Automatic learning from the web • Automatically collect new words/phrases, their pronunciation, and usage examples News articles (Yahoo! news) and web dictionaries • Add new words to ASR dictionary (0.24M words) Users can find and correct ASR errors • Original efficient error correction interface [Ogata & Goto, Interspeech 2005] • Improve retrieval performances by correct indices • Improve recognition performances by automatic learning (adaptation/training) Searching function • Full-text search of ASR results • List of speech data containing a query is displayed together with text excerpts • Excerpts can be played back individually Reading function • View the transcribed text of speech data • Each word is colored according to the degree of ASR reliability • Full text can be indexed and accessed by external search engines (e.g., Google) Annotating function (error correction) • Add "annotations" to correct ASR errors • Select the correct candidate from the list The list is generated by using a confusion network that condenses a huge internal word graph • Type in the correct text • Corrected errors can be used for improving retrieval and recognition performances Search podcast In this paper, we describe a public web service, "PodCastle", that provides full-text searching of Japanese podcasts on the basis of automatic speech recognition. This is an instance of our research approach, "Speech Recognition Research 2.0", which is aimed at providing users with a web service based on Web 2.0 so that they can experience state-of-the-art speech per-