A research for automatic structure extraction and information retrieval from video data using speech and voice
A research for automatic structure extraction and information retrieval from video data using speech and voice
批准号:
17500073
负责人:
ITOH Yoshiaki
金额:
$2.41万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2005
资助国家:
日本
项目状态:
已结题
起止时间:
2005 至 2007
中文摘要
本研究的第一个主题是实现视频的自动分割。为此,我们开发了提取视频信息局部和全局特征以及对视频数据集进行分类的方法。第二个主题是利用分割部分之间的局部和全局相似性和不相似性自动提取分割视频数据集的结构。本研究首先对视频数据集中的局部声学和图像相似性进行了分析,并提取了在音乐片段中重复出现的相似部分。然后,所开发的方法分别使用局部特征和全局特征来区分语音部分和音乐部分。我们证实了所开发的方法对真实的视频数据集效果良好。本研究还对上述方法分割的语音视频片段进行了文本检索和语音查询。在语音检索方面,提出了对查询进行任意词处理的新方法和对多个子词模型进行集成的方法,实验结果表明,该方法比以前的方法具有更好的性能。这些结果在许多国内和国际会议上进行了报道。在未来,我们将开发实际使用的方法。
英文摘要
The first theme of this research is to realize an automatic video segmentation. For this purpose,we developed the methods for extracting a local and global feature of video information and for classifying video data sets. The second theme is automatic extraction of the structure for segmented video data sets using local and global similarity and dissimilarity between the segmented sections. This research first conducted the analysis of local acoustic and image similarity in video data sets,and extracted similar partial sections that are repeated in a music piece. The developed method then discriminated speech sections and music sections respectively using a local feature and a global feature.We confirmed the developed methods worked well for real video data sets. The research also conducted speech retrieval by a text and speech query for speech video sections segmented by the above method. For speech retrieval,new technique of dealing with any words for a query and the integration method for plural subword models were proposed,and the experimental results demonstrated the method showed better performance compared to former methods. These results were reported at the many domestic and international conferences. In the future,we are going to develop the methods for the actual use.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
語彙フリー音声検索におけるサブワードの検討および災害放送検索システムへの応用
无词汇语音搜索子词研究及其在灾害广播搜索系统中的应用
DOI:
--
发表时间:
2005
期刊:
電子情報通信学会研究技術報告 SP2005-21
影响因子:
--
作者:
[岩田耕平, 伊藤慶明, 小嶋和徳, 石亀昌明, 田中和世, 李時旭]
通讯作者:
李時旭
DOI:
--
发表时间:
2005
期刊:
インタラクション2005
影响因子:
--
作者:
[S.Aubry, S.Okawa, D.Lenne, I.Thouvenin]
通讯作者:
I.Thouvenin
語彙非依存型音声文書検索のためのサブワードモデルおよび検索方式の検討
与词汇无关的口语文档检索的子词模型和搜索方法研究
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[岩田 耕平, 伊藤 慶明, 小嶋 和徳, 石亀 昌明, 田中 和世, Shi-wook Lee]
通讯作者:
Shi-wook Lee
Layered Server-Client Topology for Parallel Distributed GA on Large Problem
针对大型问题的并行分布式遗传算法的分层服务器-客户端拓扑
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[Kazunori K., Masaaki I., Shozo M]
通讯作者:
Shozo M
音声検索システムのための時間整合を考慮したサブワードモデル構築手法の検討
考虑时间对齐的语音搜索系统子词模型构建方法研究
DOI:
--
发表时间:
2006
期刊:
情報処理学会 研究技術報告、2006-SLP-062
影响因子:
--
作者:
[岩田耕平, 伊藤慶明, 小嶋和徳, 石亀昌明, 田中和世, 李時旭]
通讯作者:
李時旭
共 59 条
Tissue engineered blood vessel sheet to prevent cerebral infarction
-
批准号:23591281
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$3.66万
-
财政年份:2011
-
负责人:ITOH Yoshiaki
-
依托单位:
Random sequential packing of cubes into torus
-
批准号:23540177
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.91万
-
财政年份:2011
-
负责人:ITOH Yoshiaki
-
依托单位:
A Stochastic Construction of Golay Code
-
批准号:02640194
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$0.51万
-
财政年份:1990
-
负责人:ITOH Yoshiaki
-
依托单位:
Statistical Destribution on Symmetry Groups
-
批准号:61540171
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$0.64万
-
财政年份:1986
-
负责人:ITOH Yoshiaki
-
依托单位:
海外基金