Noise source detection and recognition for audio indexing
Noise source detection and recognition for audio indexing
批准号:
17500114
负责人:
MATSUNAGA Shoichi
金额:
$2.46万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2005
资助国家:
日本
项目状态:
已结题
起止时间:
2005 至 2007
中文摘要
我们研究了一种基于随机方法的音频源检测方法,用于检测语音、噪声、音乐和静音。我们的方法不仅使用传统的表面声学特征,如信号能量和音调频率,但也基于频谱相关性更准确的检测新功能。广播新闻的实验表明,这些特征参数使得更准确地捕获音频源成为可能。本研究亦提出一种基于噪声模型的音源侦测方法。为了准确检测,我们设计了两种方法来生成多个噪声模型,通过聚类技术。一种方法是基于逐帧数据相似性,另一种是基于噪声源相似性。前一种方法采用K均值聚类和平滑技术,以避免不准确的分割。后一种方法涉及噪声建模的基础上产生的噪声集群的逐步合并的树数据结构。分类实验表明,通过使用这些方法,音频源可以检测到更好的准确性比传统的方法实现。当使用后一种方法产生的四个噪声模型时,对于声源不重叠的时段,噪声检测性能提高了3.9%。对于包括重叠段的音频流的实验,噪声检测性能提高了1.2%,而语音检测性能没有降低。
英文摘要
We have studied an audio source detection approach based on a stochastic method to detect speech, noise, music, and silence. Our approach uses not only conventional surface acoustic features such as signal energy and pitch frequency but also new features that are based on spectral correlation for more accurate detection. The experiment with the broadcast news demonstrated that these feature parameters made it possible to capture the audio source more accurately. This research also proposed a sound source detection approach based on elaborate noise-modeling techniques for audio indexing. For accurate detection, we devised two methods to generate multiple-noise models through clustering techniques. One method is based on frame-wise data similarity, and the other is based on noise source similarity. The former method employs K-means clustering and a smoothing technique to avoid inaccurate segmentation. The latter method involves noise modeling based on a tree data structure generated by the progressive merging of noise clusters. The classification experiments show that by using these proposed methods, audio sources can be detected with better accuracy than that achieved by the conventional methods. When four noise models generated by the latter method were used, the noise detection performance increased by 3.9% for the periods in which the sound sources did not overlap. With regard to the experiments for an audio stream that included overlapped segments, the noise detection performance increased by 1.2% without a decrease in the speech detection performance.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Sound source detection using multiple noise models
使用多种噪声模型进行声源检测
DOI:
--
发表时间:
2008
期刊:
Proc.of ICASSP 2008「掲載確定」
影响因子:
--
作者:
[Sumi K., Liu C., Matsuyama T., K.Kanatani, K.Kanatani, K.Kanatani, K.Kanatani et al., K.Kanatani, K.Kanatani et al., Y.Sugaya, K.Kanatani, K.Kanatani, K.Kanatani, K.Kanatani, Y.Sugaya et al., K.Kanatani et al., K.Kanatani, K.Kanatani et al., K.Kanatani, R.Klette et al., K.Kanatani, K.Kanatani, A.Nakatsuji, K.Kanatani et al., A.Nakatsuji et al., E.Bayro Corrochano et al., K.Kanatani, K.Kanatani, K.Kanatani, R.Klette, E.Bayro Corrochano, 金谷健一, K.Kanatani, 金谷健一, Shoichi Matsunaga]
通讯作者:
Shoichi Matsunaga
特徴的音情報検出によるニューストピック分割手法の検討
基于特征声音信息检测的新闻主题切分方法研究
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[Shoichi, Matsunaga, 金城 潤]
通讯作者:
金城 潤
Emotion clustering using the results of subjective opinion tests for emotion recognition in infant's cries
使用婴儿哭声情绪识别的主观意见测试结果进行情绪聚类
DOI:
--
发表时间:
2007
期刊:
Proc.of INTERSPEECH 2007
影响因子:
--
作者:
[Sumi K., Liu C., Matsuyama T., K.Kanatani, K.Kanatani, K.Kanatani, K.Kanatani et al., K.Kanatani, K.Kanatani et al., Y.Sugaya, K.Kanatani, K.Kanatani, K.Kanatani, K.Kanatani, Y.Sugaya et al., K.Kanatani et al., K.Kanatani, K.Kanatani et al., K.Kanatani, R.Klette et al., K.Kanatani, K.Kanatani, A.Nakatsuji, K.Kanatani et al., A.Nakatsuji et al., E.Bayro Corrochano et al., K.Kanatani, K.Kanatani, K.Kanatani, R.Klette, E.Bayro Corrochano, 金谷健一, K.Kanatani, 金谷健一, Shoichi Matsunaga, Noriko Satoh]
通讯作者:
Noriko Satoh
Noise-source clustering for accurate audio source detection
噪声源聚类可实现准确的音频源检测
DOI:
--
发表时间:
2007
期刊:
影响因子:
--
作者:
[Shoichi, Matsunaga]
通讯作者:
Matsunaga
DOI:
--
发表时间:
2006
期刊:
Interspeech 2006
影响因子:
--
作者:
[Sumi K., Liu C., Matsuyama T., K.Kanatani, K.Kanatani, K.Kanatani, K.Kanatani et al., K.Kanatani, K.Kanatani et al., Y.Sugaya, K.Kanatani, K.Kanatani, K.Kanatani, K.Kanatani, Y.Sugaya et al., K.Kanatani et al., K.Kanatani, K.Kanatani et al., K.Kanatani, R.Klette et al., K.Kanatani, K.Kanatani, A.Nakatsuji, K.Kanatani et al., A.Nakatsuji et al., E.Bayro Corrochano et al., K.Kanatani, K.Kanatani, K.Kanatani, R.Klette, E.Bayro Corrochano, 金谷健一, K.Kanatani, 金谷健一, Shoichi Matsunaga, Noriko Satoh, Shoichi Matsunaga]
通讯作者:
Shoichi Matsunaga
共 13 条
Detection of unhealthy subjects based on stochastic modeling and extraction of adventitious sounds
-
批准号:17K00408
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.75万
-
财政年份:2017
-
负责人:MATSUNAGA Shoichi
-
依托单位:
Robust classification between a healthy subject and a patient with pulmonary emphysema using lung sound samples based on a stochastic approach
-
批准号:23500217
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$3.0万
-
财政年份:2011
-
负责人:MATSUNAGA Shoichi
-
依托单位:
Detection of acoustic features in human biological sounds based on statistical approach
-
批准号:20500157
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.83万
-
财政年份:2008
-
负责人:MATSUNAGA Shoichi
-
依托单位:
海外基金