课题基金 / 基金详情

FODAVA-Partner: Visualizing Audio for Anomaly Detection

FODAVA-Partner: Visualizing Audio for Anomaly Detection
FODAVA-合作伙伴:可视化音频以进行异常检测
批准号:
0807329
负责人:
Mark Hasegawa-Johnson
金额:
$45.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2014-08-31

项目摘要

项目成果

Mark Hasegawa-Johnson的其他基金

相似基金

相关文献

中文摘要
翻译
AbstractFODAVA-合作伙伴:异常检测的可视化音频Mark Hasegawa-Johnson,托马斯黄和Camille Goudeseun本提案的目标是将大型音频语料库转换为适合可视化的形式。 具体来说,该提案解决了人类数据分析师立即听到的音频异常类型:愤怒的喊叫,午夜居民区街道上的卡车,枪声。 人耳可以快速且高精度地检测到这种类型的异常。 不幸的是,数据分析师一次只能听到一种声音。 可视化显示分析师许多声音一次,可能让他或她发现一个异常几个数量级的速度比?真实的时间 该提案旨在以交互式图形的形式呈现大型音频数据集,包括数千个麦克风或数千分钟,这些图形可以一目了然地揭示重要的异常。 数据转换将包括信号处理、统计建模和可视化。 信号处理将寻求表征两个音频信号之间的差异的所有方式。重要吗?包括例如频谱差异、节奏差异和在人类收听者的听觉皮层上产生的印象的差异。 统计建模将寻求特征的音频事件的范围是?正常吗或容易解释的,这样我们就可以精确地测量潜在异常的异常或无法解释的程度。 可视化方法将以适合于快速浏览的形式呈现异常的测量以及关于每个异常的信号特征的信息。 提出了两个试验平台。 什么?多日音频时间轴将是一个便携式应用程序,视觉上类似于非线性音频编辑套件,这将允许分析师快速放大潜在的异常时间段。 什么?手机?将成为指挥和控制中心的三维可视化工具。 分散在城市或大型工业场所的1000个安全麦克风的音频记录将以颜色鲜艳的可见线的形式呈现,这些线从安全区域的地图上延伸到天空。 分析师将能够通过触摸其线来收听任何麦克风上记录的音频;通过触摸不同高度的线,分析师将能够审计不同的时间段。 每条线的亮度、颜色、粗细都会显示出音频信号在每个时间点的异常情况和信号特征。 数据转换研究将相关类型的异常映射到相关的颜色/亮度代码,以便经过培训的数据分析师立即看到重要的异常和通过多个麦克风确认的异常。这项研究旨在在安全分析领域产生广泛的影响。 摄像机通常用于工业场所、政府设施、日托中心和疗养院的安全监控。 数据分析师和警卫经常同时浏览多达20个监控摄像头记录的视频,快进到不感兴趣的时间段。 如果麦克风有用的话,它们也会被用于相同的应用中;但是目前数据分析师没有办法快速准确地审计来自许多不同麦克风的信号。 所提出的技术将为警卫和数据分析师提供新的数据转换和可视化技术,帮助他们快速识别由无法解释但异常的音频信号发出的危险情况。
英文摘要
AbstractFODAVA-Partner: Visualizing Audio for Anomaly DetectionMark Hasegawa-Johnson, Thomas Huang and Camille GoudeseuneThe goal of this proposal is to transform large audio corpora into a form suitable for visualization. Specifically, this proposal addresses the type of audio anomalies that human data analysts hear instantly: angry shouting, trucks at midnight on a residential street, gunshots. The human ear detects anomalies of this type rapidly and with high accuracy. Unfortunately, a data analyst can listen to only one sound at a time. Visualization shows the analyst many sounds at once, possibly allowing him or her to detect an anomaly several orders of magnitude faster than ?real time.? This proposal aims to render large audio data sets, comprising thousands of microphones or thousands of minutes, in the form of interactive graphics that reveal important anomalies at a glance. Data transformations will include signal processing, statistical modeling, and visualization. Signal processing will seek to characterize all of the ways in which the difference between two audio signals may be ?important,? including, for example, spectral differences, rhythmic differences, and differences in the impression made on the auditory cortex of a human listener. Statistical modeling will seek to characterize the range of audio events that are ?normal? or easily explicable, so that we may precisely measure the degree to which a potential anomaly is abnormal or inexplicable. Visualization methods will render measures of abnormality, and information about the signal characteristics of each anomaly, in a form suitable for rapid browsing. Two testbeds are proposed. The ?multi-day audio timeline? will be a portable application, visually similar to a nonlinear audio editing suite, which will allow the analyst to rapidly zoom in on potentially anomalous periods of time. The ?milliphone? will be a three-dimensional visualization tool for command and control centers. Audio recordings from one thousand security microphones scattered throughout a city or a large industrial site will be rendered in the form of brightly colored visible threads reaching skyward from a map of the secure region. The analyst will be able to listen to the audio recorded on any microphone by touching its thread; by touching the thread at different heights, the analyst will be able to audit different periods of time. The brightness, color, and thickness of each thread will display the abnormality and signal characteristics of the audio signal at each point in time. Data transformation research will map related types of abnormality to related color/brightness codes, so that important anomalies, and anomalies confirmed by multiple microphones, are immediately visible to the trained data analyst.This research seeks broad impact in the area of security analysis. Video cameras are routinely used for security monitoring of industrial sites, government installations, day care centers, and nursing homes. Data analysts and guards routinely browse the video recorded by up to twenty surveillance cameras simultaneously, fast-forwarding through uninteresting periods of time. Microphones would be used in the same applications, if they were useful; but there is at present no way for a data analyst to rapidly and accurately audit the signals from many different microphones. The proposed techniques will give guards and data analysts new data transformation and visualization techniques that will help them to rapidly identify dangerous situations signaled by inexplicable but anomalous audio signals.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
RI Medium: Audio Diarization - Towards Comprehensive Description of Audio Events
海外基金