课题基金 / 基金详情

III: Small: Collaborative Research: Algorithms for Query by Example of Audio Databases

III: Small: Collaborative Research: Algorithms for Query by Example of Audio Databases
III:小:协作研究:以音频数据库为例的查询算法
批准号:
1617107
负责人:
Zhiyao Duan
金额:
$29.98万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2020-08-31

项目摘要

项目成果

Zhiyao Duan的其他基金

相似基金

相关文献

中文摘要
翻译
随着多媒体存储库的激增和增长,寻找自动索引、标记和访问多媒体内容(如音频文档)的方法变得越来越重要。社区生成的SoundCloud存储库就是一个例子。它包含乐队、音效、播客等的录音,贡献者每分钟上传12个小时的音频。像SoundCloud这样的存储库通常会在文件级使用简短的文本标签来标记音频。使用这些标签对所需记录进行基于文本的搜索可能会有问题。在音轨内基于文本的搜索是不可能的,因为它们没有使用文件正文中的标签进行索引。在这个项目中,罗切斯特大学和西北大学的研究人员的目标是开发一种方法和系统,用于通过逐个实例查询来进行音频搜索,其中的实例在某些关键方面与数据库中所需的音频相似,但并不完全匹配。这将允许在文件中进行搜索,而不需要基于文本的标记。这个项目将专注于使用声音模仿作为搜索关键字,因为它们对人类来说是自然的,并且在互动中被广泛使用。它将开发一种新的声音搜索引擎,将声音模仿作为查询(例如,模仿鸟鸣以找到鸟鸣的录音)。为通过这种新颖的方式搜索音频/视频集合而开发的技术还将在许多其他方面有益于社会,诸如犯罪监视(例如,用于政策监控站的自动枪声或尖叫检测)、生物多样性测量(例如,在野外记录中听起来像这样的鸟类的自动ID)、对听力受损的人的环境意识(例如,当我的狗叫的时候提醒我)、电影声音设计者的制作辅助工具(例如,在数千个声音效果的数据库中寻找门砰的声音)、以及基于声音的诊断(例如,“你的车需要一个新的起动机”)。该项目将有利于科学、技术、工程和数学(STEM)教育,因为基于音频的研究已被证明是吸引不同大学生进入STEM学科的一种成功方式。声音模仿传达了丰富的信息,涵盖了许多声学方面:音高、响度、音色、它们的时间演变和节奏模式等。这使得用户可以查询使用文本标签难以搜索的准确声音。然而,出于同样的原因,声音模仿可能在许多方面与期望的目标不同。由于人类发声系统的物理限制,与要检索的声音相比,查询声音还可以位于非常受限的声音空间中。建立一个成功的模仿发声查询系统将需要研究表示音频和基于仅在其可测量维度的子集上与目标声音相似的查询来检索音频的方法。它还需要便于在非基于文本的上下文中提供查询和改进搜索结果的界面。对于前者,研究人员将研究使用深度神经网络学习特定于方面的音频表示的方法。调查人员还将开发适用于这些表示的匹配算法。调查人员将设计新颖的搜索界面,让用户反复优化他们的搜索结果。该系统将从交互中学习并调整不同声学方面的权重,以搜索所需的声音。这项研究的预期结果是:(1)突出声音查询与一般音频目标声音匹配的感知相关特征的音频表示;(2)将声音查询与一般音频匹配和对齐的算法;(3)使用声音模仿和声音实例迭代地精炼搜索结果的交互方法;(4)大型声音模拟和声音数据集;以及(5)体现这些结果的开源声音检索系统。有关该项目的更多信息,请访问项目网站(http://www.ece.rochester.edu/projects/air/projects/audiosearch).
英文摘要
Finding ways to automatically index, label, and access multimedia content (such as audio documents) is increasing in importance as multimedia repositories proliferate and grow. The community-generated SoundCloud repository is one example. It contains recordings of bands, sound effects, podcasts, etc., and contributors upload 12 hours of audio every minute. Repositories like SoundCloud typically tag audio at the file level with short text labels. Text-based search for a desired recording using these labels can be problematic. Text-based search within a track is not possible, since they are not indexed with tags in the body of the file. In this project, investigators at the University of Rochester and Northwestern University aim to develop methods and a system for audio search via query-by-example, where the example is similar, in some key way, to the desired audio in the database, but is not an exact match. This will allow search within files, bypassing the need for text-based tagging. This project will be focusing on using vocal imitations as search keys because they are natural for humans and are widely used in interaction. It will develop a novel search engine for sounds that takes vocal imitations as queries (e.g., imitation of a bird call to find recordings of the bird call). The technology developed for this novel way to search through audio/video collections will also benefit society in numerous other ways, such as crime surveillance (e.g., automated gunshot or scream detection for policy monitoring stations), biodiversity measurement (e.g., automatic ID of bird species that sound "like this" in field recordings), environmental awareness for the hearing impaired (e.g., alert me when my dog is the one barking), a production aid for a movie sound designer (e.g., finding door slam sounds in a database of thousands of sound effects), and sound-based diagnosis (e.g., "your car needs a new starter motor"). The project will benefit science technology engineering and mathematics (STEM) education as audio-based research has been shown to be a successful way to attract diverse college students into STEM disciplines.Vocal imitation conveys rich information covering many acoustic aspects: pitch, loudness, timbre, their temporal evolutions, and rhythmic patterns, etc. This lets a user query for precise sounds that are difficult to search for with text tags. For the same reason, however, vocal imitations may vary from the desired target on many dimensions. The query sound can also lies in a very constrained sound space compared to the sounds to be retrieved, due to the physical constraints of the human vocal system. Building a successful query-by-vocal-imitation system will require research into methods for representing audio and retrieving audio based on queries that are similar to target sounds only on a subset of their measurable dimensions. It will also require interfaces that facilitate providing queries and refining search results in a non-text-based context. For the former, the investigators will research on methods for learning of aspect-specific audio representations using deep neural networks. The investigators will also develop matching algorithms suitable for these representations. The investigators will design novel search interfaces that let users iteratively refine their search results. The system will learn from the interactions and adjust the weightings of different acoustic aspects to search for the wanted sound. Expected outcomes of this research are: (1) audio representations that highlight perceptually relevant features of vocal queries for matching to general audio target sounds; (2) algorithms for matching and aligning vocal queries to general audio; (3) interaction methods for iteratively refining search results using vocal imitations and sound examples; (4) a large vocal imitation and sound dataset; and (5) an open-source sound retrieval system that embodies these outcomes. More information about this project can be found at the project web site (http://www.ece.rochester.edu/projects/air/projects/audiosearch).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Human-Computer Collaborative Music Making
  • 批准号:
    1846184
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.92万
  • 财政年份:
    2019
  • 负责人:
    Zhiyao Duan
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: