课题基金 / 基金详情

CISE-ANR: HCC: Small: Omnidirectional BatVision: Learning How to Navigate from Cell Phone Audios

CISE-ANR: HCC: Small: Omnidirectional BatVision: Learning How to Navigate from Cell Phone Audios
CISE-ANR:HCC:小型:全向 BatVision:学习如何通过手机音频进行导航
批准号:
2215542
负责人:
Stella Yu
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-04-01 至 2026-03-31

项目摘要

项目成果

Stella Yu的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的目的是开发实时三维空间重建从声音捕捉不是由昂贵的专业设备,而是由普通的消费级移动的手机。 这种方法的灵感来自蝙蝠使用的回声定位,是从声音中开发出足以用于导航的3D空间地图,例如近距离避障和在拥挤的火车站中寻找远处的出口。 这项研究将使3D视觉超越视线,在低光或无光条件下的应用范围从听汽车,可以听到行人在拐角处的集体3D地图重建人群。 项目成果将有助于更好地对机器人的声音感知和有效的声音-视觉集成进行计算建模,并有助于有影响力的应用,例如视障人士的导航辅助设备和烟雾或黑暗造成的低能见度条件下的消防员。 这项工作将为视觉3D映射提供一个互补的具有成本效益的替代方案,使每个人都能成为3D内容创作者。 虽然立体声音频为水平到达方向估计提供了直接线索,但它仅在良好控制的环境中起作用。 在现实世界中,没有简单的数学模型可以将声音映射到3D空间,因为许多因素,如设备方向,房间布局,材料,背景噪音都会影响声音的传播。 该项目采用机器学习方法从手机音频中推断3D空间。 将使用带有双耳麦克风、扬声器和RGB-D立体声的传感器装置在不同环境中收集大规模视听数据集。 一个附加的智能手机将记录时间同步的数据与自己的立体声麦克风和摄像头。扬声器将发出信号以实现回声定位,但部分数据将仅包含自然发生的声音。 几个室内和室外环境与激光雷达扫描的3D模型将作为地面实况。 还将在公共街道上收集数据,以测试在无法进行激光雷达扫描的现实情况下的鲁棒性。 给定数据集,将针对视场和全360°视图制定几个3D场景重建任务,首先使用特权传感器数据,最后仅使用手机传感器。 在使用双耳麦克风和立体声摄像机在各种环境中收集大规模视听数据后,将训练模型以将声音数据映射到从视觉数据中提取的深度图。 一旦模型经过训练,它将能够仅基于声音输入“看到”3D空间。 然后,该模型将通过立体声麦克风和移动的手机上的传感器进行调整,以实现同样高质量的3D感知。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project aims to develop real-time 3D space reconstruction from sound captured not by expensive specialized equipment, but by common-place consumer-grade mobile phones. The approach, inspired by echolocation used by bats, is to develop from sound alone 3D spatial maps that are sufficient for navigation, such as close obstacle avoidance and finding distant exits in a crowded train station. The research will enable 3D vision beyond the line of sight and in low or no light conditions with applications ranging from listening cars that can hear pedestrians around the corner to collective 3D map reconstruction from crowds. Project outcomes will contribute to better computational modeling of sound perception and effective sound-vision integration in robotics, as well as to impactful applications such as navigational aids for visually impaired persons and for fire-fighters in low visibility conditions caused by smoke or darkness. The work will provide a complementary cost-effective alternative to visual 3D mapping that allows everybody to become a 3D content creator.The task of 3D perception from sound is challenging. While stereo audio provides direct cues for horizontal direction of arrival estimation, it only works in well controlled environments. There are no simple mathematical models to map sound to 3D space in real-word settings, as many factors such as device orientations, room layouts, materials, background noises shape sound propagation. This project takes a machine learning approach to infer 3D space from cell phone audios. A large-scale audio-visual dataset will be collected in different environments using a sensor-rig with a binaural microphone, a speaker and an RGB-D stereo. An attached smartphone will record time-synchronized data with its own stereo microphone and cameras. The speaker will emit signals to enable echolocation, but a part of the data will contain only naturally occurring sounds. Several indoor and outdoor environments with LiDAR scanned 3D models will serve as ground-truth. Data will also be collected in public streets to test robustness in realistic situations where LiDAR scans are not possible. Given the dataset, several 3D scene reconstruction tasks will be formulated for both the field of view and full 360° view, first with privileged sensor data and finally from cellphone sensors alone. After collecting large-scale audio-visual data in a variety of environments with binaural microphones and stereo cameras, a model will be trained to map sound data to depth maps extracted from visual data. Once the model is trained, it will be able to “see” the 3D space based on sound inputs alone. The model will then be adapted to achieve the same high quality 3D perception with stereo-microphones and sensors available on a mobile phone.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: RI: Medium: Lie group representation learning for vision
CAREER: Art and Vision: Scene Layout from Pictorial Cues
CAREER: Art and Vision: Scene Layout from Pictorial Cues
  • 批准号:
    0644204
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2007
  • 负责人:
    Stella Yu
  • 依托单位:
国内基金
海外基金
花青素还原酶(ANR)在荔枝果皮褐变底物积累中的作用
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2021
  • 负责人:
    方方
  • 依托单位:
ANR与LAR在茶树表型儿茶素生物合成中的作用机制研究
  • 批准号:
    31902070
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2019
  • 负责人:
    王培强
  • 依托单位: