课题基金 / 基金详情

场景实体有向图引导下基于听觉感知的视觉替代关键技术研究

批准号:
62103269
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
李恒
依托单位:
学科分类:
人工智能驱动的自动化
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
李恒

项目摘要

结项摘要

李恒的其他基金

相似基金

相关文献

中文摘要
失明是影响人类生活质量最严重的残疾之一,如何防盲、治盲、助盲已是全世界关注的公共卫生问题。随着人工智能技术的发展,利用盲人听觉来替代视觉的智能助盲系统已成为当前的研究热点。然而,目前的智能助盲系统大多是将基于机器视觉的目标检测和自然语言处理进行串联堆砌,反馈的信息局限于图像中的视觉元素,缺乏一定深度的推理能力。基于计算机视觉与自然语言处理领域交叉的视觉语义描述和视觉问答为解决这一问题提供了切实可行的技术途径,但实际应用中的视觉语义多样性描述和视觉问答可控性表达仍是当前亟需攻克的问题。为此,本项目拟开展场景实体有向图引导下基于听觉感知的视觉替代关键技术研究,构建视觉语义描述和视觉问答的知识表达架构与融合新方法;并通过心理物理学实验研究,探究盲人高效执行任务时的听觉感知特征规律,进而指导模型和听觉反馈策略的优化,为未来更具“智能”的视觉替代系统设计与实现提供关键的技术理论基础和科学的实验依据。
英文摘要
Blindness is one of the most serious disabilities that affect the quality of human life. How to avoid, treat, and assist blindness has become a public health issue in the current world. With the rapid development of artificial intelligence technology, intelligent blindness assistance systems that use blind people’s hearing for vision substitution have become a current research focus. However, current intelligent blindness assistance systems are relied on the technologies of machine-vision-based object detection and natural language processing, which are stacked in series. Feedback information thus is limited to the visual elements in the image, and it lacks a certain depth of reasoning ability. The visual semantic description and visual question answering technologies based on the intersection of computer vision and natural language processing provide practical technical approaches to deal with this problem. However, the description of visual semantic diversity and the controllable expression of visual question answering in practical applications are still the current highland to be occupied. In response to the above, the present project intends to carry out research on key technologies of auditory-perception-based visual substitution guided by scene entity directed graph, constructing new knowledge expression framework and fusion methods for visual semantic description and visual question answering. Furthermore, we will carry out psychophysical experimental research to explore the characteristics of auditory perception when blind people perform visual tasks efficiently, and then guide the adjustment and optimization of model algorithms and auditory feedback strategies. These will provide key technical theoretical supports and scientific experimental evidence for the future design and development of a ‘smarter’ noninvasive visual substitution system.
本项目聚焦非侵入式盲人视觉替代技术,围绕多样性语义描述和可控性视觉问答等关键技术展开研究,并针对在研究开展过程中发现的新问题,又相继深入开展了面向认知智能的多模态融合机制、影响盲人高效执行视觉任务时的影响因素、面向输入扰动的模型鲁棒性训练方法等基础研究。并在此基础上,将相关研究理论和成果扩展应用至智能化检测领域。.在多样性视觉语义描述方面,构建了基于视觉先验知识的显著性目标排序数据集与模型架构。通过多种方式引导视觉语义描述模型输出多样化内容,并提出基于强化学习的显著性目标排序方法,取得 SOTA 性能。相关成果在智能助盲和智能化检测领域成功应用,如帮助盲人更好理解周围环境、提高检测准确性等。在可控性视觉问答方面,深入研究多模态融合机制,剖析多种特征融合结构,发现解耦性高的结构在可控性视觉问答任务中性能更优,并开展了相关应用研究。同时探究了盲人执行视觉任务的影响因素,通过实验明确了提高助盲系统可用性的关键在于建立准确、可交互且信息量大的模型,并指出基于视频的问答系统是未来发展方向。还针对图像质量问题开展研究,提出鲁棒性目标检测方法和构建视觉问答数据集,提升模型在复杂图像条件下的性能,为盲人视觉辅助提供支持。此外,本课题还利用多物理场仿真软件和动物实验开展了无创/微创电刺激视觉替代技术的前沿探索,为视网膜疾病治疗和视觉功能修复提供理论支撑和潜在方案。.本项目研究成果为盲人视觉替代技术的发展提供了有力的推动,在提升盲人生活质量、促进其融入社会等方面具有重要意义,也为相关领域的进一步研究积累了宝贵经验与成果,有望在未来实现更广泛的应用与突破,持续助力盲人辅助技术的创新发展。
面向盲人视觉替代:基于多模态认知智能的视频“时、空”抗扰互补性关键技术研究
  • 批准号:
    --
  • 项目类别:
    省市级项目
  • 资助金额:
    0.0万元
  • 批准年份:
    2025
  • 负责人:
    李恒
  • 依托单位:
国内基金
海外基金