Enhanced AI Perception through Unified Joint Embedding of Multimodal Sensory Data
Enhanced AI Perception through Unified Joint Embedding of Multimodal Sensory Data
批准号:
2874479
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
人工智能(AI)是一个旨在创造能够像人类一样思考和学习的机器的领域,它的子集深度学习(Deep Learning)使用多层神经网络来代表人类大脑,严重依赖于大量标记数据来训练机器。这种依赖关系通常会阻碍它们在动态、真实场景中的应用。相比之下,人类天生处理和交织多种感官——从听到城市声音的交响乐,到感受物体的精细纹理,再到视觉上解读距离。这种融合我们的感官和解读周围环境的自然能力可能是人工智能进化中缺失的一环。这就提出了一个核心的研究问题:多感官数据的整合能否缩小人类认知和机器学习之间的差距,从而使机器能够更多地从自然感官体验中学习,而从大量标记数据中学习得更少?本研究的目的如下:首先,本研究旨在开发能够从不同感官输入中提取结构的多模态计算模型。为了解决标记的多模态数据集的稀缺性,我们利用自然发生的成对数据来区分特定模态的信息并整合它们,从由视听线索发起的人类学习过程中汲取灵感(例如,我们将特定鸟类的声音与其照片相匹配)。其次,本研究试图通过结合新的感官模式(包括热数据、触觉信号和空间深度)来扩展人工智能的感知范围。这种整合旨在扩大人工智能的感知范围,超越视觉和音频的传统模式。第三,本研究旨在建立一个整体的感知系统。开发的多模态计算模型将被训练来同时处理广泛的感官模式,包括文本和视觉线索、听觉信号、空间深度、热感觉、IMU读数、触觉信号等。研究方法的新颖性如下。首先,本研究使用了对比学习。该技术帮助模型发现来自不同模态的数据点之间的相似性和差异性。因此,模型可以关联模式并跨多个模态建立连接。其次,本研究探索并关联了以往未被广泛研究的新的情态对,如视觉深度和视觉触觉。此外,本研究旨在超越传统的双模态嵌入,从多种自然共存的模式中发展出统一的嵌入景观。在这种情况下,嵌入是为机器解释简化复杂数据的基本数学表示。通过利用最先进的大型语言模型(llm)和视觉语言模型(vlm),我们的目标是识别特定于每种罕见模态的特征,匹配它们对,并在它们之间建立联系。llm / vlm的零学习能力促进了这一点,该功能使模型能够解释和执行它们以前从未遇到过的任务。因此,该模型可以更好地将来自各种感觉模态的信息编码到统一的、联合的嵌入空间中。该项目属于EPSRC的人工智能技术研究领域。我们渴望通过探索多模态感知数据的联合嵌入来提高人工智能的感知能力。我们的改进方法旨在开发一个更加互联和丰富的嵌入式空间,增强其对各种下游任务的适应性。例如,在机器人自动化领域,本研究衍生的统一、联合嵌入有可能显著提高机器人的感知能力,并彻底改变不同场景下的操作效率。
英文摘要
Artificial Intelligence (AI), a field aiming to create machines that can think and learn like humans, and its subset, Deep Learning, which uses multi-layered neural networks to represent human brains, rely heavily on large collections of labeled data to train the machines. This dependency often hinders their application in dynamic, real-world scenarios. In contrast, humans natively process and intertwine multiple senses - from hearing the symphony of urban sounds to feeling the object's fine textures and interpreting distances visually. This natural ability to blend our senses and interpret our surroundings can potentially be the missing link in AI's evolution. This raises the central research question: can the integration of multisensory data close the gap between human cognition and machine learning, so that machines can learn more from natural sensory experiences and less from extensive labeled data?The objectives of this research are as follows. First, this research aims to develop multimodal computational models capable of extracting structures from diverse sensory inputs. To address the scarcity of labeled multimodal datasets, we leverage naturally occurring paired data to distinguish between modality-specific information and integrate them, drawing inspiration from human learning processes initiated with audio-visual cues (e.g., we match the sound of a specific bird to its photo). Second, this research seeks to extend AI's perception range by incorporating novel sensory modalities, including thermal data, tactile signals, and spatial depth. This integration aims to augment AI's perceptual range beyond the conventional modalities of vision and audio. Third, this research aims to establish a holistic perception system. The developed multimodal computational model will be trained to process a wide array of sensory modalities concurrently, covering text and visual cues, auditory signals, spatial depth, thermal sensations, IMU readings, tactile signals, etc.The novelty of the research methodology is as follows. First, this research uses contrastive learning. This technique helps models to find similarities and differences across data points from different modalities. Consequently, the models can correlate patterns and build connections across multiple modalities. Second, this research explores and associates novel modality pairs such as visual-depth and visual-touch, which have not been extensively researched before. Moreover, this research aims to move beyond the traditional dual-modality embeddings and develops a unified embedding landscape from diverse naturally co-occurring modalities. In this context, embeddings are essential mathematical representations that simplify complex data for machine interpretation. By leveraging state-of-the-art Large Language Models (LLMs) and Vision-Language Models (VLMs), we aim to identify features specific to each rare modality, match the pairs, and establish connections between them. This is facilitated by the zero-shot learning capabilities of LLMs/VLMs, a feature that enables models to interpret and execute tasks they have never encountered before. Consequently, the model can better encode information from various sensory modalities into a unified, joint embedded space.This project falls within the EPSRC's Artificial Intelligence Technologies research area. We aspire to improve AI's perception through the exploration of joint embedding of multimodal sensory data. Our refined approach aims to develop a more interconnected and enriched embedded space, enhancing its adaptivity to a variety of downstream tasks. For example, in the field of robotic automation, the unified, joint embedding derived from this research have the potential to significantly improve robots' perception and revolutionize operational efficiency across different scenarios.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
面向AI驱动的信息化工程监管与自动化测试平台研发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:刘登志
-
依托单位:
建筑-音乐跨模态AI生成平台研发与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:许蕴彰
-
依托单位:
适用于AI眼镜的横向错位光学变焦系统技术开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:窦健泰
-
依托单位:
AI赋能中国传统壁画大模型开发与数字再生展示
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:朱亮亮
-
依托单位:
基于协同创新视角下AI赋能课程体系的模块化开发与应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:吴惠玲
-
依托单位:
AI赋能未成年人心理健康应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:傅绪荣
-
依托单位:
备多分AI智能研学系统开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:常直杨
-
依托单位:
带阻尼的弹簧型减振系统的虚拟建模、能控性分析及AI数智教育技术的开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:王成强
-
依托单位:
基于大数据分析与AI算力的民营教培企业提档升级内控管理系统研发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:卞禹臣
-
依托单位:
智能吊篮AI检测盒子开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:田申
-
依托单位: