课题基金 / 基金详情

Combining Semantic Technologies and Machine Learning for the (semi-) automatic annotation of data.

Combining Semantic Technologies and Machine Learning for the (semi-) automatic annotation of data.
结合语义技术和机器学习来(半)自动注释数据。
批准号:
2248787
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
背景:识别不同数据格式的真实实体及其结构关系对于包括数据集成、数据清理、数据挖掘和知识发现在内的许多应用程序至关重要。在许多领域,由于元数据丢失、不完整或混淆,这仍然是一项非常耗费人力的任务。全球数据量的增加使得手动注释数据不再可行,需要更具可扩展性的解决方案。因此,用于(半自动)数据标注的新技术将在这些领域中增加许多价值。目标/目标:该项目的目标是研究使用描述底层结构和语义的有意义的元数据对数据进行(半自动)注释。牛津大学信息系统小组的初步研究表明,语义技术(如大型在线知识图和本体)与最先进的机器学习技术(如深度神经网络和语义嵌入)的组合为关系数据的标注带来了非常有希望的结果。这表明了关于关系数据上更复杂的结构关系的注释以及其他结构化、半结构化或非结构化数据格式的注释的进一步研究的巨大潜力。研究方法的新颖性:只有数量有限的出版物使用语义技术(例如在线知识图谱和本体)和深度学习(例如深度神经网络和语义嵌入)来自动标注数据。此外,这些工作主要集中在关系数据上。技术的结合和对不同数据格式的应用构成了该项目独有的一种新的研究方法。与EPSRC的战略和研究领域(项目涉及的EPSRC研究领域)保持一致:“本项目属于EPSRC信息和通信技术研究领域”,尤其与以下主题相关:信息系统数据库机器学习/人工智能公司合作者:牛津的信息系统小组与慕尼黑的西门子研究小组就该项目进行合作。西门子团队在这一领域贡献了现实的用例以及大量的研究专业知识和资源。
英文摘要
Context:Identifying real-world entities and their structural relations in different data formats is of crucial importance for many applications including data integration, data cleaning, data mining, and knowledge discovery. In many domains, this is still a very labour-intensive task due to missing, incomplete or obfuscated metadata. The increase in the global data volume makes the manual annotation of data no longer feasible and requires a more scalable solution. New techniques for the (semi-) automatic data annotation would, therefore, add a lot of value in these domains. Goals/Objectives:The goal of the project is to investigate the (semi-) automatic annotation of data with meaningful metadata that describes the underlying structure and semantics. Preliminary research from the Information Systems group at the University of Oxford has shown that the combination of semantic technologies (such as large online knowledge graphs and ontologies) and state-of-the-art machine learning techniques (such as deep neural networks and semantic embedding) delivered very promising results for the annotation of relational data. This indicates great potential for further research regarding the annotation of more complex structural relationships on relational data as well as the annotation of other structured, semi-structured or unstructured data formats. Novelty of the research methodology:There exists only a limited number of publications that use a combination of semantic technologies (e.g. online knowledge graphs and ontologies) and deep learning (e.g. deep neural networks and semantic embedding) for the automatic annotation of data. Additionally, most of this work focuses on relational data. The combination of technologies and the application on different data formats constitute a novel research methodology unique to this project. Alignment to EPSRC's strategies and research areas (which EPSRC research area the project relates to):'This project falls within the EPSRC Information and Communication Technologies research area'It is especially related to topics such as:Information SystemsDatabasesMachine Learning/Artificial Intelligence Company collaborators:The Information Systems group in Oxford works in collaboration with the Siemens research group in Munich on this project. The Siemens team contributes realistic use-cases as well as considerable research expertise and resources in this area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金