课题基金 / 基金详情

Negative Knowledge at Web Scale

Negative Knowledge at Web Scale
网络规模的负面知识
批准号:
453095897
负责人:
Dr. Simon Razniewski
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2021
资助国家:
德国
项目状态:
已结题
起止时间:
2020-12-31 至 2023-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
结构化知识在问答、对话或推荐系统等一系列应用中至关重要。所需的知识通常存储在知识库(知识库)中,近年来人们对知识库的构建、查询和维护越来越感兴趣。一些KBs侧重于词汇信息,另一些侧重于地理空间知识、活动或常识。但最突出的是,KBs获取了百科全书式的知识,其中著名的项目是Wikidata、DBpedia或b谷歌知识图谱。这些知识库存储了诸如“saarbrcken是Saarland的首府”之类的正面陈述,并且是许多知识密集型人工智能应用程序的关键资产。所有这些KBs的一个主要限制是它们无法处理负面信息。目前,所有主要的知识库都只包含积极的信息,而像汤姆克鲁斯没有赢得奥斯卡这样的陈述只能通过需要大量假设的推论来推断出来。由于KBs通常只包含真实内容的子集,用户经常不得不猜测KBs中不包含的信息是假的,还是知识库不知道真相。不能正式区分一个陈述是假的还是未知的,在各种应用中都带来了挑战。例如,在医学上,区分知道物质之间没有生化反应和根本不知道它的存在是很重要的。在企业诚信方面,重要的是要知道一个人是否从未受雇于某个竞争对手,而在反腐败调查中,需要确定是否存在家庭关系。在(假)新闻领域,真相不明的谣言(如“马来亚航空370航班被劫持”)和已被证实的谣言(如“奥巴马出生在肯尼亚”)之间有一个重要的区别。虽然负面信息在逻辑学和数据库理论中受到了很大的关注,但在目前的网络知识库中仍然缺乏负面信息。例如,Wikidata, DBpedia和YAGO都只包含正面信息,并且最多允许通过模式约束对否定进行有限的推断。同样,到目前为止,文本提取和统计推断只处理了积极的信息。在这个项目中,我们的目标是通过研究克服目前知识库对正面信息的限制,该研究包括三个组成部分:(i)生成负面信息的统计推理技术,(ii)解决矛盾和不一致的网络验证和联合整合技术,以及(iii)允许检索负面信息的排名技术。
英文摘要
Structured knowledge is crucial in a range of applications such as question answering, dialogue or recommender systems. The required knowledge is usually stored in knowledge bases (KBs), and recent years have seen a rise of interest in KB construction, querying and maintenance. Some KBs focus on lexical information, others on geospatial knowledge, activities, or common sense. But most prominently, KBs capture encyclopedic knowledge, with notable projects being Wikidata, DBpedia, or the Google Knowledge Graph. These KBs store positive statements such as “Saarbrücken is the capital of the Saarland”, and are a key asset for many knowledge-intensive AI applications.A major limitation of all these KBs is their inability to deal with negative information. At present, all major knowledge bases only contain positive information, whereas statements such as that Tom Cruise did not win an Oscar can only be deduced by inferences that require substantial assumptions. As KBs generally only contain subsets of what is true, users often have to guess whether information not contained in a KBs is false, or truth is merely unknown to the KB. Not being able to formally distinguish whether a statement is false or unknown poses challenges in a variety of applications. In medicine, for instance, it is important to distinguish between knowing about the absence of a biochemical reaction between substances, and not knowing about its existence at all. In corporate integrity, it is important to know whether a person was never employed by a certain competitor, while in anti-corruption investigations, absence of family relations needs to be ascertained. In the domain of (fake) news, there is an important distinction between rumors whose truth is unknown (such as “Malayan Airlines 370 was hijacked”), and those established to be false (“Obama was born in Kenya”).While negative information has received great attention in logics and database theory, it is still absent from current web-scale knowledge bases. For instance, Wikidata, DBpedia and YAGO all only contain positive information, and at best allow limited inferences about negation via schema constraints. Similarly, text extraction and statistical inferences so far have only tackled positive information. In this project we aim to overcome the current restriction of knowledge bases to positive information by research that encompasses three components: (i) statistical inferencing techniques for generating negative information, (ii) web-validation and joint consolidation techniques for resolving contradictions and inconsistencies, and (iii) ranking techniques that allow to retrieve negative information as relevant in specific use cases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金