课题基金 / 基金详情

Negative Knowledge at Web Scale

Negative Knowledge at Web Scale
网络规模的负面知识
批准号:
453095897
负责人:
Dr. Simon Razniewski
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2021
资助国家:
德国
项目状态:
已结题
起止时间:
2020-12-31 至 2023-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
结构化知识在问题回答、对话或推荐系统等一系列应用中至关重要。所需的知识通常存储在知识库(KBS)中,近年来人们对知识库的构建、查询和维护产生了越来越大的兴趣。一些知识库侧重于词汇信息,另一些侧重于地理空间知识、活动或常识。但最突出的是,KBS获取百科全书式的知识,值得注意的项目是维基数据、数据库百科或谷歌知识图谱。这些知识库存储着一些正面的说法,如“萨尔布吕肯是萨尔州的首府”,是许多知识密集型人工智能应用的关键资产。所有这些知识库的一个主要限制是它们无法处理负面信息。目前,所有主要知识库都只包含正面信息,而像汤姆·克鲁斯没有获得奥斯卡奖这样的说法只能通过需要大量假设的推理来推断。由于KBS通常只包含真值的子集,用户经常不得不猜测KBS中未包含的信息是假的,还是KBS仅仅不知道真相。不能正式区分一项陈述是错误的还是未知的,这在各种应用程序中都会带来挑战。例如,在医学上,区分知道物质之间没有生化反应和根本不知道物质的存在是很重要的。在企业诚信方面,重要的是知道某人是否从未受雇于某个竞争对手,而在反腐败调查中,需要确定没有家庭关系。在(假)新闻领域,真相不明的谣言(如马航370被劫持)和被确定为虚假的谣言(如奥巴马出生在肯尼亚)之间存在着重要的区别。尽管负面信息在逻辑和数据库理论中受到了极大的关注,但目前网络规模的知识库中仍然没有负面信息。例如,Wikidata、DBpedia和Yago都只包含正面信息,充其量也只能通过模式约束进行有限的否定推断。同样,到目前为止,文本提取和统计推断只处理了积极的信息。在这个项目中,我们的目标是通过研究克服目前知识库对积极信息的限制,研究包括三个组成部分:(1)用于生成负面信息的统计推理技术,(2)用于解决矛盾和不一致的网络验证和联合合并技术,以及(3)允许检索在特定用例中相关的负面信息的排名技术。
英文摘要
Structured knowledge is crucial in a range of applications such as question answering, dialogue or recommender systems. The required knowledge is usually stored in knowledge bases (KBs), and recent years have seen a rise of interest in KB construction, querying and maintenance. Some KBs focus on lexical information, others on geospatial knowledge, activities, or common sense. But most prominently, KBs capture encyclopedic knowledge, with notable projects being Wikidata, DBpedia, or the Google Knowledge Graph. These KBs store positive statements such as “Saarbrücken is the capital of the Saarland”, and are a key asset for many knowledge-intensive AI applications.A major limitation of all these KBs is their inability to deal with negative information. At present, all major knowledge bases only contain positive information, whereas statements such as that Tom Cruise did not win an Oscar can only be deduced by inferences that require substantial assumptions. As KBs generally only contain subsets of what is true, users often have to guess whether information not contained in a KBs is false, or truth is merely unknown to the KB. Not being able to formally distinguish whether a statement is false or unknown poses challenges in a variety of applications. In medicine, for instance, it is important to distinguish between knowing about the absence of a biochemical reaction between substances, and not knowing about its existence at all. In corporate integrity, it is important to know whether a person was never employed by a certain competitor, while in anti-corruption investigations, absence of family relations needs to be ascertained. In the domain of (fake) news, there is an important distinction between rumors whose truth is unknown (such as “Malayan Airlines 370 was hijacked”), and those established to be false (“Obama was born in Kenya”).While negative information has received great attention in logics and database theory, it is still absent from current web-scale knowledge bases. For instance, Wikidata, DBpedia and YAGO all only contain positive information, and at best allow limited inferences about negation via schema constraints. Similarly, text extraction and statistical inferences so far have only tackled positive information. In this project we aim to overcome the current restriction of knowledge bases to positive information by research that encompasses three components: (i) statistical inferencing techniques for generating negative information, (ii) web-validation and joint consolidation techniques for resolving contradictions and inconsistencies, and (iii) ranking techniques that allow to retrieve negative information as relevant in specific use cases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金