CN-DBpedia: A Never-Ending Chinese Knowledge Extraction System

CN-DBpedia: A Never-Ending Chinese Knowledge Extraction System
复制标题

DOI:
10.1007/978-3-319-60045-1_44
复制
发表时间:
2017-06
期刊:
--
影响因子:
--
通讯作者:
Bo Xu;Yong Xu;Jiaqing Liang;Chenhao Xie;Bin Liang;Wanyun Cui;Yanghua Xiao
Bo Xu;Yong Xu;Jiaqing Liang;Chenhao Xie;Bin Liang;Wanyun Cui;Yanghua Xiao
中科院分区:
其他
文献类型:
--
作者:
Bo Xu;Yong Xu;Jiaqing Liang;Chenhao Xie;Bin Liang;Wanyun Cui;Yanghua Xiao

文献摘要

被引文献

相似文献

人们一直在努力从在线百科全书中获取知识库。这些知识库在使机器理解文本方面发挥着重要作用。然而,目前大多数知识库都是英文的,非英文的知识库,尤其是中文知识库,还非常少。许多以前的系统,从在线百科全书中提取知识,虽然适用于建立一个中文知识库,仍然受到两个挑战。首先,它需要大量的人力来构建本体和建立一个监督的知识提取模型。二是知识库的更新频率很慢。为了解决这些问题,我们提出了一个永不停止的中文知识抽取系统,CN-DBpedia,它可以自动生成一个知识库,这是不断增加的规模和不断更新。特别是,我们通过重用现有知识库的本体和建立一个端到端的事实抽取模型来减少人力成本。我们进一步提出了一个智能的主动更新策略,以保持我们的知识库的新鲜度,很少的人力成本。发布的服务的1.64亿个API调用证明了我们系统的成功。
Great efforts have been dedicated to harvesting knowledge bases from online encyclopedias. These knowledge bases play important roles in enabling machines to understand texts. However, most current knowledge bases are in English and non-English knowledge bases, especially Chinese ones, are still very rare. Many previous systems that extract knowledge from online encyclopedias, although are applicable for building a Chinese knowledge base, still suffer from two challenges. The first is that it requires great human efforts to construct an ontology and build a supervised knowledge extraction model. The second is that the update frequency of knowledge bases is very slow. To solve these challenges, we propose a never-ending Chinese Knowledge extraction system,CN-DBpedia, which can automatically generate a knowledge base that is of ever-increasing in size and constantly updated. Specially, we reduce the human costs by reusing the ontology of existing knowledge bases and building an end-to-end facts extraction model. We further propose a smart active update strategy to keep the freshness of our knowledge base with little human costs. The 164 million API calls of the published services justify the success of our system.