Knowledge Graph Completion for the Chinese Text of Cultural Relics Based on Bidirectional Encoder Representations from Transformers with Entity-Type Information.

Knowledge Graph Completion for the Chinese Text of Cultural Relics Based on Bidirectional Encoder Representations from Transformers with Entity-Type Information.
复制标题

基于带有实体类型信息的 Transformer 双向编码器表示的文物中文文本知识图谱补全

DOI:
10.3390/e22101168
复制
发表时间:
2020-10-16
期刊:
Entropy (Basel, Switzerland)
影响因子:
--
通讯作者:
Jia H
Jia H
中科院分区:
其他
文献类型:
--
作者:
Zhang M;Geng G;Zeng S;Jia H

文献摘要

参考文献

相似文献

知识图补全可以使知识图更加完整,这是一个有意义的研究课题。然而,现有的方法没有充分利用实体的语义信息。另一个挑战是深度模型需要大规模手动标记的数据,这大大增加了手工劳动。为了缓解文物领域标注数据的稀缺性,获取实体丰富的语义信息,提出了一种基于双向编码器表示(Bidirectional Encoder Representations from Transformers,BERT)和实体类型信息的文物中文文本知识图补全模型。在这项工作中,知识图完成任务被视为一个分类任务,而实体,关系和实体类型的信息被集成为一个文本序列,和汉字被用作一个令牌单元,其中的输入表示是通过求和令牌,段和位置嵌入构建。少量的标记数据用于预训练模型,然后,大量的未标记数据用于微调预训练模型。实验结果表明,引入实体类型信息的BERT-KGC模型能够丰富实体的语义信息,在一定程度上降低实体和关系的歧义程度,在使用35%的文物标注数据进行三重分类、链接预测和关系预测时,取得了比基线更好的性能。
Knowledge graph completion can make knowledge graphs more complete, which is a meaningful research topic. However, the existing methods do not make full use of entity semantic information. Another challenge is that a deep model requires large-scale manually labelled data, which greatly increases manual labour. In order to alleviate the scarcity of labelled data in the field of cultural relics and capture the rich semantic information of entities, this paper proposes a model based on the Bidirectional Encoder Representations from Transformers (BERT) with entity-type information for the knowledge graph completion of the Chinese texts of cultural relics. In this work, the knowledge graph completion task is treated as a classification task, while the entities, relations and entity-type information are integrated as a textual sequence, and the Chinese characters are used as a token unit in which input representation is constructed by summing token, segment and position embeddings. A small number of labelled data are used to pre-train the model, and then, a large number of unlabelled data are used to fine-tune the pre-training model. The experiment results show that the BERT-KGC model with entity-type information can enrich the semantics information of the entities to reduce the degree of ambiguity of the entities and relations to some degree and achieve more effective performance than the baselines in triple classification, link prediction and relation prediction tasks using 35% of the labelled data of cultural relics.
DOI: 10.1109/tkde.2017.2754499
发表时间: 2017-12-01
影响因子: 8.9
作者:
Wang, Quan;Mao, Zhendong;Guo, Li
通讯作者: Guo, Li
DOI: 10.1002/asi.23837
发表时间: 2017-08-01
影响因子: 3.5
作者:
Minkov, Einat;Kahanov, Keren;Kuflik, Tsvi
通讯作者: Kuflik, Tsvi