An unsupervised approach for learning a Chinese IS-A taxonomy from an unstructured corpus

An unsupervised approach for learning a Chinese IS-A taxonomy from an unstructured corpus
复制标题

一种从非结构化语料库学习中文 IS-A 分类的无监督方法

DOI:
10.1016/j.knosys.2019.07.032
复制
发表时间:
2019-10
影响因子:
8.8
通讯作者:
Shengwei Gu
Shengwei Gu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Subin Huang;Xiangfeng Luo;Jing Huang;Yike Guo;Shengwei Gu

文献摘要

参考文献

被引文献

相似文献

分类法在各种自然语言处理(NLP)任务(例如,文本分类、信息提取和知识推理)。然而,由于中文自然语言的复杂性和灵活性,准确地从非结构化语料库中学习中文IS-A(C-IS-A)分类是具有挑战性的。在本文中,我们提出了一个无监督的C-IS-A分类学习方法,通过分析一个给定的非结构化语料库。我们的方法使用三个主要步骤来自动学习C-IS-A分类法。首先,我们的方法提取高质量的C-IS-A种子关系,通过语义迭代模式为基础的匹配和句法方法。其次,我们的方法利用了一个无监督的分类语义团为基础的方法,以增加C-IS-A分类的覆盖范围。作为我们的方法的核心组成部分,我们利用提取的C-IS-A种子关系构建分类语义集团,并使用集团的上下文和多概念共现信息来推断潜在的新的C-IS-A关系。最后,提出了一种两步关系检测策略,以消除潜在的不正确的C-IS-A关系,这可以大大提高学习分类的准确性。我们在四个中文非结构化语料库上实现了我们的方法,并从精度、覆盖率、时间成本和子成分的效果等方面对其进行了评估。评估结果表明,我们的方法是一种有效的方法,优于国家的最先进的比较方法。
Taxonomies play an important role in various Natural Language Processing (NLP) tasks (e.g., text classification, information extraction and knowledge inference). However, due to the complexity and flexibility of the Chinese natural language, it is challenging to accurately learn a Chinese IS-A (C-IS-A) taxonomy from an unstructured corpus. In this paper, we propose an unsupervised C-IS-A taxonomy learning approach by analyzing a given unstructured corpus. Our approach uses three main steps to automatically learn a C-IS-A taxonomy. First, our approach extracts high-quality C-IS-A seed relations via semantic iterative pattern-based matching and syntactic methods. Second, our approach utilizes an unsupervised taxonomic semantic-clique-based method to increase the coverage of the C-IS-A taxonomy. As the core component of our approach, we exploit the extracted C-IS-A seed relations to construct taxonomic semantic cliques and use the context of the cliques and multi-concept co-occurrence information to infer potential novel C-IS-A relations. Last, a two-step relation detection strategy is proposed to remove potentially incorrect C-IS-A relations, which can substantially improve the accuracy of the learned taxonomy. We implement our approach on four Chinese unstructured corpora and evaluate it in terms of precision, coverage, time cost and the effects of the subcomponents. The evaluation results demonstrate that our approach is an effective method that outperforms the state-of-the-art compared approaches.
DOI: 10.1007/s10994-017-5638-4
发表时间: 2017-08
期刊: Machine Learning
影响因子: 7.5
作者:
Junyu Xuan;Jie Lu;Guangquan Zhang;R. Xu;Xiangfeng Luo
通讯作者: Junyu Xuan;Jie Lu;Guangquan Zhang;R. Xu;Xiangfeng Luo
DOI: 10.1016/j.artint.2011.01.003
发表时间: 2011-06
期刊: Artif. Intell.
影响因子: --
作者:
Simone Paolo Ponzetto;M. Strube
通讯作者: Simone Paolo Ponzetto;M. Strube
DOI: 10.1016/j.websem.2009.07.002
发表时间: 2009-09-01
影响因子: 2.5
作者:
Bizer, Christian;Lehmann, Jens;Hellmann, Sebastian
通讯作者: Hellmann, Sebastian
DOI: 10.1007/978-3-319-60045-1_44
发表时间: 2017-06
期刊: --
影响因子: --
作者:
Bo Xu;Yong Xu;Jiaqing Liang;Chenhao Xie;Bin Liang;Wanyun Cui;Yanghua Xiao
通讯作者: Bo Xu;Yong Xu;Jiaqing Liang;Chenhao Xie;Bin Liang;Wanyun Cui;Yanghua Xiao
DOI: 10.1016/j.websem.2018.05.002
发表时间: 2018-05
期刊: J. Web Semant.
影响因子: --
作者:
Tianxing Wu;Haofen Wang;G. Qi;Jiangang Zhu;Tong Ruan
通讯作者: Tianxing Wu;Haofen Wang;G. Qi;Jiangang Zhu;Tong Ruan