Expanding Science and Technology Thesauri from Bibliographic Datasets Using Word Embedding

Expanding Science and Technology Thesauri from Bibliographic Datasets Using Word Embedding
复制标题

DOI:
10.1109/ictai.2016.0133
复制
发表时间:
2016-11
期刊:
2016 IEEE 28th International Conference on Tools with Artificial Intelligence (ICTAI)
影响因子:
--
通讯作者:
Takahiro Kawamura;Kouji Kozaki;Tatsuya Kushida;Katsutaro Watanabe;Katsuji Matsumura
Takahiro Kawamura;Kouji Kozaki;Tatsuya Kushida;Katsutaro Watanabe;Katsuji Matsumura
中科院分区:
其他
文献类型:
--
作者:
Takahiro Kawamura;Kouji Kozaki;Tatsuya Kushida;Katsutaro Watanabe;Katsuji Matsumura

文献摘要

相似文献

科学计量学中科技信息的叙词表和分类法的使用引起了人们的注意。然而,人工构建和维护叙词表的成本很高,需要大量的时间,因此,人们正在积极研究半自动构建和维护的方法。我们提出了一种方法,使用有限结构化信息的最先进技术领域的文章摘要来扩展现有的同义词词典。具体地说,我们考虑了一种使用快速演变的词嵌入将新术语适当地分配到现有同义词词典的层次结构的方法。在一项实验中,从567,000篇生物医学文章中构建了500度的单词向量,并使用主成分分析进行降维后对其进行聚类。然后,基于新术语与叙词表中的任何术语之间的空间关系来估计语义关系。然后,我们对来自三位专家的结果进行了比较。未来,我们将开发一个与现有术语相关的新术语推荐系统,以支持半自动的词库维护。
The use of thesauri and taxonomies for science and technology information in scientometrics has been attracting attention. However, manual construction and maintenance of thesauri is expensive and requires significant time, thus, methods for semi-automatic construction and maintenance are being actively studied. We propose a method to expand an existing thesaurus using the abstracts of articles from state-of-the-art technological domains with limited structured information. Specifically, we consider a method for properly allocating new terms to the hierarchical structures of an existing thesaurus using rapidly evolving word embedding. In an experiment, word vectors of 500 degrees are constructed from 567,000 biomedical articles and are clustered after dimension reduction using principal component analysis. Then, semantic relations are estimated based on the spatial relations between the new term and any of the terms in the thesaurus. We then conducted a comparison of the results obtained from three experts. In future, we will develop a recommendation system for new terms related to the existing terms to support semi-automatic thesaurus maintenance.