On Efficiency of Semantic Relation Extraction through Low-dimensional Distributed Representations for Substrings

On Efficiency of Semantic Relation Extraction through Low-dimensional Distributed Representations for Substrings
复制标题

DOI:
10.1109/hpcc-css-icess.2015.267
复制
发表时间:
2015-08
期刊:
2015 IEEE 17th International Conference on High Performance Computing and Communications, 2015 IEEE 7th International Symposium on Cyberspace Safety and Security, and 2015 IEEE 12th International Conference on Embedded Software and Systems
影响因子:
--
通讯作者:
Z. Jin;Chihiro Shibata;Jingtao Sun;Kazuya Tago
Z. Jin;Chihiro Shibata;Jingtao Sun;Kazuya Tago
中科院分区:
其他
文献类型:
--
作者:
Z. Jin;Chihiro Shibata;Jingtao Sun;Kazuya Tago

文献摘要

相似文献

凭借机器学习技术的最新发展,现在可以从大数据中提取更高级别的信息。为了分析大数据,通过使用足够快的算法以及高度准确的结果来实现数据的高效和智能表示非常重要。本文通过对文本数据中子串的高效低维表达,采用轻量化处理方法提取多种语义关系。我们提出了一种方法来建立功能的关系分类组成的只有低维向量表示两个词之间的子串,称为子串向量。实验结果表明,使用有效的低维表示的数据,并在一个小的计算成本,我们的方法实现了足够高的准确性,优于大多数现有的方法。此外,通过实验,我们确保了将子串映射到足够低维的空间在准确性和效率方面都会产生更好的结果。
By virtue of recent developments in machine learning techniques, higher-level information can now to be extracted from big data. To analyze big data, efficient and smart representations of data achieved by using sufficiently fast algorithms, as well as highly accurate results, are important. In this paper, we focus on extracting multiple semantic relations using light-weight processing through the efficient low-dimensional expression of substrings in text data. We propose an approach to build features for relation classification consisting of only low-dimensional vectors representing substrings between two words, called substring vectors. The experimental results show that, using efficient low-dimensional representations of data and at a small computational cost, our approach achieves a sufficiently high accuracy that is better than most existing approaches. In addition, through experiments, we ensured that mapping substrings to a sufficiently low dimensional space yields better results in terms of both accuracy and efficiency.