Automatic Translation of Scholarly Terms into Patent Terms Using Synonym Extraction Techniques

Automatic Translation of Scholarly Terms into Patent Terms Using Synonym Extraction Techniques
复制标题

DOI:
--
复制
发表时间:
2012-05
期刊:
--
影响因子:
--
通讯作者:
Hidetsugu Nanba;T. Takezawa;Kiyoko Uchiyama;Akiko Aizawa
Hidetsugu Nanba;T. Takezawa;Kiyoko Uchiyama;Akiko Aizawa
中科院分区:
其他
文献类型:
--
作者:
Hidetsugu Nanba;T. Takezawa;Kiyoko Uchiyama;Akiko Aizawa

文献摘要

相似文献

检索研究论文和专利对于任何研究人员评估具有高度工业相关性的领域范围都很重要。然而,专利中使用的术语往往比研究论文中使用的术语更抽象或更具创造性,因为它们旨在扩大权利要求的范围。因此,需要一种将学术术语翻译成专利术语的方法。在本文中,我们提出了六种方法翻译学术术语到专利术语使用两种同义词提取方法:统计机器翻译(SMT)为基础的方法和分布相似性(DS)为基础的方法。我们使用NTCIR-7研讨会的专利挖掘任务数据集进行了实验,以确认我们的方法的有效性。本研究的目的是使用IPC系统对日语研究论文(标题和摘要对)进行子类(第三级)、主组(第四级)和子组(第五级和最详细的级别)的分类。结果表明,基于SMT的方法(SMT_ABST+IDF)在子组水平上表现最好,而基于DS的方法(DS+IDF)在子类水平上表现最好。
Retrieving research papers and patents is important for any researcher assessing the scope of a field with high industrial relevance. However, the terms used in patents are often more abstract or creative than those used in research papers, because they are intended to widen the scope of claims. Therefore, a method is required for translating scholarly terms into patent terms. In this paper, we propose six methods for translating scholarly terms into patent terms using two synonym extraction methods: a statistical machine translation (SMT)-based method and a distributional similarity (DS)-based method. We conducted experiments to confirm the effectiveness of our method using the dataset of the Patent Mining Task from the NTCIR-7 Workshop. The aim of the task was to classify Japanese language research papers (pairs of titles and abstracts) using the IPC system at the subclass (third level), main group (fourth level), and subgroup (the fifth and most detailed level). The results showed that an SMT-based method (SMT_ABST+IDF) performed best at the subgroup level, whereas a DS-based method (DS+IDF) performed best at the subclass level.