Building a Non-Trivial Paraphrase Corpus Using Multiple Machine Translation Systems

Building a Non-Trivial Paraphrase Corpus Using Multiple Machine Translation Systems
复制标题

DOI:
10.18653/v1/p17-3007
复制
发表时间:
2017-07
期刊:
--
影响因子:
--
通讯作者:
Yui Suzuki;Tomoyuki Kajiwara;Mamoru Komachi
Yui Suzuki;Tomoyuki Kajiwara;Mamoru Komachi
中科院分区:
其他
文献类型:
--
作者:
Yui Suzuki;Tomoyuki Kajiwara;Mamoru Komachi

文献摘要

相似文献

我们提出了一种新的句子释义习得方法。为了建立一个平衡良好的意译识别语料库,我们特别关注获取非平凡的积极和消极实例。我们使用多个机器翻译系统生成正面候选词,并使用单语语料库提取负面候选词。为了收集非平凡的实例,候选对象按单词重叠率统一采样。最后,注释者判断候选词是积极的还是消极的。利用该方法构建并发布了首个日语释义识别评价语料库,该语料库包含655个句子对。
We propose a novel sentential paraphrase acquisition method. To build a well-balanced corpus for Paraphrase Identifi-cation, we especially focus on acquiring both non-trivial positive and negative instances. We use multiple machine translation systems to generate positive candidates and a monolingual corpus to extract negative candidates. To collect non-trivial instances, the candidates are uniformly sampled by word overlap rate. Finally, annotators judge whether the candidates are either positive or negative. Using this method, we built and released the first evaluation corpus for Japanese paraphrase identification, which comprises 655 sentence pairs.