Building a Non-Trivial Paraphrase Corpus Using Multiple Machine Translation Systems
Building a Non-Trivial Paraphrase Corpus Using Multiple Machine Translation Systems
复制标题
DOI:
10.18653/v1/p17-3007
复制
发表时间:
2017-07
期刊:
影响因子:
--
通讯作者:
Yui Suzuki;Tomoyuki Kajiwara;Mamoru Komachi
中科院分区:
文献类型:
--
作者:
Yui Suzuki;Tomoyuki Kajiwara;Mamoru Komachi
We propose a novel sentential paraphrase acquisition method. To build a well-balanced corpus for Paraphrase Identifi-cation, we especially focus on acquiring both non-trivial positive and negative instances. We use multiple machine translation systems to generate positive candidates and a monolingual corpus to extract negative candidates. To collect non-trivial instances, the candidates are uniformly sampled by word overlap rate. Finally, annotators judge whether the candidates are either positive or negative. Using this method, we built and released the first evaluation corpus for Japanese paraphrase identification, which comprises 655 sentence pairs.