Multi-level Bootstrapping For Extracting Parallel Sentences From a Quasi-Comparable Corpus

Multi-level Bootstrapping For Extracting Parallel Sentences From a Quasi-Comparable Corpus
复制标题

DOI:
10.3115/1220355.1220506
复制
发表时间:
2004-08
期刊:
Proceedings of the 20th international conference on Computational Linguistics - COLING '04
影响因子:
--
通讯作者:
Pascale Fung;Percy Cheung
Pascale Fung;Percy Cheung
中科院分区:
其他
文献类型:
--
作者:
Pascale Fung;Percy Cheung

文献摘要

被引文献

相似文献

我们提出了一种完全无监督的方法,用于从具有不同大小的准可比双语文本中挖掘并列句子,这些文本包括话题内和话题外的文档。我们讨论和分析了具有不同程度相似性的不同双语语料库。我们认为,虽然更好的文档匹配会带来更好的平行句子提取,但更好的句子匹配也会带来更好的文档匹配。在此基础上,我们使用多级自举来迭代地改进文档、句子和双语词对之间的对齐。我们的方法是第一种不依赖于任何有监督的训练数据(例如句子对齐的语料库)或时间信息(例如新闻文章的发布日期)的方法。实验结果表明,该方法比无多层自举的方法提高了23%。
We propose a completely unsupervised method for mining parallel sentences from quasi-comparable bilingual texts which have very different sizes, and which include both in-topic and off-topic documents. We discuss and analyze different bilingual corpora with various levels of comparability. We propose that while better document matching leads to better parallel sentence extraction, better sentence matching also leads to better document matching. Based on this, we use multi-level bootstrapping to improve the alignments between documents, sentences, and bilingual word pairs, iteratively. Our method is the first method that does not rely on any supervised training data, such as a sentence-aligned corpus, or temporal information, such as the publishing date of a news article. It is validated by experimental results that show a 23% improvement over a method without multilevel bootstrapping.