No Longer Lost in Translation: Evidence that Google Translate Works for Comparative Bag-of-Words Text Applications

No Longer Lost in Translation: Evidence that Google Translate Works for Comparative Bag-of-Words Text Applications
复制标题

DOI:
10.1017/pan.2018.26
复制
发表时间:
2018-10-01
期刊:
影响因子:
5.4
通讯作者:
Schumacher, Gijs
Schumacher, Gijs
中科院分区:
法学1区
文献类型:
--
作者:
de Vries, Erik;Schoonvelde, Martijn;Schumacher, Gijs

文献摘要

被引文献

相似文献

自动文本分析允许研究人员分析大量文本。然而,比较研究人员面临着一个巨大的挑战:不同国家的人们说不同的语言。为了解决这个问题,一些分析师建议在开始分析之前使用谷歌翻译将所有文本转换为英语(Lucas et al. 2015)。但这样做时,我们会迷失在翻译中吗?本文评估了机器翻译对于词袋模型(例如主题模型)的有用性。我们使用 europarl 数据集并比较术语文档矩阵 (TDM) 以及黄金标准翻译文本和机器翻译文本的主题模型结果。我们在文档和语料库级别评估结果。我们首先发现两个文本语料库的 TDM 高度相似,不同语言之间存在微小差异。更重要的是,我们发现人工翻译和机器翻译文本生成的特征集有相当大的重叠。关于 LDA 主题模型,我们发现主题流行度和主题内容高度相似,不同语言之间也只有很小的差异。我们得出的结论是,在使用词袋文本模型时,谷歌翻译对于比较研究人员来说是一个有用的工具。
Automated text analysis allows researchers to analyze large quantities of text. Yet, comparative researchers are presented with a big challenge: across countries people speak different languages. To address this issue, some analysts have suggested using Google Translate to convert all texts into English before starting the analysis (Lucas et al. 2015). But in doing so, do we get lost in translation? This paper evaluates the usefulness of machine translation for bag-of-words models such as topic models. We use the europarl dataset and compare term-document matrices (TDMs) as well as topic model results from gold standard translated text and machine-translated text. We evaluate results at both the document and the corpus level. We first find TDMs for both text corpora to be highly similar, with minor differences across languages. What is more, we find considerable overlap in the set of features generated from human-translated and machine-translated texts. With regard to LDA topic models, we find topical prevalence and topical content to be highly similar with again only small differences across languages. We conclude that Google Translate is a useful tool for comparative researchers when using bag-of-words text models.