Last Words: Googleology is Bad Science

Last Words: Googleology is Bad Science
复制标题

遗言:谷歌学是糟糕的科学

DOI:
10.1162/coli.2007.33.1.147
复制
发表时间:
2021
期刊:
影响因子:
7.3
通讯作者:
A. Kilgarriff
A. Kilgarriff
中科院分区:
生物学1区
文献类型:
--
作者:
A. Kilgarriff

文献摘要

被引文献

相似文献

万维网是巨大的、免费的、可立即获得的,而且主要是语言。我们发现,在越来越多的方面,语言分析和生成受益于大数据,因此使用Web作为数据源变得很有吸引力。那么,问题是如何做到这一点。使用网络的低成本方式是通过商业搜索引擎。如果目标是寻找某些感兴趣的现象的频率或概率,我们可以使用搜索引擎命中页面中给出的命中数来进行估计。人们这样做已经有一段时间了。使用点击数的早期工作包括Grefenstette(1999),他确定了组合短语的可能翻译,Turney(2001),他发现了同义词;也许被引用最多的研究是Keller和Lapata(2003),他们通过对人类受试者的实验,确定了以这种方式收集的频率的有效性。最近的主要工作包括Nakov和Hearst(2005),他们建立了名词复合括号模型。
The World Wide Web is enormous, free, immediately available, and largely linguistic. As we discover, on ever more fronts, that language analysis and generation benefit from big data, so it becomes appealing to use the Web as a data source. The question, then, is how.The low-entry-cost way to use the Web is via a commercial search engine. If the goal is to find frequencies or probabilities for some phenomenon of interest, we can use the hit count given in the search engine’s hits page to make an estimate. People have been doing this for some time now. Early work using hit counts include Grefenstette (1999), who identified likely translations for compositional phrases, and Turney (2001), who found synonyms; perhaps the most cited study is Keller and Lapata (2003), who established the validity of frequencies gathered in this way using experiments with human subjects. Leading recent work includes Nakov and Hearst (2005), who build models of noun compound bracketing.