Last Words: Googleology is Bad Science
Last Words: Googleology is Bad Science
复制标题
遗言:谷歌学是糟糕的科学
DOI:
10.1162/coli.2007.33.1.147
复制
发表时间:
2021
影响因子:
7.3
通讯作者:
A. Kilgarriff
中科院分区:
文献类型:
--
作者:
A. Kilgarriff
The World Wide Web is enormous, free, immediately available, and largely linguistic. As we discover, on ever more fronts, that language analysis and generation benefit from big data, so it becomes appealing to use the Web as a data source. The question, then, is how.The low-entry-cost way to use the Web is via a commercial search engine. If the goal is to find frequencies or probabilities for some phenomenon of interest, we can use the hit count given in the search engine’s hits page to make an estimate. People have been doing this for some time now. Early work using hit counts include Grefenstette (1999), who identified likely translations for compositional phrases, and Turney (2001), who found synonyms; perhaps the most cited study is Keller and Lapata (2003), who established the validity of frequencies gathered in this way using experiments with human subjects. Leading recent work includes Nakov and Hearst (2005), who build models of noun compound bracketing.