Accurate and efficient general-purpose boilerplate detection for crawled web corpora

Accurate and efficient general-purpose boilerplate detection for crawled web corpora
复制标题

用于爬行网络语料库的准确高效的通用样板检测

DOI:
10.1007/s10579-016-9359-2
复制
发表时间:
2017
影响因子:
2.7
通讯作者:
Roland
Roland
中科院分区:
计算机科学4区
文献类型:
--
作者:
Schäfer;Roland

文献摘要

参考文献

被引文献

相似文献

模板的删除是web语料库构建和web索引的基本任务之一。样板(冗余和自动插入的材料,如菜单、版权声明、导航元素等)通常被认为在语言上不适合包含在web语料库中。此外,搜索引擎不应该索引这样的材料,因为如果这些内容出现在网页的样板区域中,可能会导致虚假的搜索结果。在本文中,我提出并评估了一种监督机器学习方法,该方法使用多层感知器(mlp)对基于拉丁字母的语言进行通用样板检测。它既高效又准确(根据输入语言的不同,分类在95%和正确率之间)。我展示了特定于语言的分类器极大地提高了样板检测器的准确性。用于分类的单个特征是根据它们对分类的贡献来评估的。此外,我还展示了MLP的准确性与许多其他分类器的准确性相当。我的方法已经在开源的exexxweb页面清理软件中实现,并且使用它构建的大型语料库可以从COW计划中获得,包括从CommonCrawl数据集创建的CommonCOW语料库。
Removal of boilerplate is one of the essential tasks in web corpus construction and web indexing. Boilerplate (redundant and automatically inserted material like menus, copyright notices, navigational elements, etc.) is usually considered to be linguistically unattractive for inclusion in a web corpus. Also, search engines should not index such material because it can lead to spurious results for search terms if these terms appear in boilerplate regions of the web page. In this paper, I present and evaluate a supervised machine-learning approach to general-purpose boilerplate detection for languages based on Latin alphabets using Multi-Layer Perceptrons (MLPs). It is both very efficient and very accurate (between 95 % andcorrect classifications, depending on the input language). I show that language-specific classifiers greatly improve the accuracy of boilerplate detectors. The single features used for the classification are evaluated with regard to the merit they contribute to the classification. Furthermore, I show that the accuracy of the MLP is on a par with that of a wide range of other classifiers. My approach has been implemented in the open-sourcetexrexweb page cleaning software, and large corpora constructed using it are available from the COW initiative, including the CommonCOW corpora created from CommonCrawl datasets.
DOI: --
发表时间: 2014
期刊: Journal of management science
影响因子: --
作者:
อนิรุธ สืบสิงห์
通讯作者: อนิรุธ สืบสิงห์
使用新的高效工具链从网络构建大型语料库
DOI: --
发表时间: 2012
期刊: International Conference on Language Resources and Evaluation
影响因子: --
作者:
R. Schäfer;Felix Bildhauer
通讯作者: Felix Bildhauer
DOI: 10.1016/j.chemolab.2005.09.003
发表时间: 2006-03-15
影响因子: 3.9
作者:
Üstün, B;Melssen, WJ;Buydens, LMC
通讯作者: Buydens, LMC
构建社交媒体文本语料库的增强网页清理
DOI: 10.1007/978-3-662-46578-3_78
发表时间: 2015
期刊: International Journal of Insect Morphology & Embryology
影响因子: --
作者:
M. Neunerdt;Eva Reimer;M. Reyer;R. Mathar
通讯作者: R. Mathar
样板检测和重新编码
DOI: 10.1007/978-3-319-06028-6_42
发表时间: 2014
期刊: International Journal of Insect Morphology & Embryology
影响因子: --
作者:
Matthias Gallé;J. Renders
通讯作者: J. Renders