Enhanced Web Page Cleaning for Constructing Social Media Text Corpora

Enhanced Web Page Cleaning for Constructing Social Media Text Corpora
复制标题

构建社交媒体文本语料库的增强网页清理

DOI:
10.1007/978-3-662-46578-3_78
复制
发表时间:
2015
期刊:
International Journal of Insect Morphology & Embryology
影响因子:
--
通讯作者:
R. Mathar
R. Mathar
中科院分区:
--
文献类型:
--
作者:
M. Neunerdt;Eva Reimer;M. Reyer;R. Mathar

文献摘要

被引文献

相似文献

网页清理是网络语料库建设中的一项重要工作。这样做的目的是将主要内容与导航元素、模板和广告(通常称为样板)分开。在本文中,我们特别增强了应用于包含评论的页面的网页清理,并为此引入了新的训练语料库。除了通过评论分类器扩展现有的样板检测算法之外,我们还在扩展的特征集上训练和测试不同的分类器,以解决我们和现有基准语料库上的两类问题(内容与样板)。结果表明,该方法优于现有的方法,特别是在不同领域的评论页面。最后,我们指出,我们训练的分类器是领域独立的,只有小的调整,可转移到其他语言。
Web page cleaning is one of the most essential tasks in Web corpus construction. The intention is to separate the main content from navigational elements, templates, and advertisements, often referred to as boilerplate. In this paper, we particularly enhance Web page cleaning applied to pages containing comments and introduce a new training corpus for that purpose. Beside extending an existing boilerplate detection algorithm by means of a comment classifier, we train and test different classifiers on extended feature sets solving a two-class problem (content vs. boilerplate) on our and an existing benchmark corpus. Results show that the proposed approach outperforms existing methods, particularly on comment pages from different domains. Finally, we point out that our trained classifiers are domain independent and with small adjustments only transferable to other languages.