Enhanced Web Page Cleaning for Constructing Social Media Text Corpora
Enhanced Web Page Cleaning for Constructing Social Media Text Corpora
复制标题
构建社交媒体文本语料库的增强网页清理
DOI:
10.1007/978-3-662-46578-3_78
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
R. Mathar
中科院分区:
文献类型:
--
作者:
M. Neunerdt;Eva Reimer;M. Reyer;R. Mathar
Web page cleaning is one of the most essential tasks in Web corpus construction. The intention is to separate the main content from navigational elements, templates, and advertisements, often referred to as boilerplate. In this paper, we particularly enhance Web page cleaning applied to pages containing comments and introduce a new training corpus for that purpose. Beside extending an existing boilerplate detection algorithm by means of a comment classifier, we train and test different classifiers on extended feature sets solving a two-class problem (content vs. boilerplate) on our and an existing benchmark corpus. Results show that the proposed approach outperforms existing methods, particularly on comment pages from different domains. Finally, we point out that our trained classifiers are domain independent and with small adjustments only transferable to other languages.