Boilerplate Detection and Recoding

Boilerplate Detection and Recoding
复制标题

样板检测和重新编码

DOI:
10.1007/978-3-319-06028-6_42
复制
发表时间:
2014
期刊:
International Journal of Insect Morphology & Embryology
影响因子:
--
通讯作者:
J. Renders
J. Renders
中科院分区:
--
文献类型:
--
作者:
Matthias Gallé;J. Renders

文献摘要

被引文献

相似文献

许多信息访问应用程序必须处理自然语言文本,这些文本包含很大比例的重复和基本不变的模式(称为样板),例如自动模板,标题,签名和表格格式。这些特定领域的标准公式通常比传统的搭配或标准名词短语长得多,通常覆盖一个或多个句子。这样的主题显然具有非合成的意义,理想的文档表示应该反映这种现象。 我们在这里提出了一种方法,自动检测和无监督的方式,这样的图案,并丰富了这些图案的具体功能,包括文件表示。我们的实验表明,这种文件重新编码策略导致不同的集合,以改善分类。
Many information access applications have to tackle natural language texts that contain a large proportion of repeated and mostly invariable patterns --- called boilerplates ---, such as automatic templates, headers, signatures and table formats. These domain-specific standard formulations are usually much longer than traditional collocations or standard noun phrases and typically cover one or more sentences. Such motifs clearly have a non-compositional meaning and an ideal document representation should reflect this phenomenon. We propose here a method that detects automatically and in an unsupervised way such motifs; and enriches the document representation by including specific features for these motifs. We experimentally show that this document recoding strategy leads to improved classification on different collections.