Automatic identification of genre in Web pages

Automatic identification of genre in Web pages
复制标题

自动识别网页流派

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Marina Santini
Marina Santini
中科院分区:
--
文献类型:
--
作者:
Marina Santini

文献摘要

被引文献

相似文献

类型是一个复杂但直观理解的概念。主页、常见问题解答、博客等是当前网络上蓬勃发展的类型的示例。自动识别网络类型将帮助我们找到与我们的信息需求更相关的文档。本书描述的研究目的是开发自动流派分类算法。然而,存在一些影响这些算法建模的挑战。首先,网络上的流派在网页中实例化,网页可以被视为一种新型文档,比纸质文档更加不可预测和个性化。其次,网络是不稳定和流动的,正在经历快节奏的演变,因此流派识别受到小说流派的形成、流派混合、个性化、流派内和流派间变异等现象的影响。最后,迄今为止使用的自动可提取的流派揭示特征不足以定义现有的和新颖的网络流派。作者认为,网页流派的自动识别需要更灵活的流派分类方案。本书的主体描述了支持这一主张的实验。
Genre is a complex but intuitively understood concept. Home pages, FAQs, blogs, etc. are examples of genres currently thriving on the web. Automatically identifying web genres would help us find documents that are more relevant to our information needs. The aim of the research described in this book is to develop automatic genre classification algorithms. There are several challenges, however, that affect the modelling of these algorithms. First, genres on the web are instantiated in web pages, which can be considered documents of a new type, much more unpredictable and individualised than documents on paper. Second, the web is unstable and fluid, undergoing a fast-paced evolution, so genre identification is influenced by phenomena such as the formation of novel genres, genre hybridism, individualisation, intra-genre and inter-genre variation. Finally, the automatically extractable genre-revealing features used up to now are not adequate to define existing and novel web genres. The author argues that automatic identification of genre in web pages needs more flexible genre classification schemes. The main body of the book describes experiments that support this claim.