BlogVox: Separating Blog Wheat from Blog Chaff

BlogVox: Separating Blog Wheat from Blog Chaff
复制标题

DOI:
--
复制
发表时间:
2007-01
期刊:
--
影响因子:
--
通讯作者:
Akshay Java;Pranam Kolari;Timothy W. Finin;J. Mayfield;A. Joshi;Justin Martineau
Akshay Java;Pranam Kolari;Timothy W. Finin;J. Mayfield;A. Joshi;Justin Martineau
中科院分区:
其他
文献类型:
--
作者:
Akshay Java;Pranam Kolari;Timothy W. Finin;J. Mayfield;A. Joshi;Justin Martineau

文献摘要

被引文献

相似文献

博客文章通常写得不正式,结构糟糕,充斥着拼写和语法错误,并以非传统内容为特色。这些特征使它们难以用标准语言分析工具进行处理。对博客进行语言分析受到两个额外问题的困扰:(i)垃圾博客和垃圾评论的存在;(ii)无关的非内容,包括博客卷、链接卷、广告和侧边栏。我们描述了作为BlogVox系统(我们为2006 TREC博客跟踪开发的博客分析引擎)的一部分而开发的用于消除嘈杂博客数据的技术。本文的研究结果强调了从博客收集中删除虚假内容的重要性。
Blog posts are often informally written, poorly structured, rife with spelling and grammatical errors, and feature non-traditional content. These characteristics make them difficult to process with standard language analysis tools. Performing linguistic analysis on blogs is plagued by two additional problems: (i) the presence of spam blogs and spam comments and (ii) extraneous non-content including blog-rolls, link-rolls, advertisements and sidebars. We describe techniques designed to eliminate noisy blog data developed as part of the BlogVox system - a blog analytics engine we developed for the 2006 TREC Blog Track. The findings in this paper underscore the importance of removing spurious content from blog collections.