Model-based clustering and visualization of navigation patterns on a web site

Model-based clustering and visualization of navigation patterns on a web site
复制标题

DOI:
10.1023/a:1024992613384
复制
发表时间:
2003-10-01
影响因子:
4.8
通讯作者:
White, S
White, S
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cadez, I;Heckerman, D;White, S

文献摘要

被引文献

相似文献

我们提出了一种新的方法来探索和分析网站上的导航模式。可以分析的模式由用户遍历的URL类别序列组成。在我们的方法中,我们首先将站点用户划分为集群,这样具有相似导航路径的用户将被放置到同一个集群中。然后,对于每个集群,我们为该集群中的用户显示这些路径。我们采用的聚类方法是基于模型的(而不是基于距离的),并根据用户请求网页的顺序对用户进行分区。特别是,我们通过学习使用期望最大化算法的一阶马尔可夫模型的混合物来聚类用户。我们的算法的运行时间与集群的数量和数据的大小成线性关系;我们的实现可以轻松地处理内存中的数十万个用户会话。在本文中,我们描述了我们的方法和可视化工具的基础上,它被称为WebCANVAS的细节。我们说明了使用我们的方法对msnbc.com的用户流量数据。
We present a new methodology for exploring and analyzing navigation patterns on a web site. The patterns that can be analyzed consist of sequences of URL categories traversed by users. In our approach, we first partition site users into clusters such that users with similar navigation paths through the site are placed into the same cluster. Then, for each cluster, we display these paths for users within that cluster. The clustering approach we employ is model-based ( as opposed to distance-based) and partitions users according to the order in which they request web pages. In particular, we cluster users by learning a mixture of first-order Markov models using the Expectation-Maximization algorithm. The runtime of our algorithm scales linearly with the number of clusters and with the size of the data; and our implementation easily handles hundreds of thousands of user sessions in memory. In the paper, we describe the details of our method and a visualization tool based on it called WebCANVAS. We illustrate the use of our approach on user-traffic data from msnbc.com.