Probabilistic Topic Models for Web Services Clustering and Discovery

Probabilistic Topic Models for Web Services Clustering and Discovery
复制标题

DOI:
10.1007/978-3-642-40651-5_3
复制
发表时间:
2013-09
期刊:
--
影响因子:
--
通讯作者:
Mustapha Aznag;M. Quafafou;E. M. Rochd;Zahi Jarir
Mustapha Aznag;M. Quafafou;E. M. Rochd;Zahi Jarir
中科院分区:
其他
文献类型:
--
作者:
Mustapha Aznag;M. Quafafou;E. M. Rochd;Zahi Jarir

文献摘要

被引文献

相似文献

在信息检索中,概率主题模型最初被开发并用于主题提取和文档建模。在本文中,我们探讨了几个概率主题模型:概率潜在语义分析(PLSA),潜在狄利克雷分配(LDA)和相关主题模型(CTM)提取潜在的因素,从Web服务描述。这些提取的潜在因素,然后使用分组的服务到集群。在我们的方法中,主题模型被用来作为有效的降维技术,这是能够捕捉词主题和主题服务解释的概率分布方面的语义关系。为了解决基于关键字的查询的局限性,我们表示为一个向量空间的Web服务描述,我们引入了一种新的方法来发现Web服务使用潜在的因素。在我们的实验中,我们比较了三个概率聚类算法(PLSA,LDA和CTM)的准确性与经典的聚类算法。我们还通过计算精度(P@n)和归一化折扣累积增益(NDCGn)来评估我们的服务发现方法。实验结果表明,基于CTM和LDA的两种方法都优于其他搜索方法。
In Information Retrieval the Probabilistic Topic Models were originally developed and utilized for topic extraction and document modeling. In this paper, we explore several probabilistic topic models: Probabilistic Latent Semantic Analysis (PLSA), Latent Dirichlet Allocation (LDA) and Correlated Topic Model (CTM) to extract latent factors from web service descriptions. These extracted latent factors are then used to group the services into clusters. In our approach, topic models are used as efficient dimension reduction techniques, which are able to capture semantic relationships between word-topic and topic-service interpreted in terms of probability distributions. To address the limitation of keywords-based queries, we represent web service description as a vector space and we introduce a new approach for discovering web services using latent factors. In our experiment, we compared the accuracy of the three probabilistic clustering algorithms (PLSA, LDA and CTM) with that of a classical clustering algorithm. We evaluated also our service discovery approach by calculating the precision (P@n) and normalized discounted cumulative gain (NDCGn). The results show that both approaches based on CTM and LDA perform better than other search methods.