https://doi.org/10.1137/1.9781611977172.17

https://doi.org/10.1137/1.9781611977172.17
复制标题

https://doi.org/10.1137/1.9781611977172.17

DOI:
10.1137/1.9781611977172.17
复制
发表时间:
2022
期刊:
Proceedings of the 2022 SIAM International Conference on Data Mining (SDM
影响因子:
--
通讯作者:
Ilya Amburg, Nate Veldt
Ilya Amburg, Nate Veldt
中科院分区:
--
文献类型:
--
作者:
Ilya Amburg, Nate Veldt

文献摘要

相似文献

在组建团队或小组时,一个人的目标通常是平衡一个主要关注领域的专业知识,同时鼓励每个团队的技能多样性。在本文中,我们将在给定任务中找到具有专业知识的不同群体的问题建模为具有异构边缘类型的超图上的聚类问题。在这里,超边缘类型编码过去经验类型的组,聚类的输出是个体组(节点)。与寻求公平或平衡集群(例如,在一些受保护的节点属性方面)的互补问题不同,我们的模型通过在边缘类型的节点参与方面在经验和多样性之间取得平衡,鼓励这些组中过去经验的多样性。我们证明了朴素目标不会导致多样性与经验的权衡,这激发了我们基于正则化基于边的超图聚类目标的改进模型。虽然优化我们的目标是np困难的,但我们设计了一个2-近似算法,该算法适用于更一般的问题类别,其中每个节点被允许对特定集群具有偏好,并说明了计算正则化强度界限的技术,该技术揭示了有意义的多样性/经验权衡机制。我们说明了我们的框架在几个现实生活数据集上的效用-最值得注意的是在线评论平台数据-为给定类型的产品策划评论集,这些评论集展示了评论者经验,或熟悉产品类型,经验,或评论者也评论相关产品类型的倾向之间的权衡。在允许节点偏好的设置中,我们展示了我们的框架发现对用户偏好敏感的评论集。
In forming teams or groups, one often aims to balance expertise in a main focus area while also encouraging diversity of skills in each team. In this paper we model the problem of finding diverse groups of individuals who have expertise in a given task as a clustering problem on hypergraphs with heterogeneous edge types. Here, the hyperedge types encode past experience types of groups, and the output of the clustering is groups of individuals (nodes). Unlike complementary problems that seek to find fair or balanced clusters (e.g., in terms of some protected node attributes), our model encourages diversity of pastexperiencewithin these groups by striking a balance between experience and diversity with respect to node participation in edge types. We show that naive objectives lead to no diversity-experience tradeoff, which motivates our refined model based on regularizing an edge-based hypergraph clustering objective. While optimizing our objective is NP-hard, we design a 2-approximation algorithm that works for a more general class of problems where each node is allowed to have a preference for a particular cluster, and illustrate a technique for computing regularization strength bounds that reveal meaningful diversity/experience tradeoff regimes. We illustrate the utility of our framework on several real-life datasets – most notably to online review platform data – to curate sets of reviews for a given type of product which exhibit a tradeoff between reviewer experience, or familiarity with a product type, and experience, or the reviewer's tendency to also review related product types. In the setting allowing for node preferences, we show that our framework discovers sets of reviews sensitive to user preference.