Efficient Maximal Biclique Enumeration for Large Sparse Bipartite Graphs
Efficient Maximal Biclique Enumeration for Large Sparse Bipartite Graphs
复制标题
大型稀疏二部图的高效最大 Biclique 枚举
DOI:
10.14778/3529337.3529341
复制
发表时间:
2022-04
期刊:
影响因子:
--
通讯作者:
Jianxin Li
中科院分区:
文献类型:
--
作者:
Lu Chen;Chengfei Liu;Rui Zhou;Jiajie Xu;Jianxin Li
Maximal bicliques are effective to reveal meaningful information hidden in bipartite graphs. Maximal biclique enumeration (MBE) is challenging since the number of the maximal bicliques grows exponentially w.r.t. the number of vertices in a bipartite graph in the worst case. However, a large bipartite graph is usually very sparse, which is against the worst case and may lead to fast MBE algorithms. The uncharted opportunity is taking advantage of the sparsity to substantially improve the MBE efficiency for large sparse bipartite graphs. We observe that for a large sparse bipartite graph, a vertex
u
may converge to a few vertices in the same vertex set as
u
via its neighbours, which reveals that the enumeration scope for a vertex could be very small. Based on this observation, we propose novel concepts: unilateral coreness for individual vertices, unilateral order for each vertex set and unilateral convergence (ζ) for a large sparse bipartite graph, ζ could be a few thousand for a large sparse bipartite graph with hundreds of million edges. Using the unilateral order, every vertex with τ unilateral coreness only needs to check at most 2
τ
combinations so that all maximal bicliques can be enumerated and τ is bounded by ζ, which leads to a novel MBE algorithm running in
O
*
(2
ζ
). We then propose a batch-pivots technique to eliminate all enumerations resulting in non-maximal bicliques, which guarantees that every maximal biclique is reported in
O
(ζ
e
)-delay, where
e
is the number of edges. We devise novel data structures that allow storing subgraphs at omissible space for further speeding up MBE. Extensive experiments are conducted on synthetic and real large datasets to justify that our proposed algorithm is faster and more scalable than the existing algorithms.
登录
查看更多内容
DOI:
10.1109/icde48307.2020.00063
发表时间:
2020-01
期刊:
2020 IEEE 36th International Conference on Data Engineering (ICDE)
影响因子:
--
作者:
Kai Wang;Xuemin Lin;Lu Qin;Wenjie Zhang;Ying Zhang
通讯作者:
Kai Wang;Xuemin Lin;Lu Qin;Wenjie Zhang;Ying Zhang
DOI:
10.1007/978-3-030-03599-0
发表时间:
2018-12
期刊:
--
影响因子:
--
作者:
Lijun Chang;Lu Qin
通讯作者:
Lijun Chang;Lu Qin
DOI:
10.1109/icde.2019.00017
发表时间:
2019-04
期刊:
2019 IEEE 35th International Conference on Data Engineering (ICDE)
影响因子:
--
作者:
Lu Chen;Chengfei Liu;Kewen Liao;Jianxin Li;Rui Zhou
通讯作者:
Lu Chen;Chengfei Liu;Kewen Liao;Jianxin Li;Rui Zhou
DOI:
10.1007/978-3-030-39219-2_10
发表时间:
2020-02
期刊:
--
影响因子:
--
作者:
Arti Pandey;Gopika Sharma;N. Jain
通讯作者:
Arti Pandey;Gopika Sharma;N. Jain
影响因子:
14.9
作者:
Bu, DB;Zhao, Y;Chen, RS
通讯作者:
Chen, RS