Graggle: A Graph-based Approach to Document Clustering

Graggle: A Graph-based Approach to Document Clustering
复制标题

DOI:
10.1109/bigdata55660.2022.10020645
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
I. J. King;H. H. Huang-H.
I. J. King;H. H. Huang-H.
中科院分区:
其他
文献类型:
--
作者:
I. J. King;H. H. Huang-H.

文献摘要

相似文献

文档推荐系统传统上依赖于高维向量表示,在具有不同词汇表的语料库中伸缩性差。现有的基于图的方法侧重于文档的元数据,不幸的是,忽略了论文的内容。在这项工作中,我们设计并实现了一个新的系统,我们称之为Graggle,它建立一个图来建模语料库。节点是纸张,边表示它们之间共享的重要单词。然后,我们利用现代图学习技术将该图转换为高效的降维工具。文档表示为低维矢量嵌入,由图形自编码器生成。我们的实验表明,该方法在标记数据上优于传统的基于文档向量和文本自动编码方法。此外,我们还将该技术应用于关于新型冠状病毒的未标记研究文件的存储库,以证明其作为现实世界工具的有效性。
Document recommendation systems have traditionally relied upon high-dimensional vector representations that scale poorly in corpora with diverse vocabularies. Existing graph-based approaches focus on the metadata of documents and, unfortunately, ignore the content of the papers. In this work, we have designed and implemented a new system we call Graggle, which builds a graph to model a corpus. Nodes are papers, and edges represent significant words shared between them. We then leverage modern graph learning techniques to turn this graph into a highly efficient tool for dimensionality reduction. Documents are represented as low-dimensional vector embeddings generated with a graph autoencoder. Our experiments show that this approach outperforms traditional document vector-based and text autoencoding approaches on labeled data. Additionally, we have applied this technique to a repository of unlabeled research documents about the novel coronavirus to demonstrate its effectiveness as a real-world tool.