A content-based literature recommendation system for datasets to improve data reusability - A case study on Gene Expression Omnibus (GEO) datasets

A content-based literature recommendation system for datasets to improve data reusability - A case study on Gene Expression Omnibus (GEO) datasets
复制标题

DOI:
10.1016/j.jbi.2020.103399
复制
发表时间:
2020-04-01
影响因子:
4.5
通讯作者:
Wu, Hulin
Wu, Hulin
中科院分区:
医学3区
文献类型:
--
作者:
Patra, Braja Gopal;Maroufy, Vahed;Wu, Hulin

文献摘要

被引文献

相似文献

目的:数据对生物医学研究的中心地位很难被低估,生物医学文献在传播针对此类数据提出的科学问题的实证结果方面的重要性也是如此。但文献和相关数据集之间的联系往往很弱,这阻碍了科学家在现有数据集和现有发现之间轻松移动以得出新的科学假设的能力。这项工作旨在推荐数据集的相关文献文章,最终目标是提高研究人员的生产力。我们的数据集文献推荐方法是休斯顿德克萨斯大学健康科学中心开发的数据集可重用性平台的一部分,用于与基因表达相关的数据集。该平台整合了来自 Gene Expression Omnibus (GEO) 的数据集。过去五年(即 2014 年至 2018 年)平均每天有 34 个数据集添加到 GEO,这表明需要自动方法将这些数据集与相关文献连接起来。给定数据集的相关文献可以描述该数据集,提供基于该数据集的科学发现,甚至描述数据集用户感兴趣的数据集主题的先前和相关工作。材料和方法:我们采用信息检索范式进行文献推荐。在我们的实验中,分布式语义特征是根据 MEDLINE 文章的标题和摘要创建的。然后,为 GEO 中的数据集识别相关文章。我们评估了多种分布式方法,例如 TF-IDF、BM25、潜在语义分析、潜在狄利克雷分配、word2vec 和 doc2vec。使用数据集的向量表示和每篇论文的向量表示之间的余弦相似度为每个数据集推荐最相似的论文。我们还提出了几种新颖的嵌入重排序和归一化方法来改进推荐。结果:基于对 36 个数据集的手动评估,使用 BM25,表现最好的文献推荐技术实现了 10 的严格精度为 0.8333,部分精度为 10 的 0.9000。通过强调数据集和文章标题的相似性,对更大的自动收集基准的评估显示出小但一致的收益。结论:这项工作是通过推荐数据集的相关文献来开发文献推荐工具的第一步。这有望带来更好的数据重用体验。
Objective: The centrality of data to biomedical research is difficult to understate, and the same is true for the importance of the biomedical literature in disseminating empirical findings to scientific questions made on such data. But the connections between the literature and related datasets are often weak, hampering the ability of scientists to easily move between existing datasets and existing findings to derive new scientific hypotheses. This work aims to recommend relevant literature articles for datasets with the ultimate goal of increasing the productivity of researchers. Our approach to literature recommendation for datasets is a part of the dataset reusability platform developed at the University Texas Health Science Center at Houston for datasets related to gene expression. This platform incorporates datasets from Gene Expression Omnibus (GEO). An average of 34 datasets were added to GEO daily in the last five years (i.e. 2014 to 2018), demonstrating the need for automatic methods to connect these datasets with relevant literature. The relevant literature for a given dataset may describe that dataset, provide a scientific finding based on that dataset, or even describe prior and related work to the dataset's topic that is of interest to users of the dataset.Materials and methods: We adopt an information retrieval paradigm for literature recommendation. In our experiments, distributional semantic features are created from the title and abstract of MEDLINE articles. Then, related articles are identified for datasets in GEO. We evaluate multiple distributional methods such as TF-IDF, BM25, Latent Semantic Analysis, Latent Dirichlet Allocation, word2vec, and doc2vec. Top similar papers are recommended for each dataset using cosine similarity between the dataset's vector representation and every paper's vector representation. We also propose several novel re-ranking and normalization methods over embeddings to improve the recommendations.Results: The top-performing literature recommendation technique achieved a strict precision at 10 of 0.8333 and a partial precision at 10 of 0.9000 using BM25 based on a manual evaluation of 36 datasets. Evaluation on a larger, automatically-collected benchmark shows small but consistent gains by emphasizing the similarity of dataset and article titles.Conclusion: This work is the first step toward developing a literature recommendation tool by recommending relevant literature for datasets. This will hopefully lead to better data reuse experience.