Leveraging word embeddings and medical entity extraction for biomedical dataset retrieval using unstructured texts

Leveraging word embeddings and medical entity extraction for biomedical dataset retrieval using unstructured texts
复制标题

DOI:
10.1093/database/bax091
复制
发表时间:
2017-12-20
影响因子:
5.8
通讯作者:
Liu, Hongfang
Liu, Hongfang
中科院分区:
生物学4区
文献类型:
--
作者:
Wang, Yanshan;Rastegar-Mojarad, Majid;Liu, Hongfang

文献摘要

被引文献

相似文献

最近在生物医学领域朝着开放数据的方向发展,产生了大量可公开获取的数据集。大数据到知识数据索引项目,生物医学和健康数据发现索引生态系统(bioCADDIE),将这些数据集收集在一站式门户网站中,旨在促进其重复使用,以加速科学进步。然而,随着存储和索引的生物医学数据集的数量增加,根据研究人员的查询检索相关数据集变得越来越具有挑战性。在这篇文章中,我们提出了一个信息检索(IR)系统来解决这个问题,并实现它的bioCADDIE数据集检索挑战。该系统利用每个数据集的非结构化文本,包括数据集的标题和描述,并利用最先进的IR模型,医学命名实体提取技术,基于深度学习的词嵌入的查询扩展和重新排序策略来提高检索性能。在实证实验中,我们使用bioCADDIE数据集检索挑战数据集将所提出的系统与11个基线系统进行了比较。实验结果表明,该系统在推理平均精度和推理归一化折扣累积增益方面优于其他系统,表明该系统是生物医学数据集检索的一种可行选择。
The recent movement towards open data in the biomedical domain has generated a large number of datasets that are publicly accessible. The Big Data to Knowledge data indexing project, biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE), has gathered these datasets in a one-stop portal aiming at facilitating their reuse for accelerating scientific advances. However, as the number of biomedical datasets stored and indexed increases, it becomes more and more challenging to retrieve the relevant datasets according to researchers' queries. In this article, we propose an information retrieval (IR) system to tackle this problem and implement it for the bioCADDIE Dataset Retrieval Challenge. The system leverages the unstructured texts of each dataset including the title and description for the dataset, and utilizes a state-of-the-art IR model, medical named entity extraction techniques, query expansion with deep learning-based word embeddings and a re-ranking strategy to enhance the retrieval performance. In empirical experiments, we compared the proposed system with 11 baseline systems using the bioCADDIE Dataset Retrieval Challenge datasets. The experimental results show that the proposed system outperforms other systems in terms of inference Average Precision and inference normalized Discounted Cumulative Gain, implying that the proposed system is a viable option for biomedical dataset retrieval.