Domain-Specific Topic Model for Knowledge Discovery in Computational and Data-Intensive Scientific Communities

Domain-Specific Topic Model for Knowledge Discovery in Computational and Data-Intensive Scientific Communities
复制标题

DOI:
10.1109/tkde.2021.3093350
复制
发表时间:
2023-02
影响因子:
8.9
通讯作者:
Yuanxun Zhang;P. Calyam;T. Joshi;Satish Nair;Dong Xu
Yuanxun Zhang;P. Calyam;T. Joshi;Satish Nair;Dong Xu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yuanxun Zhang;P. Calyam;T. Joshi;Satish Nair;Dong Xu

文献摘要

相似文献

缩短知识发现和调整现有领域知识的时间是计算和数据密集型社区的挑战,生物信息学和神经科学。领域科学家面临的挑战在于通过从包含广泛主题的不同文本语料库中查询大量信息来获得指导的行动:研究新方法,开发新工具或集成数据集。在本文中,我们提出了一种新的“特定领域的主题模型”(DSTM),发现潜在的知识模式之间的关系的研究主题,工具和数据集的示范科学领域。我们的DSTM是一个生成模型,它扩展了潜在狄利克雷分配(LDA)模型,并使用马尔可夫链蒙特卡罗(MCMC)算法来推断在一个特定的域中的潜在模式在无监督的方式。我们将DSTM应用于生物信息学和神经科学领域的大量数据集,其中包括过去十年中超过25,000篇论文,其中包括相关研究中常用的数百种工具和数据集。基于泛化和信息检索指标的评估实验表明,我们的模型具有更好的性能比国家的最先进的基线模型发现高度特定的潜在主题在一个领域。最后,我们展示了从我们的DSTM中受益的应用程序,以发现域内,跨域和趋势知识模式。
Shortened time to knowledge discovery and adapting prior domain knowledge is a challenge for computational and data-intensive communities such as e.g., bioinformatics and neuroscience. The challenge for a domain scientist lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics when: investigating new methods, developing new tools, or integrating datasets. In this paper, we propose a novel “domain-specific topic model” (DSTM) to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplary scientific domains. Our DSTM is a generative model that extends the Latent Dirichlet Allocation (LDA) model and uses the Markov chain Monte Carlo (MCMC) algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include more than 25,000 of papers over the last ten years, featuring hundreds of tools and datasets that are commonly used in relevant studies. Evaluation experiments based on generalization and information retrieval metrics show that our model has better performance than the state-of-the-art baseline models for discovering highly-specific latent topics within a domain. Lastly, we demonstrate applications that benefit from our DSTM to discover intra-domain, cross-domain and trend knowledge patterns.