Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities

Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities
复制标题

DOI:
10.1109/bigdata.2018.8622309
复制
发表时间:
2018-12
期刊:
--
影响因子:
--
通讯作者:
Yuanxun Zhang;P. Calyam;T. Joshi;S. Nair;Dong Xu
Yuanxun Zhang;P. Calyam;T. Joshi;S. Nair;Dong Xu
中科院分区:
其他
文献类型:
--
作者:
Yuanxun Zhang;P. Calyam;T. Joshi;S. Nair;Dong Xu

文献摘要

被引文献

相似文献

作为大数据分析基础的机器学习技术有可能使数据密集型社区受益,例如,生物信息学和神经科学领域的科学。今天,这些领域的创新进步越来越多地建立在多学科知识发现和跨领域合作的基础上。因此,在研究新方法、开发新工具或集成数据集时,缩短知识发现时间是一个挑战。领域科学家的挑战尤其在于通过从由广泛的主题集合组成的不同文本语料库中查询大量信息来获得指导的行动。在本文中,我们提出了一种新的“特定领域的主题模型”(DSTM),可以驱动会话代理用户发现潜在的知识模式之间的关系的研究主题,工具和数据集的范例科学领域。DSTM的目标是执行数据挖掘,通过聊天机器人为领域科学家选择与解决当前计算和数据密集型研究问题相关的相关工具或数据集提供有意义的指导。我们的DSTM是一个贝叶斯分层模型,扩展了潜在的狄利克雷分配(LDA)模型,并使用马尔可夫链蒙特卡罗算法来推断在一个特定的域中的潜在模式在无监督的方式。我们将DSTM应用于生物信息学和神经科学领域的大量数据集,其中包括来自知名期刊档案的数百篇论文,数百种工具和数据集。通过一个困惑度的评价实验,我们表明,我们的模型具有更好的泛化性能在一个领域内发现高度具体的潜在主题。
Machine learning techniques underlying Big Data analytics have the potential to benefit data intensive communities in e.g., bioinformatics and neuroscience domain sciences. Today’s innovative advances in these domain communities are increasingly built upon multi-disciplinary knowledge discovery and cross-domain collaborations. Consequently, shortened time to knowledge discovery is a challenge when investigating new methods, developing new tools, or integrating datasets. The challenge for a domain scientist particularly lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics. In this paper, we propose a novel "domain-specific topic model" (DSTM) that can drive conversational agents for users to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplar scientific domains. The goal of DSTM is to perform data mining to obtain meaningful guidance via a chatbot for domain scientists to choose the relevant tools or datasets pertinent to solving a computational and data intensive research problem at hand. Our DSTM is a Bayesian hierarchical model that extends the Latent Dirichlet Allocation (LDA) model and uses a Markov chain Monte Carlo algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include hundreds of papers from reputed journal archives, hundreds of tools and datasets. Through evaluation experiments with a perplexity metric, we show that our model has better generalization performance within a domain for discovering highly specific latent topics.