Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities
Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities
复制标题
DOI:
10.1109/bigdata.2018.8622309
复制
发表时间:
2018-12
期刊:
影响因子:
--
通讯作者:
Yuanxun Zhang;P. Calyam;T. Joshi;S. Nair;Dong Xu
中科院分区:
文献类型:
--
作者:
Yuanxun Zhang;P. Calyam;T. Joshi;S. Nair;Dong Xu
Machine learning techniques underlying Big Data analytics have the potential to benefit data intensive communities in e.g., bioinformatics and neuroscience domain sciences. Today’s innovative advances in these domain communities are increasingly built upon multi-disciplinary knowledge discovery and cross-domain collaborations. Consequently, shortened time to knowledge discovery is a challenge when investigating new methods, developing new tools, or integrating datasets. The challenge for a domain scientist particularly lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics. In this paper, we propose a novel "domain-specific topic model" (DSTM) that can drive conversational agents for users to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplar scientific domains. The goal of DSTM is to perform data mining to obtain meaningful guidance via a chatbot for domain scientists to choose the relevant tools or datasets pertinent to solving a computational and data intensive research problem at hand. Our DSTM is a Bayesian hierarchical model that extends the Latent Dirichlet Allocation (LDA) model and uses a Markov chain Monte Carlo algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include hundreds of papers from reputed journal archives, hundreds of tools and datasets. Through evaluation experiments with a perplexity metric, we show that our model has better generalization performance within a domain for discovering highly specific latent topics.