Structured Correspondence Topic Models for Mining Captioned Figures in Biological Literature.

Structured Correspondence Topic Models for Mining Captioned Figures in Biological Literature.
复制标题

DOI:
10.1145/1557019.1557031
复制
发表时间:
2009
期刊:
KDD : proceedings. International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Murphy RF
Murphy RF
中科院分区:
其他
文献类型:
--
作者:
Ahmed A;Xing EP;Cohen WW;Murphy RF

文献摘要

被引文献

相似文献

科学期刊、论文集和书籍中学术文章的一个主要信息来源(通常是最关键和信息量最大的部分)是直接提供关键实验结果和其他科学内容的图像和其他图形说明的数字。在生物学文章中,一个典型的图通常包括多个面板,并伴有范围或全局标题文本。此外,标题中的文本包含重要的语义实体,如蛋白质名称、基因本体、组织标签等,与图中的图像相关。近年来,随着生物学文献的大量涌现以及各种生物成像技术的日益普及,从生物学文献中自动检索和摘要生物信息已成为生命科学领域计算知识提取和管理的一个重要挑战。我们提出了一个新的结构化概率主题模型建立在一个现实的数字生成方案,以模拟结构化注释的生物数字,我们推导出一个有效的推理算法的基础上崩溃吉布斯抽样的信息检索和可视化。由此产生的程序构成了我们的SLIF系统中的关键IR引擎之一,该系统最近进入了生命科学知识增强Elsevier Grand Challenge的最后一轮(70个竞争系统中的4个)。在这里,我们提出了一些数据挖掘任务的各种评价,以说明我们的方法。
A major source of information (often the most crucial and informative part) in scholarly articles from scientific journals, proceedings and books are the figures that directly provide images and other graphical illustrations of key experimental results and other scientific contents. In biological articles, a typical figure often comprises multiple panels, accompanied by either scoped or global captioned text. Moreover, the text in the caption contains important semantic entities such as protein names, gene ontology, tissues labels, etc., relevant to the images in the figure. Due to the avalanche of biological literature in recent years, and increasing popularity of various bio-imaging techniques, automatic retrieval and summarization of biological information from literature figures has emerged as a major unsolved challenge in computational knowledge extraction and management in the life science. We present a new structured probabilistic topic model built on a realistic figure generation scheme to model the structurally annotated biological figures, and we derive an efficient inference algorithm based on collapsed Gibbs sampling for information retrieval and visualization. The resulting program constitutes one of the key IR engines in our SLIF system that has recently entered the final round (4 out 70 competing systems) of the Elsevier Grand Challenge on Knowledge Enhancement in the Life Science. Here we present various evaluations on a number of data mining tasks to illustrate our method.