Statistical modeling of biomedical corpora: mining the Caenorhabditis Genetic Center Bibliography for genes related to life span

Statistical modeling of biomedical corpora: mining the Caenorhabditis Genetic Center Bibliography for genes related to life span
复制标题

DOI:
10.1186/1471-2105-7-250
复制
发表时间:
2006-05-08
期刊:
影响因子:
3
通讯作者:
Mian, I. S.
Mian, I. S.
中科院分区:
生物学4区
文献类型:
--
作者:
Blei, D. M.;Franks, K.;Mian, I. S.

文献摘要

被引文献

相似文献

背景:生物医学语料库的统计建模可以对生物现象产生综合的、从粗到细的观点,补充从分子序列分析和分析数据中获得的发现。在这里,通过使用统计信息检索技术检查Caenorhabditis Genetic Center (CGC) Bibliography中的5,225个自由文本项目,证明了这种建模的潜力。CGC生物医学文本语料库中的条目使用潜在狄利克雷分配(Latent Dirichlet Allocation, LDA)模型建模。LDA是一种层次贝叶斯模型,它将文档表示为潜在主题的随机混合;每个主题的特征是单词的分布。结果:从CGC项目估计的LDA模型比使用相同数据训练的两个标准模型(单图和混合单图)具有更好的预测性能。为了说明LDA模型在生物医学语料库中的实际应用,我们使用一个经过训练的CGC LDA模型对已知与寿命改变相关的线虫基因进行了回顾性研究。语料库、文档级和词级LDA参数与基因本体的术语相结合,以增强CGC LDA模型的解释价值,并为年龄相关基因提供额外的候选基因。提出了一种基于主题单纯形的后验分布的新型两两文档相似度度量方法,并将其用于在CGC数据库中搜索讨论寿命修饰clk-2基因的“查询”文档的“同源物”。对这些文件同源物的检查能够并促进了关于clk-2的功能和作用的假设的产生。结论:与其他遗传、基因组和其他类型生物数据的图形模型一样,LDA提供了一种提取意想不到的见解并生成可用于后续实验验证的预测的方法。
Background: The statistical modeling of biomedical corpora could yield integrated, coarse-to-fine views of biological phenomena that complement discoveries made from analysis of molecular sequence and profiling data. Here, the potential of such modeling is demonstrated by examining the 5,225 free-text items in the Caenorhabditis Genetic Center (CGC) Bibliography using techniques from statistical information retrieval. Items in the CGC biomedical text corpus were modeled using the Latent Dirichlet Allocation (LDA) model. LDA is a hierarchical Bayesian model which represents a document as a random mixture over latent topics; each topic is characterized by a distribution over words.Results: An LDA model estimated from CGC items had better predictive performance than two standard models (unigram and mixture of unigrams) trained using the same data. To illustrate the practical utility of LDA models of biomedical corpora, a trained CGC LDA model was used for a retrospective study of nematode genes known to be associated with life span modification. Corpus, document-, and word-level LDA parameters were combined with terms from the Gene Ontology to enhance the explanatory value of the CGC LDA model, and to suggest additional candidates for age-related genes. A novel, pairwise document similarity measure based on the posterior distribution on the topic simplex was formulated and used to search the CGC database for "homologs" of a "query" document discussing the life span-modifying clk-2 gene. Inspection of these document homologs enabled and facilitated the production of hypotheses about the function and role of clk-2.Conclusion: Like other graphical models for genetic, genomic and other types of biological data, LDA provides a method for extracting unanticipated insights and generating predictions amenable to subsequent experimental validation.