Gene Ontology density estimation and discourse analysis for automatic GeneRiF extraction.

Gene Ontology density estimation and discourse analysis for automatic GeneRiF extraction.
复制标题

DOI:
10.1186/1471-2105-9-s3-s9
复制
发表时间:
2008-04-11
期刊:
影响因子:
3
通讯作者:
Ruch P
Ruch P
中科院分区:
生物学4区
文献类型:
--
作者:
Gobeill J;Tbahriti I;Ehrler F;Mottaz A;Veuthey AL;Ruch P

文献摘要

被引文献

相似文献

本文描述并评估了一个句子选择引擎,该引擎根据 MEDLINE 记录提取 ENTREZ-Gene 中定义的 GeneRiF(函数中的基因引用)。此任务的输入包括基因和指向 MEDLINE 参考的指针。在建议的方法中,我们合并了两个独立的句子提取策略。第一个提出的策略(LASt)使用论证特征,其灵感来自于话语分析模型。第二种提取方案(GOEx)使用自动文本分类器来估计每个句子中基因本体类别的密度;从而提供所有可能的候选 GeneRiF 的完整排名。提出了两种方法的组合,其目的还在于通过过滤掉不包含内容的修辞短语来减小所选片段的大小。基于用于 GeneRiF 鉴定的 TREC-2003 基因组集合,LASt 提取策略已经具有竞争力(52.78%)。当采用组合方法时,提取任务明显显示出改进,实现了超过 57% (+10%) 的 Dice 分数。使用基因本体内容的论证性表示水平和概念密度估计似乎与蛋白质组学中的功能注释互补。
This paper describes and evaluates a sentence selection engine that extracts a GeneRiF (Gene Reference into Functions) as defined in ENTREZ-Gene based on a MEDLINE record. Inputs for this task include both a gene and a pointer to a MEDLINE reference. In the suggested approach we merge two independent sentence extraction strategies. The first proposed strategy (LASt) uses argumentative features, inspired by discourse-analysis models. The second extraction scheme (GOEx) uses an automatic text categorizer to estimate the density of Gene Ontology categories in every sentence; thus providing a full ranking of all possible candidate GeneRiFs. A combination of the two approaches is proposed, which also aims at reducing the size of the selected segment by filtering out non-content bearing rhetorical phrases. Based on the TREC-2003 Genomics collection for GeneRiF identification, the LASt extraction strategy is already competitive (52.78%). When used in a combined approach, the extraction task clearly shows improvement, achieving a Dice score of over 57% (+10%). Argumentative representation levels and conceptual density estimation using Gene Ontology contents appear complementary for functional annotation in proteomics.