Retrieval with gene queries.

Retrieval with gene queries.
复制标题

DOI:
10.1186/1471-2105-7-220
复制
发表时间:
2006-04-21
期刊:
影响因子:
3
通讯作者:
Srinivasan P
Srinivasan P
中科院分区:
生物学4区
文献类型:
--
作者:
Sehgal AK;Srinivasan P

文献摘要

参考文献

被引文献

相似文献

从MEDLINE检索用于基因查询的文档的准确性对于生物信息学中的许多应用是至关重要的。我们探索了五种基于信息检索的方法来对PubMed基因查询检索到的人类基因组文档进行排序。其目的是将相关文档在检索的列表中排在更高的位置。我们解决了由于基因命名中的歧义而面临的特殊挑战:指多个基因的基因术语,也是英语单词的基因术语,以及具有其他生物学含义的基因术语。我们的两种基准排名策略在性能上非常相似。我们的三个基于LocusLink的策略中有两个提供了重大改进。这些方法即使在基因术语有歧义的情况下也能很好地发挥作用。我们的最佳排名策略在三种不同类型的歧义方面比我们的两个基线策略有了显着的改善(根据基线的不同,改进幅度从15.9%到17.7%和11.7%到13.3%不等)。对于大多数基因来说,最好的排序查询是根据LocusLink(现在的Entrez基因)摘要和产品信息以及基因名称和别名构建的查询。对其他人来说,基因名称和别名就足够了。我们还提出了一种方法,对于给定的基因,成功地预测这两个排序查询中的哪一个更合适。我们探讨了不同的后期检索策略对PubMed针对人类基因查询返回的文档进行排序的影响。我们已经成功地应用了其中一些策略来提高检索到的集合中相关文档的排名。即使在遇到各种模棱两可的情况下,这一点也是正确的。我们认为,在PubMed搜索结果上应用我们的策略将非常有用,因为这些搜索结果没有以任何方式按相关性排序。对于检索大量文档的查询尤其如此。
Accuracy of document retrieval from MEDLINE for gene queries is crucially important for many applications in bioinformatics. We explore five information retrieval-based methods to rank documents retrieved by PubMed gene queries for the human genome. The aim is to rank relevant documents higher in the retrieved list. We address the special challenges faced due to ambiguity in gene nomenclature: gene terms that refer to multiple genes, gene terms that are also English words, and gene terms that have other biological meanings. Our two baseline ranking strategies are quite similar in performance. Two of our three LocusLink-based strategies offer significant improvements. These methods work very well even when there is ambiguity in the gene terms. Our best ranking strategy offers significant improvements on three different kinds of ambiguities over our two baseline strategies (improvements range from 15.9% to 17.7% and 11.7% to 13.3% depending on the baseline). For most genes the best ranking query is one that is built from the LocusLink (now Entrez Gene) summary and product information along with the gene names and aliases. For others, the gene names and aliases suffice. We also present an approach that successfully predicts, for a given gene, which of these two ranking queries is more appropriate. We explore the effect of different post-retrieval strategies on the ranking of documents returned by PubMed for human gene queries. We have successfully applied some of these strategies to improve the ranking of relevant documents in the retrieved sets. This holds true even when various kinds of ambiguity are encountered. We feel that it would be very useful to apply strategies like ours on PubMed search results as these are not ordered by relevance in any way. This is especially so for queries that retrieve a large number of documents.
DOI: 10.1186/1471-2105-6-149
发表时间: 2005-06-16
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Schijvenaars, BJA;Mons, B;Kors, JA
通讯作者: Kors, JA
DOI: 10.1006/jbin.2001.1023
发表时间: 2001-08-01
影响因子: 4.5
作者:
Liu, HF;Lussier, YA;Friedman, C
通讯作者: Friedman, C
生物公约概述:生物学信息提取的批判性评估。
DOI: 10.1186/1471-2105-6-s1-s1
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Hirschman L;Yeh A;Blaschke C;Valencia A
通讯作者: Valencia A
DOI: 10.1093/bioinformatics/bth496
发表时间: 2005-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chen, LF;Liu, HF;Friedman, C
通讯作者: Friedman, C
DOI: 10.1016/s1532-0464(03)00014-5
发表时间: 2002-08-01
影响因子: 4.5
作者:
Hirschman, L;Morgan, AA;Yeh, AS
通讯作者: Yeh, AS