Enhancing navigation in biomedical databases by community voting and database-driven text classification.

Enhancing navigation in biomedical databases by community voting and database-driven text classification.
复制标题

通过社区投票和数据库驱动的文本分类增强生物医学数据库中的导航。

DOI:
10.1186/1471-2105-10-317
复制
发表时间:
2009-10-03
期刊:
影响因子:
3
通讯作者:
Weissleder R
Weissleder R
中科院分区:
生物学4区
文献类型:
--
作者:
Duchrow T;Shtatland T;Guettler D;Pivovarov M;Kramer S;Weissleder R

文献摘要

参考文献

被引文献

相似文献

The breadth of biological databases and their information content continues to increase exponentially. Unfortunately, our ability to query such sources is still often suboptimal. Here, we introduce and apply community voting, database-driven text classification, and visual aids as a means to incorporate distributed expert knowledge, to automatically classify database entries and to efficiently retrieve them. Using a previously developed peptide database as an example, we compared several machine learning algorithms in their ability to classify abstracts of published literature results into categories relevant to peptide research, such as related or not related to cancer, angiogenesis, molecular imaging, etc. Ensembles of bagged decision trees met the requirements of our application best. No other algorithm consistently performed better in comparative testing. Moreover, we show that the algorithm produces meaningful class probability estimates, which can be used to visualize the confidence of automatic classification during the retrieval process. To allow viewing long lists of search results enriched by automatic classifications, we added a dynamic heat map to the web interface. We take advantage of community knowledge by enabling users to cast votes in Web 2.0 style in order to correct automated classification errors, which triggers reclassification of all entries. We used a novel framework in which the database "drives" the entire vote aggregation and reclassification process to increase speed while conserving computational resources and keeping the method scalable. In our experiments, we simulate community voting by adding various levels of noise to nearly perfectly labelled instances, and show that, under such conditions, classification can be improved significantly. Using PepBank as a model database, we show how to build a classification-aided retrieval system that gathers training data from the community, is completely controlled by the database, scales well with concurrent change events, and can be adapted to add text classification capability to other biomedical databases. The system can be accessed at .
DOI: 10.1023/a:1007515423169
发表时间: 1999-07-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Bauer, E;Kohavi, R
通讯作者: Kohavi, R
DOI: 10.1186/1471-2105-4-11
发表时间: 2003-03-27
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Donaldson, I;Martin, J;de Bruijn, B;Wolting, C;Lay, V;Tuekam, B;Zhang, SD;Baskin, B;Bader, GD;Michalickova, K;Pawson, T;Hogue, CWV
通讯作者: Hogue, CWV
DOI: 10.1038/sj.neo.7900044
发表时间: 1999-11-01
期刊: Neoplasia (New York)
影响因子: --
作者:
Bogdanov, A., Jr.;Marecos, E.;Weissleder, R.
通讯作者: Weissleder, R.
DOI: 10.1093/bioinformatics/btg1011
发表时间: 2003-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Dobrokhotov, Pavel B.;Goutte, Cyril;Gaussier, Eric
通讯作者: Gaussier, Eric
DOI: 10.1093/nar/25.17.3389
发表时间: 1997-09-01
影响因子: 14.9
作者:
Altschul, SF;Madden, TL;Lipman, DJ
通讯作者: Lipman, DJ