Learning to rank figures within a biomedical article.

Learning to rank figures within a biomedical article.
复制标题

DOI:
10.1371/journal.pone.0061567
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Yu H
Yu H
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Liu F;Yu H

文献摘要

参考文献

被引文献

相似文献

生物医学文献中有数以亿计的数字,代表着重要的生物医学实验证据。不断增加的数据量使得科学家很难有效和准确地获取他们感兴趣的数据,这一过程对于验证研究事实和制定或测试新的研究假设至关重要。目前的数字搜索应用程序不能完全满足这一挑战,因为“数字袋”的假设没有考虑到数字之间的关系。在我们之前的研究中,数百名生物医学研究人员在他们担任通讯作者的文章中进行了注释。他们根据一个数字的重要性对论文中的每个数字进行排名,称为“数字排名”。使用这个注释数据的集合,我们研究了自动排列数字的计算方法。我们利用并扩展了最先进的列表学习排名算法,并开发了一个新的监督学习模型BioFigRank。交叉验证结果表明,BioFigRank产生了最好的性能相比,其他国家的最先进的计算模型,贪婪的特征选择可以进一步提高排名性能显着。此外,我们通过将BioFigRank与三级竞争领域特定人类专家进行比较来进行评估:(1)第一作者,(2)非作者领域内专家谁不是作者或共同作者的文章,但谁在同一领域的文章的相应作者的作品,以及(3)非作者领域外专家,他不是文章的作者或合著者,并且可能与文章的相应作者在同一领域工作,也可能不工作。我们的研究结果表明,BioFigRank优于非作者域外专家和非作者域内专家。尽管BioFigRank的表现不如第一作者,但由于大多数生物医学研究人员都是文章的域内或域外专家,我们得出结论,BioFigRank代表了一个人工智能系统,它提供专家级的智能,帮助生物医学研究人员有效地浏览日益激增的大数据。
Hundreds of millions of figures are available in biomedical literature, representing important biomedical experimental evidence. This ever-increasing sheer volume has made it difficult for scientists to effectively and accurately access figures of their interest, the process of which is crucial for validating research facts and for formulating or testing novel research hypotheses. Current figure search applications can't fully meet this challenge as the “bag of figures” assumption doesn't take into account the relationship among figures. In our previous study, hundreds of biomedical researchers have annotated articles in which they serve as corresponding authors. They ranked each figure in their paper based on a figure's importance at their discretion, referred to as “figure ranking”. Using this collection of annotated data, we investigated computational approaches to automatically rank figures. We exploited and extended the state-of-the-art listwise learning-to-rank algorithms and developed a new supervised-learning model BioFigRank. The cross-validation results show that BioFigRank yielded the best performance compared with other state-of-the-art computational models, and the greedy feature selection can further boost the ranking performance significantly. Furthermore, we carry out the evaluation by comparing BioFigRank with three-level competitive domain-specific human experts: (1) First Author, (2) Non-Author-In-Domain-Expert who is not the author nor co-author of an article but who works in the same field of the corresponding author of the article, and (3) Non-Author-Out-Domain-Expert who is not the author nor co-author of an article and who may or may not work in the same field of the corresponding author of an article. Our results show that BioFigRank outperforms Non-Author-Out-Domain-Expert and performs as well as Non-Author-In-Domain-Expert. Although BioFigRank underperforms First Author, since most biomedical researchers are either in- or out-domain-experts for an article, we conclude that BioFigRank represents an artificial intelligence system that offers expert-level intelligence to help biomedical researchers to navigate increasingly proliferated big data efficiently.
DOI: 10.1093/bioinformatics/btm301
发表时间: 2007-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hearst, Marti A.;Divoli, Anna;Ye, Jerry
通讯作者: Ye, Jerry
DOI: 10.2214/ajr.06.1740
发表时间: 2007-06-01
影响因子: 5
作者:
Kahn, Charles E., Jr.;Thao, Cheng
通讯作者: Thao, Cheng
DOI: 10.1023/b:inrt.0000009438.69013.fa
发表时间: 2004-01-01
期刊: INFORMATION RETRIEVAL
影响因子: --
作者:
Braschler, M;Peters, C
通讯作者: Peters, C
DOI: 10.1162/0891201053630273
发表时间: 2005-03-01
影响因子: 9.3
作者:
Collins, M;Koo, T
通讯作者: Koo, T
DOI: 10.1093/bib/2.4.363
发表时间: 2001-12-01
影响因子: 9.5
作者:
Marcotte, E;Date, S
通讯作者: Date, S