Overview of BioCreative II gene normalization.

Overview of BioCreative II gene normalization.
复制标题

DOI:
10.1186/gb-2008-9-s2-s3
复制
发表时间:
2008
期刊:
影响因子:
12.3
通讯作者:
Hirschman L
Hirschman L
中科院分区:
生物学1区
文献类型:
--
作者:
Morgan AA;Lu Z;Wang X;Cohen AM;Fluck J;Ruch P;Divoli A;Fundel K;Leaman R;Hakenberg J;Sun C;Liu HH;Torres R;Krauthammer M;Lau WW;Liu H;Hsu CN;Schuemie M;Cohen KB;Hirschman L

文献摘要

参考文献

被引文献

相似文献

基因归一化任务的目标是将文献中提及的基因或基因产物与生物数据库相链接。这是准确检索生物文献的关键步骤。即使对于人类专家来说,这也是一项具有挑战性的任务;基因常常是被描述而非通过基因符号被提及,而且令人困惑的是,一个基因名称可能指的是不同的基因(通常来自不同的生物体)。对于BioCreative II,任务是列出PubMed/MEDLINE摘要中提及的人类基因或基因产物的Entrez基因标识符。我们选择了与先前为人类基因整理的文章相关的摘要。我们提供了281篇由专家标注的摘要,其中包含684个基因标识符用于训练,以及一个包含262份文档的盲测集,其中包含785个标识符,并有专家标注者创建的金标准。标注者间的一致性测量值超过90%。 20个小组各自提交了1到3次运行结果,总共54次运行。三个系统实现了介于0.80到0.81之间的F值(平衡的准确率和召回率)。使用简单的投票方案和分类器组合系统输出获得了改进的结果;最佳的组合系统在10倍交叉验证下实现了0.92的F值。一个基于所有参与者汇总响应的“最大召回率”系统给出了0.97的召回率(准确率为0.23),在785个标识符中识别出763个。 BioCreative II基因归一化任务的主要进展包括更广泛的参与(20个小组对比8个小组)以及与人类专家相当的组合系统性能,一致性超过90%。这些结果显示出作为将文献与生物数据库相链接的工具的前景。
The goal of the gene normalization task is to link genes or gene products mentioned in the literature to biological databases. This is a key step in an accurate search of the biological literature. It is a challenging task, even for the human expert; genes are often described rather than referred to by gene symbol and, confusingly, one gene name may refer to different genes (often from different organisms). For BioCreative II, the task was to list the Entrez Gene identifiers for human genes or gene products mentioned in PubMed/MEDLINE abstracts. We selected abstracts associated with articles previously curated for human genes. We provided 281 expert-annotated abstracts containing 684 gene identifiers for training, and a blind test set of 262 documents containing 785 identifiers, with a gold standard created by expert annotators. Inter-annotator agreement was measured at over 90%. Twenty groups submitted one to three runs each, for a total of 54 runs. Three systems achieved F-measures (balanced precision and recall) between 0.80 and 0.81. Combining the system outputs using simple voting schemes and classifiers obtained improved results; the best composite system achieved an F-measure of 0.92 with 10-fold cross-validation. A 'maximum recall' system based on the pooled responses of all participants gave a recall of 0.97 (with precision 0.23), identifying 763 out of 785 identifiers. Major advances for the BioCreative II gene normalization task include broader participation (20 versus 8 teams) and a pooled system performance comparable to human experts, at over 90% agreement. These results show promise as tools to link the literature with biological databases.
DOI: 10.1186/1471-2105-6-s1-s15
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Fundel K;Güttler D;Zimmer R;Apostolakis J
通讯作者: Apostolakis J
DOI: 10.1186/1471-2105-6-s1-s13
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Crim J;McDonald R;Pereira F
通讯作者: Pereira F
DOI: 10.1186/1471-2105-6-s1-s12
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Colosimo ME;Morgan AA;Yeh AS;Colombe JB;Hirschman L
通讯作者: Hirschman L
DOI: 10.1006/jcss.1997.1504
发表时间: 1997-08-01
影响因子: 1.1
作者:
Freund, Y;Schapire, RE
通讯作者: Schapire, RE
DOI: 10.1093/nar/gki470
发表时间: 2005-07-01
影响因子: 14.9
作者:
Doms A;Schroeder M
通讯作者: Schroeder M