A dictionary to identify small molecules and drugs in free text

A dictionary to identify small molecules and drugs in free text
复制标题

DOI:
10.1093/bioinformatics/btp535
复制
发表时间:
2009-11-15
期刊:
影响因子:
5.8
通讯作者:
Kors, Jan A.
Kors, Jan A.
中科院分区:
生物学3区
文献类型:
--
作者:
Hettne, Kristina M.;Stierum, Rob H.;Kors, Jan A.

文献摘要

被引文献

相似文献

动机:从科学界,在文本中正确识别基因和蛋白质名称的正确识别上花费了很多精力,而在正确识别化学名称上的努力减少了。基于字典的术语识别有能力识别文献中化学信息的多样化表示,并将化学物质映射到其数据库标识符。回报:我们开发了一个词典,用于识别文本中的小分子和药物,结合UMLS的信息,结合UMLS的信息,网格,Chebi,药品库,KEGG,HMDB和Chemidplus。采用了基于规则的术语过滤,对高度频繁的条款进行手动检查和歧义规则。我们测试了从带注释的语料库上的单个资源得出的联合词典和词典,并得出以下结论:(i)每个不同的处理步骤都会随着召回的少量损失而提高了精度; (ii)合并词典的总体性能是可以接受的(精度为0.67,召回0.40(琐碎的名称为0.80);(iii)联合词典的执行效果优于化学识别器oscar3中的字典;(iv)词典的性能仅基于ChemIdplus,与联合字典的性能相当。
Motivation: From the scientific community, a lot of effort has been spent on the correct identification of gene and protein names in text, while less effort has been spent on the correct identification of chemical names. Dictionary-based term identification has the power to recognize the diverse representation of chemical information in the literature and map the chemicals to their database identifiers.Results: We developed a dictionary for the identification of small molecules and drugs in text, combining information from UMLS, MeSH, ChEBI, DrugBank, KEGG, HMDB and ChemIDplus. Rule-based term filtering, manual check of highly frequent terms and disambiguation rules were applied. We tested the combined dictionary and the dictionaries derived from the individual resources on an annotated corpus, and conclude the following: (i) each of the different processing steps increase precision with a minor loss of recall; (ii) the overall performance of the combined dictionary is acceptable (precision 0.67, recall 0.40 (0.80 for trivial names); (iii) the combined dictionary performed better than the dictionary in the chemical recognizer OSCAR3; (iv) the performance of a dictionary based on ChemIDplus alone is comparable to the performance of the combined dictionary.