A simple approach for protein name identification: prospects and limits.

A simple approach for protein name identification: prospects and limits.
复制标题

DOI:
10.1186/1471-2105-6-s1-s15
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Apostolakis J
Apostolakis J
中科院分区:
生物学4区
文献类型:
--
作者:
Fundel K;Güttler D;Zimmer R;Apostolakis J

文献摘要

被引文献

相似文献

生物知识的重要部分仅以无结构文本的形式存在于生物医学期刊的文章中。通过自动识别基因和基因产物(蛋白质)名称,并将其映射到唯一的数据库标识符,就有可能从文章和各种数据源中提取和整合信息。 我们提出了一种简单有效的方法,该方法可识别文本中的基因和蛋白质名称,并为匹配项返回数据库标识符。它在最近由一个独立评审团进行的BioCreAtIvE实体提取和提及归一化任务中得到了评估。 我们的方法基于使用同义词列表,这些列表将每个基因/蛋白质的唯一数据库标识符映射到不同的同义词名称。对于酵母和小鼠,使用了组织者提供的同义词列表,这些组织者是从公共模式生物数据库中生成的。果蝇的同义词列表是直接从相应的生物数据库中生成的。然后,这些列表在很大程度上通过自动化程序进行了广泛的整理,并通过精确的文本匹配与MEDLINE摘要进行匹配。设计并应用了基于规则和基于支持向量机的后置过滤器以提高精度。 我们的方法在BioCreAtIvE评估(任务1B)中对酵母的F值为0.897,对小鼠为0.764/0.773,在后续评估中对果蝇为0.768,显示出较高的召回率和精度。 结果接近所有提交结果中的最佳值。根据同义词的特性,考虑上下文并过滤掉错误匹配可能至关重要。这对于果蝇尤其重要,因为果蝇在蛋白质名称识别任务中具有极具挑战性的命名法。在这里,基于支持向量机的后置过滤器被证明是非常有效的。
Significant parts of biological knowledge are available only as unstructured text in articles of biomedical journals. By automatically identifying gene and gene product (protein) names and mapping these to unique database identifiers, it becomes possible to extract and integrate information from articles and various data sources. We present a simple and efficient approach that identifies gene and protein names in texts and returns database identifiers for matches. It has been evaluated in the recent BioCreAtIvE entity extraction and mention normalization task by an independent jury. Our approach is based on the use of synonym lists that map the unique database identifiers for each gene/protein to the different synonym names. For yeast and mouse, synonym lists were used as provided by the organizers who generated them from public model organism databases. The synonym list for fly was generated directly from the corresponding organism database. The lists were then extensively curated in largely automated procedure and matched against MEDLINE abstracts by exact text matching. Rule-based and support vector machine-based post filters were designed and applied to improve precision. Our procedure showed high recall and precision with F-measures of 0.897 for yeast and 0.764/0.773 for mouse in the BioCreAtIvE assessment (Task 1B) and 0.768 for fly in a post-evaluation. The results were close to the best over all submissions. Depending on the synonym properties it can be crucial to consider context and to filter out erroneous matches. This is especially important for fly, which has a very challenging nomenclature for the protein name identification task. Here, the support vector machine-based post filter proved to be very effective.