Benchmarking natural-language parsers for biological applications using dependency graphs.

Benchmarking natural-language parsers for biological applications using dependency graphs.
复制标题

DOI:
10.1186/1471-2105-8-24
复制
发表时间:
2007-01-25
期刊:
影响因子:
3
通讯作者:
Shepherd, Adrian J
Shepherd, Adrian J
中科院分区:
生物学4区
文献类型:
--
作者:
Clegg, Andrew B;Shepherd, Adrian J

文献摘要

被引文献

相似文献

在句法解析器在生物学中的自然语言处理问题上的应用中,人们的兴趣正在增加,但是评估其表现很困难,因为语言惯例的差异似乎是错误的。我们提出了一种使用基于依赖图的中间表示来评估其准确性的方法,其中在大多数信息提取任务中重要的语义关系更靠近表面。我们还展示了如何轻松地针对各种应用程序驱动的标准量身定制此方法。 我们使用Genia语料库作为黄金标准,测试了四个在生物信息学项目中使用的开源解析器。我们首先介绍整体性能指标,并在量身定制的子任务上测试了两个领先的Charniak租赁和Bikel Parsers,以反映用于提取基因表达关系的系统的要求。这两个工具在评估中显然超过了其他解析器,并且在先前的生物学评估中,在类似任务上可相当或超过本机依赖解析器的精度水平。 使用依赖图评估可以根据特定生物应用的语义轻松地根据选择的标准来轻松测试解析器,引起人们对重要错误的注意,并吸收许多微不足道的差异,否则这些差异将被报告为错误。从短语结构解析器的输出中生成高准确的依赖图图还提供了对几种自然语言处理技术中使用的更详细的语法树的访问。
Interest is growing in the application of syntactic parsers to natural language processing problems in biology, but assessing their performance is difficult because differences in linguistic convention can falsely appear to be errors. We present a method for evaluating their accuracy using an intermediate representation based on dependency graphs, in which the semantic relationships important in most information extraction tasks are closer to the surface. We also demonstrate how this method can be easily tailored to various application-driven criteria. Using the GENIA corpus as a gold standard, we tested four open-source parsers which have been used in bioinformatics projects. We first present overall performance measures, and test the two leading tools, the Charniak-Lease and Bikel parsers, on subtasks tailored to reflect the requirements of a system for extracting gene expression relationships. These two tools clearly outperform the other parsers in the evaluation, and achieve accuracy levels comparable to or exceeding native dependency parsers on similar tasks in previous biological evaluations. Evaluating using dependency graphs allows parsers to be tested easily on criteria chosen according to the semantics of particular biological applications, drawing attention to important mistakes and soaking up many insignificant differences that would otherwise be reported as errors. Generating high-accuracy dependency graphs from the output of phrase-structure parsers also provides access to the more detailed syntax trees that are used in several natural-language processing techniques.