MedScan, a natural language processing engine for MEDLINE abstracts

MedScan, a natural language processing engine for MEDLINE abstracts
复制标题

DOI:
10.1093/bioinformatics/btg207
复制
发表时间:
2003-09-01
期刊:
影响因子:
5.8
通讯作者:
Daraselia, N
Daraselia, N
中科院分区:
生物学3区
文献类型:
--
作者:
Novichkova, S;Egorov, S;Daraselia, N

文献摘要

被引文献

相似文献

动机:从科学出版物中提取生物医学信息的重要性得到了广泛认可。已经报道了许多针对生物医学领域的信息提取系统,但它们都没有被广泛用于实际应用中。迄今为止,大多数建议对自然语言的句法方面做出了相当简单的假设。迫切需要一个具有广泛覆盖范围并在实际文本应用程序中表现良好的系统:我们提出了一种名为MedScan的一般生物医学域的NLP引擎,该引擎称为MedScan,该引擎有效地从Medline摘要中处理句子并生产一套正则化的逻辑结构表示每个句子的含义。该发动机利用特殊开发的无上下文语法和词典。对系统的性能,准确性和覆盖范围的初步评估表现出令人鼓舞的结果。讨论了进一步的方法,以增加发动机的覆盖范围和减少解析歧义及其信息提取的应用。
Motivation: The importance of extracting biomedical information from scientific publications is well recognized. A number of information extraction systems for the biomedical domain have been reported, but none of them have become widely used in practical applications. Most proposals to date make rather simplistic assumptions about the syntactic aspect of natural language. There is an urgent need for a system that has broad coverage and performs well in real-text applications.Results: We present a general biomedical domain-oriented NLP engine called MedScan that efficiently processes sentences from MEDLINE abstracts and produces a set of regularized logical structures representing the meaning of each sentence. The engine utilizes a specially developed context-free grammar and lexicon. Preliminary evaluation of the system's performance, accuracy, and coverage exhibited encouraging results. Further approaches for increasing the coverage and reducing parsing ambiguity of the engine, as well as its application for information extraction are discussed.