A Data-Driven Approach for Extracting "the Most Specific Term" for Ontology Development

A Data-Driven Approach for Extracting "the Most Specific Term" for Ontology Development
复制标题

一种数据驱动的方法,为本体开发提取“最具体的术语”

DOI:
--
复制
发表时间:
2003
期刊:
American Medical Informatics Association Annual Symposium
影响因子:
--
通讯作者:
C. Chute
C. Chute
中科院分区:
--
文献类型:
--
作者:
G. Savova;M. Harris;Thomas M. Johnson;Serguei V. S. Pakhomov;C. Chute

文献摘要

参考文献

相似文献

我们提出了一个数据驱动的方法来提取“最具体”的功能,残疾和健康的本体相关的条款。该算法是统计和语言方法的组合。统计过滤器是基于在一个给定的文本字符串中的内容词的频率;语言启发式是现有算法的扩展,但超越了名词短语,并制定为一个“完整的语法节点”。因此,它可以应用于特定域中感兴趣的任何句法节点。两个测试集由三位专家评分。测试集1是来自疼痛摘要的构造良好的文本;测试集2是实际的医疗报告。结果报告为召回率、精确率、F分数和假阳性中的有效术语率。目前研究的一个局限是测试集相对较小。
We present a data-driven approach to extract the "most specific" terms relevant to an ontology of functioning, disability and health. The algorithm is a combination of statistical and linguistic approaches. The statistical filter is based on the frequency of the content words in a given text string; the linguistic heuristic is an extension of existing algorithms but goes beyond noun phrases and is formulated as a "complete syntactic node". Thus, it can be applied to any syntactic node of interest in the particular domain. Two test sets were marked by three experts. Test set 1 is a well-constructed text from pain abstracts; test set 2 is actual medical reports. Results are reported as recall, precision, F-score and rate of valid terms in false positives. A limitation of the current research is the relatively small test set.
DOI: 10.1016/0277-9536(94)90294-1
发表时间: 1994-01-01
影响因子: 5.4
作者:
VERBRUGGE, LM;JETTE, AM
通讯作者: JETTE, AM