A new algorithm for construction specific field terms using co-occurrence words information

A new algorithm for construction specific field terms using co-occurrence words information
复制标题

一种利用共现词信息构建特定领域术语的新算法

DOI:
10.1109/mwscas.2003.1562453
复制
发表时间:
2003
期刊:
2003 46th Midwest Symposium on Circuits and Systems
影响因子:
--
通讯作者:
J. Aoe
J. Aoe
中科院分区:
--
文献类型:
--
作者:
E. Atlam;E. Ghada;M. Fuketa;J. Aoe

文献摘要

参考文献

被引文献

相似文献

读者可以通过阅读一些特定的词来了解许多文档字段的主题,这些词被称为字段关联(FA)术语。构造这些FA项对于从部分文档中的少量文字信息中正确确定文档字段是非常重要的。如果这些FA项的数量多,频率高,则可以有效地确定字段。如果第一级FA词(直接连接到终端字段的词)的数量有限,那么传统的方法就不能方便、快速地确定平铺的文档,特别是在语料库文档数量较少的情况下。本文提出了一种新的确定FA术语的方法,该方法利用与窄关联类别相关的同现词和偏误词的权重来消除FA术语的歧义。此外,有效的FA条款是很难被提取出来的唯一的信息,他们的频率。本文提出了一种新的有效的方法,使用新的共现词的权重,使准确率和召回率比频率的情况下更高。
Readers can know the subject of many document fields by reading only some specific words called field association (FA) terms. It is very important to construct these FA terms to decide correctly the document fields from few words information in part of file. The field can be decided efficiency if the number of these FA terms is many and the frequency rate is high. If the number of level I (words that direct connect to terminal fields) FA word is limited, old methods can not determine the documents tiled easily and fast, special when there is a small number of corpus documents. This paper proposes a new method for deciding FA terms using the weight of co-occurrence words and declinable words which related to a narrow association category with eliminating FA terms ambiguity. Moreover, efficient FA terms are difficult to be extracted only by the information of the frequency of them. This paper proposed a new efficient method using new cooccurrence words weight which makes precision and recall are higher than the case of degree of frequency.
DOI: 10.1016/s0306-4573(03)00019-0
发表时间: 2003-11-01
影响因子: 8.6
作者:
Atlam, ES;Fuketa, M;Aoe, J
通讯作者: Aoe, J