Exploiting MeSH indexing in MEDLINE to generate a data set for word sense disambiguation.

Exploiting MeSH indexing in MEDLINE to generate a data set for word sense disambiguation.
复制标题

DOI:
10.1186/1471-2105-12-223
复制
发表时间:
2011-06-02
期刊:
影响因子:
3
通讯作者:
Aronson AR
Aronson AR
中科院分区:
生物学4区
文献类型:
--
作者:
Jimeno-Yepes AJ;McInnes BT;Aronson AR

文献摘要

参考文献

被引文献

相似文献

词义消歧(WSD)方法在生物医学领域的评价是困难的,因为可用的资源要么太小,要么过于集中于特定类型的实体(例如疾病或基因)。我们提出了一种方法,可用于使用统一医学语言系统(UMLS)元词典和MEDLINE的手动MeSH索引自动开发WSD测试集。我们演示了这种方法的使用,通过开发这样一个数据集,称为MSH WSD。在我们的方法中,元词库首先筛选,以确定可能的含义包括两个或多个MeSH标题的歧义术语。然后,我们使用每个模糊的术语及其相应的MeSH标题来提取MEDLINE引文,其中术语和只有一个MeSH标题共同出现。在MEDLINE引文中找到的术语将自动分配到链接到MeSH标题的UMLS CUI。已为每个实例指定了一个UMLS概念唯一标识符(CUI)。我们比较的MSH WSD数据集的特点,以前现有的NLM WSD数据集。由此产生的MSH WSD数据集包括106个歧义缩写,88个歧义术语和9个两者的组合,共203个歧义实体。对于每个模糊术语/缩写,数据集包含从MEDLINE获得的每个含义最多100个实例。我们评估了可靠性的MSH WSD数据集使用现有的知识为基础的方法,并比较其性能,以前通过这些算法获得的结果对预先存在的数据集,NLM WSD。我们发现,基于知识的方法实现不同的结果,但保持其相对性能,除了期刊描述符索引(JDI)的方法,其性能低于其他方法。MSH WSD数据集允许在生物医学领域评估WSD算法。与以前现有的数据集相比,MSH WSD包含了大量的生物医学术语/缩写,并涵盖了最大的UMLS语义类型集。此外,MSH WSD数据集已自动生成重用已经存在的注释,因此,可以从后续的UMLS版本重新生成。
Evaluation of Word Sense Disambiguation (WSD) methods in the biomedical domain is difficult because the available resources are either too small or too focused on specific types of entities (e.g. diseases or genes). We present a method that can be used to automatically develop a WSD test collection using the Unified Medical Language System (UMLS) Metathesaurus and the manual MeSH indexing of MEDLINE. We demonstrate the use of this method by developing such a data set, called MSH WSD. In our method, the Metathesaurus is first screened to identify ambiguous terms whose possible senses consist of two or more MeSH headings. We then use each ambiguous term and its corresponding MeSH heading to extract MEDLINE citations where the term and only one of the MeSH headings co-occur. The term found in the MEDLINE citation is automatically assigned the UMLS CUI linked to the MeSH heading. Each instance has been assigned a UMLS Concept Unique Identifier (CUI). We compare the characteristics of the MSH WSD data set to the previously existing NLM WSD data set. The resulting MSH WSD data set consists of 106 ambiguous abbreviations, 88 ambiguous terms and 9 which are a combination of both, for a total of 203 ambiguous entities. For each ambiguous term/abbreviation, the data set contains a maximum of 100 instances per sense obtained from MEDLINE. We evaluated the reliability of the MSH WSD data set using existing knowledge-based methods and compared their performance to that of the results previously obtained by these algorithms on the pre-existing data set, NLM WSD. We show that the knowledge-based methods achieve different results but keep their relative performance except for the Journal Descriptor Indexing (JDI) method, whose performance is below the other methods. The MSH WSD data set allows the evaluation of WSD algorithms in the biomedical domain. Compared to previously existing data sets, MSH WSD contains a larger number of biomedical terms/abbreviations and covers the largest set of UMLS Semantic Types. Furthermore, the MSH WSD data set has been generated automatically reusing already existing annotations and, therefore, can be regenerated from subsequent UMLS versions.
DOI: 10.1186/1471-2105-9-s3-s3
发表时间: 2008-04-11
期刊: BMC bioinformatics
影响因子: 3
作者:
Jimeno A;Jimenez-Ruiz E;Lee V;Gaudan S;Berlanga R;Rebholz-Schuhmann D
通讯作者: Rebholz-Schuhmann D
DOI: 10.1197/jamia.m1101
发表时间: 2002-11-01
影响因子: 6.4
作者:
Liu, HF;Johnson, SB;Friedman, C
通讯作者: Friedman, C
DOI: 10.1006/jbin.2001.1023
发表时间: 2001-08-01
影响因子: 4.5
作者:
Liu, HF;Lussier, YA;Friedman, C
通讯作者: Friedman, C
生物公约概述:生物学信息提取的批判性评估。
DOI: 10.1186/1471-2105-6-s1-s1
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Hirschman L;Yeh A;Blaschke C;Valencia A
通讯作者: Valencia A
DOI: 10.1016/j.ijmedinf.2005.03.013
发表时间: 2005-08-01
影响因子: 4.9
作者:
Leroy, G;Rindflesch, TC
通讯作者: Rindflesch, TC