Multilabel associative classification categorization of MEDLINE articles into MeSH keywords - An intelligent data mining technique to more accurately classify large volumes of documents
Multilabel associative classification categorization of MEDLINE articles into MeSH keywords - An intelligent data mining technique to more accurately classify large volumes of documents
复制标题
DOI:
10.1109/memb.2007.335581
复制
发表时间:
2007-03-01
影响因子:
--
通讯作者:
Reformat, Marek
中科院分区:
文献类型:
--
作者:
Rak, Rafal;Kurgan, Lukasz A.;Reformat, Marek
NLM’s tools and databases have attracted significant attention in recent years. The Text Analysis and Knowledge Mining for Biomedical Documents (MedTAKIMI), which is an application to facilitate knowledge discovery from very large text databases such as the MEDLINE, has been developed and described in [3]. According to the authors, the application dynamically mines documents to obtain their characteristic features and uses categories such as MeSH keywords for term extraction and interactive series of drill-down queries. Another application, called MedMeSH Summarizer, uses MeSH keywords to annotate a set of genes obtained from DNA microarrays by summarizing all the terms tagged to MEDLINE article references that are related to a gene in a user-defined query [4]. Exploration of relationships between features used to represent text, with application to MEDLINE, has been studied in [5]. The method uses association rules and compares three different semantic levels: words, MeSH keywords, and automatically selected concepts coming from NLM’s Unified Medical Language System (UMLS). The authors were especially interested in plausibility and usefulness of the three levels. In our research we use OHSUMED, a corpus subset of the MEDLINE database. This collection has been also used by many researchers to perform classification using MeSH