Using text mining techniques to extract phenotypic information from the PhenoCHF corpus.

Using text mining techniques to extract phenotypic information from the PhenoCHF corpus.
复制标题

DOI:
10.1186/1472-6947-15-s2-s3
复制
发表时间:
2015
影响因子:
3.5
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
医学3区
文献类型:
--
作者:
Alnazzawi N;Thompson P;Batista-Navarro R;Ananiadou S

文献摘要

被引文献

相似文献

锁定在非结构化的叙述性文本的表型信息提出了显着的障碍,信息的可访问性,无论是临床医生和用于临床研究目的的计算机化应用程序。文本挖掘(TM)技术已成功地应用于从生物医学领域的文本中提取不同类型的信息。它们有可能被扩展到允许从自由文本中提取与表型相关的信息。为了刺激TM系统的发展,能够从文本中提取表型信息,我们已经创建了一个新的语料库(PhenoCHF),由领域专家注释与充血性心力衰竭有关的几种类型的表型信息。为了确保使用语料库开发的系统对多种文本类型具有鲁棒性,它集成了来自异构源的文本,即,电子健康记录(EHR)和文献中的科学文章。我们已经开发了几种不同的表型提取方法来证明语料库的实用性,并在另一个语料库上测试了这些方法,即,Share/CLEF 2013.对我们的自动化方法的评估表明,PhenoCHF可以促进可靠的表型提取系统的训练,该系统对文本类型的变化具有鲁棒性。这些结果通过在ShARe/CLEF语料库上评估我们的训练系统得到了加强,其中包含各种类型的临床记录。与生物医学领域的其他研究一样,我们发现,当与丰富的特征集相结合时,基于条件随机场的解决方案会产生最佳结果。PhenoCHF是第一个旨在编码详细表型信息的注释语料库。语料库的独特的异质组成已被证明是有利的,在训练系统,可以准确地提取表型信息,从一系列不同的文本类型。虽然我们的注释范围目前仅限于单一疾病,但所取得的有希望的结果可以刺激进一步的工作,以提取其他疾病的表型信息。PhenoCHF注释指南和注释可在https://code.google.com/p/phenochf-corpus上公开获得。
Phenotypic information locked away in unstructured narrative text presents significant barriers to information accessibility, both for clinical practitioners and for computerised applications used for clinical research purposes. Text mining (TM) techniques have previously been applied successfully to extract different types of information from text in the biomedical domain. They have the potential to be extended to allow the extraction of information relating to phenotypes from free text. To stimulate the development of TM systems that are able to extract phenotypic information from text, we have created a new corpus (PhenoCHF) that is annotated by domain experts with several types of phenotypic information relating to congestive heart failure. To ensure that systems developed using the corpus are robust to multiple text types, it integrates text from heterogeneous sources, i.e., electronic health records (EHRs) and scientific articles from the literature. We have developed several different phenotype extraction methods to demonstrate the utility of the corpus, and tested these methods on a further corpus, i.e., ShARe/CLEF 2013. Evaluation of our automated methods showed that PhenoCHF can facilitate the training of reliable phenotype extraction systems, which are robust to variations in text type. These results have been reinforced by evaluating our trained systems on the ShARe/CLEF corpus, which contains clinical records of various types. Like other studies within the biomedical domain, we found that solutions based on conditional random fields produced the best results, when coupled with a rich feature set. PhenoCHF is the first annotated corpus aimed at encoding detailed phenotypic information. The unique heterogeneous composition of the corpus has been shown to be advantageous in the training of systems that can accurately extract phenotypic information from a range of different text types. Although the scope of our annotation is currently limited to a single disease, the promising results achieved can stimulate further work into the extraction of phenotypic information for other diseases. The PhenoCHF annotation guidelines and annotations are publicly available at https://code.google.com/p/phenochf-corpus.