Supporting the annotation of chronic obstructive pulmonary disease (COPD) phenotypes with text mining workflows.
Supporting the annotation of chronic obstructive pulmonary disease (COPD) phenotypes with text mining workflows.
复制标题
DOI:
10.1186/s13326-015-0004-6
复制
发表时间:
2015
影响因子:
1.9
通讯作者:
Ananiadou S
中科院分区:
文献类型:
--
作者:
Fu X;Batista-Navarro R;Rak R;Ananiadou S
Chronic obstructive pulmonary disease (COPD) is a life-threatening lung disorder whose recent prevalence has led to an increasing burden on public healthcare. Phenotypic information in electronic clinical records is essential in providing suitable personalised treatment to patients with COPD. However, as phenotypes are often “hidden” within free text in clinical records, clinicians could benefit from text mining systems that facilitate their prompt recognition. This paper reports on a semi-automatic methodology for producing a corpus that can ultimately support the development of text mining tools that, in turn, will expedite the process of identifying groups of COPD patients. A corpus of 30 full-text papers was formed based on selection criteria informed by the expertise of COPD specialists. We developed an annotation scheme that is aimed at producing fine-grained, expressive and computable COPD annotations without burdening our curators with a highly complicated task. This was implemented in the Argo platform by means of a semi-automatic annotation workflow that integrates several text mining tools, including a graphical user interface for marking up documents. When evaluated using gold standard (i.e., manually validated) annotations, the semi-automatic workflow was shown to obtain a micro-averaged F-score of 45.70% (with relaxed matching). Utilising the gold standard data to train new concept recognisers, we demonstrated that our corpus, although still a work in progress, can foster the development of significantly better performing COPD phenotype extractors. We describe in this work the means by which we aim to eventually support the process of COPD phenotype curation, i.e., by the application of various text mining tools integrated into an annotation workflow. Although the corpus being described is still under development, our results thus far are encouraging and show great potential in stimulating the development of further automatic COPD phenotype extractors. The online version of this article (doi:10.1186/s13326-015-0004-6) contains supplementary material, which is available to authorized users.
登录
查看更多内容
影响因子:
4.3
作者:
Cunningham H;Tablan V;Roberts A;Bontcheva K
通讯作者:
Bontcheva K
影响因子:
14.9
作者:
Hastings J;de Matos P;Dekker A;Ennis M;Harsha B;Kale N;Muthukrishnan V;Owen G;Turner S;Williams M;Steinbeck C
通讯作者:
Steinbeck C
影响因子:
14.9
作者:
Köhler S;Doelken SC;Mungall CJ;Bauer S;Firth HV;Bailleul-Forestier I;Black GC;Brown DL;Brudno M;Campbell J;FitzPatrick DR;Eppig JT;Jackson AP;Freson K;Girdea M;Helbig I;Hurst JA;Jähn J;Jackson LG;Kelly AM;Ledbetter DH;Mansour S;Martin CL;Moss C;Mumford A;Ouwehand WH;Park SM;Riggs ER;Scott RH;Sisodiya S;Van Vooren S;Wapner RJ;Wilkie AO;Wright CF;Vulto-van Silfhout AT;de Leeuw N;de Vries BB;Washingthon NL;Smith CL;Westerfield M;Schofield P;Ruef BJ;Gkoutos GV;Haendel M;Smedley D;Lewis SE;Robinson PN
通讯作者:
Robinson PN
影响因子:
3
作者:
Funk C;Baumgartner W Jr;Garcia B;Roeder C;Bada M;Cohen KB;Hunter LE;Verspoor K
通讯作者:
Verspoor K
影响因子:
3.7
作者:
Dahdul WM;Balhoff JP;Engeman J;Grande T;Hilton EJ;Kothari C;Lapp H;Lundberg JG;Midford PE;Vision TJ;Westerfield M;Mabee PM
通讯作者:
Mabee PM