Automated Extraction of Reported Statistical Analyses: Towards a Logical Representation of Clinical Trial Literature
Automated Extraction of Reported Statistical Analyses: Towards a Logical Representation of Clinical Trial Literature
复制标题
自动提取报告的统计分析:实现临床试验文献的逻辑表示
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
R. Taira
中科院分区:
文献类型:
--
作者:
William Hsu;W. Speier;R. Taira
Randomized controlled trials are an important source of evidence for guiding clinical decisions when treating a patient. However, given the large number of studies and their variability in quality, determining how to summarize reported results and formalize them as part of practice guidelines continues to be a challenge. We have developed a set of information extraction and annotation tools to automate the identification of key information from papers related to the hypothesis, sample size, statistical test, confidence interval, significance level, and conclusions. We adapted the Automated Sequence Annotation Pipeline to map extracted phrases to relevant knowledge sources. We trained and tested our system on a corpus of 42 full-text articles related to chemotherapy of non-small cell lung cancer. On our test set of 7 papers, we obtained an overall precision of 86%, recall of 78%, and an F-score of 0.82 for classifying sentences. This work represents our efforts towards utilizing this information for quality assessment, meta-analysis, and modeling.
DOI:
10.1056/nejmsa1012065
发表时间:
2011-03-03
期刊:
The New England journal of medicine
影响因子:
--
作者:
Zarin DA;Tse T;Williams RJ;Califf RM;Ide NC
通讯作者:
Ide NC