Syntactic parsing of clinical text: guideline and corpus development with handling ill-formed sentences

Syntactic parsing of clinical text: guideline and corpus development with handling ill-formed sentences
复制标题

DOI:
10.1136/amiajnl-2013-001810
复制
发表时间:
2013-11-01
影响因子:
6.4
通讯作者:
Huang, Yang
Huang, Yang
中科院分区:
管理学2区
文献类型:
--
作者:
Fan, Jung-wei;Yang, Elly W.;Huang, Yang

文献摘要

被引文献

相似文献

目的开发、评估和共享:(1)临床文本句法分析指南,以及处理不良句子的新方法;(2)根据指南注释的临床树库。为有类似兴趣的读者记录过程和发现。方法使用来自共享自然语言处理挑战数据集的随机样本,我们基于两个机构之间的迭代注释和裁定开发了一本领域定制句法解析指南手册。在处理临床文本中常见的格式不良句子的指南中纳入了特殊考虑。注释者内和注释者间的一致率用于评价遵循指南的一致性。定量和定性属性的注释Treebank,以及它的使用重新训练的统计parser,reported.Results的补充Penn Treebank II指南注释临床句子。在对450个句子进行了三次迭代的注释和裁定之后,注释者在最终的独立集上达到了0.930的F-测量一致率(而注释者内部一致率为0.948)。总共有1100个句子的进展记录进行了注释,表现出特定领域的语言特征。一个统计解析器用组合的一般英语(主要是新闻文本)注释进行了重新训练,我们的注释达到了0.811的准确率(高于单纯用一般或临床句子训练的模型)。指南和句法注释均可在https://sourceforge.net/projects/medicaltreebank.Conclusions上获得。我们制定了用于解析临床文本的指南,并相应地注释了语料库。注释者内部和注释者之间的高一致率显示出遵循指南的良好一致性。语料库被证明是有用的,在重新训练的统计分析器,达到中等精度。
Objective To develop, evaluate, and share: (1) syntactic parsing guidelines for clinical text, with a new approach to handling ill-formed sentences; and (2) a clinical Treebank annotated according to the guidelines. To document the process and findings for readers with similar interest.Methods Using random samples from a shared natural language processing challenge dataset, we developed a handbook of domain-customized syntactic parsing guidelines based on iterative annotation and adjudication between two institutions. Special considerations were incorporated into the guidelines for handling ill-formed sentences, which are common in clinical text. Intra- and inter-annotator agreement rates were used to evaluate consistency in following the guidelines. Quantitative and qualitative properties of the annotated Treebank, as well as its use to retrain a statistical parser, were reported.Results A supplement to the Penn Treebank II guidelines was developed for annotating clinical sentences. After three iterations of annotation and adjudication on 450 sentences, the annotators reached an F-measure agreement rate of 0.930 (while intra-annotator rate was 0.948) on a final independent set. A total of 1100 sentences from progress notes were annotated that demonstrated domain-specific linguistic features. A statistical parser retrained with combined general English (mainly news text) annotations and our annotations achieved an accuracy of 0.811 (higher than models trained purely with either general or clinical sentences alone). Both the guidelines and syntactic annotations are made available at https://sourceforge.net/projects/medicaltreebank.Conclusions We developed guidelines for parsing clinical text and annotated a corpus accordingly. The high intra- and inter-annotator agreement rates showed decent consistency in following the guidelines. The corpus was shown to be useful in retraining a statistical parser that achieved moderate accuracy.