Building a specialized lexicon for breast cancer clinical trial subject eligibility analysis

Building a specialized lexicon for breast cancer clinical trial subject eligibility analysis
复制标题

DOI:
10.1177/1460458221989392
复制
发表时间:
2021-01-01
影响因子:
3
通讯作者:
Gaudioso, Carmelo
Gaudioso, Carmelo
中科院分区:
医学3区
文献类型:
--
作者:
Jung, Euisung;Jain, Hemant;Gaudioso, Carmelo

文献摘要

被引文献

相似文献

自然语言处理(NLP)应用程序需要复杂的词汇资源来支持其处理目标。在医疗保健信息学文献中已经提出了不同的解决方案,例如词典查找和MetaMap,以识别具有一个以上单词的疾病术语(多个Gram疾病命名实体)。虽然生物医学领域在蛋白质和基因命名实体的识别方面已经做了大量的工作,但在临床试验受试者资格分析中术语的识别和解析方面的研究还很少。在这项研究中,我们开发了一个专门的词典来改进乳腺癌领域的自然语言处理和文本挖掘分析,并将其与医学临床术语系统化命名法(SNOMED CT)进行了比较。我们使用一种混合方法,它结合了领域专家的知识、来自多个在线词典的术语,以及从样本临床试验中挖掘文本。使用我们的方法引入了4243个唯一词典条目,使二元语法实体匹配提高了38.6%,三元语法实体匹配提高了41%。我们的词典增加了大量的新术语,对于根据资格匹配自动将患者与临床试验匹配非常有用。除了临床试验匹配,本研究开发的专业词典可以作为未来医疗文本挖掘应用的基础。
A natural language processing (NLP) application requires sophisticated lexical resources to support its processing goals. Different solutions, such as dictionary lookup and MetaMap, have been proposed in the healthcare informatics literature to identify disease terms with more than one word (multi-gram disease named entities). Although a lot of work has been done in the identification of protein- and gene-named entities in the biomedical field, not much research has been done on the recognition and resolution of terminologies in the clinical trial subject eligibility analysis. In this study, we develop a specialized lexicon for improving NLP and text mining analysis in the breast cancer domain, and evaluate it by comparing it with the Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT). We use a hybrid methodology, which combines the knowledge of domain experts, terms from multiple online dictionaries, and the mining of text from sample clinical trials. Use of our methodology introduces 4243 unique lexicon items, which increase bigram entity match by 38.6% and trigram entity match by 41%. Our lexicon, which adds a significant number of new terms, is very useful for matching patients to clinical trials automatically based on eligibility matching. Beyond clinical trial matching, the specialized lexicon developed in this study could serve as a foundation for future healthcare text mining applications.