Enhancing phenotype recognition in clinical notes using large language models: PhenoBCBERT and PhenoGPT.

Enhancing phenotype recognition in clinical notes using large language models: PhenoBCBERT and PhenoGPT.
复制标题

DOI:
10.1016/j.patter.2023.100887
复制
发表时间:
2024-01-12
期刊:
影响因子:
6.5
通讯作者:
Wang, Kai
Wang, Kai
中科院分区:
其他
文献类型:
--
作者:
Yang, Jingye;Liu, Cong;Deng, Wendy;Wu, Da;Weng, Chunhua;Zhou, Yunyun;Wang, Kai

文献摘要

参考文献

相似文献

为了提高遗传疾病临床记录中的表型识别,我们开发了两个模型- phenobcbert和phenogpt,用于扩展人类表型本体(HPO)术语的词汇量。虽然HPO提供了表型的标准化词汇表,但由于传统的启发式或基于规则的方法的限制,现有工具通常无法捕获表型的全部范围。我们的模型利用大型语言模型来自动检测表型术语,包括那些不在当前HPO中的术语。我们将这些模型与另一个HPO识别工具PhenoTagger进行了比较,发现我们的模型识别了更广泛的表型概念,包括以前未表征的概念。我们的模型在生物医学文献的案例研究中也表现出很强的性能。我们在架构和准确性等方面评估了基于BERT和基于gpt的模型的优缺点。总的来说,我们的模型增强了临床文本的自动表型检测,改善了对人类疾病的下游分析。本研究使用llm来有效地检测临床记录中的异常表型,它评估了开源和闭源模型的各种参数,我们涵盖了将通用llm用于特定领域研究的重要细节。大型语言模型(llm)是一种人工智能模型,它利用大量训练数据集来实现“理解”和生成自然语言内容的广泛能力,例如,以文本的形式。商业公司和学术研究人员都在积极开发法学硕士,为许多流行的人工智能工具提供动力。鉴于它们的广泛功能,人们对如何将这些模型适应于更具体的技术任务产生了广泛的兴趣。其中一个感兴趣的领域是法学硕士的开发,法学硕士可以理解、处理和揭示来自电子健康记录数据的新见解,从而有助于改善医疗保健服务或指导治疗。这项研究阐明了大语言模型(llm)在识别关键医学信息方面的潜力,特别是临床表型,在罕见遗传疾病患者的复杂笔记中。通过比较不同的法学硕士方法及其在这项精确任务中的功效,研究强调了它们在操作方面的显著差异。至关重要的是,这些发现可以作为不同领域的灯塔,指导先进的法学硕士适应专业的、现实世界的应用,从而弥合人工智能技术与特定领域挑战之间的差距。
To enhance phenotype recognition in clinical notes of genetic diseases, we developed two models—PhenoBCBERT and PhenoGPT—for expanding the vocabularies of Human Phenotype Ontology (HPO) terms. While HPO offers a standardized vocabulary for phenotypes, existing tools often fail to capture the full scope of phenotypes due to limitations from traditional heuristic or rule-based approaches. Our models leverage large language models to automate the detection of phenotype terms, including those not in the current HPO. We compare these models with PhenoTagger, another HPO recognition tool, and found that our models identify a wider range of phenotype concepts, including previously uncharacterized ones. Our models also show strong performance in case studies on biomedical literature. We evaluate the strengths and weaknesses of BERT- and GPT-based models in aspects such as architecture and accuracy. Overall, our models enhance automated phenotype detection from clinical texts, improving downstream analyses on human diseases. This study uses LLMs to effectively detect abnormal phenotypes in clinical notes It evaluates various parameters of open-source and closed-source models We cover significant details on adapting general LLMs for domain-specific research Large language models (LLMs) are types of artificial intelligence models that leverage massive training datasets to achieve broad capabilities to “understand” and generate natural language content, for example, in the form of text. LLMs are being actively developed by both commercial companies and academic researchers to power many popular AI tools. Given their broad capabilities, there is widespread interest in how these models can be adapted to more specific technical tasks. One such area of interest is the development of LLMs that can understand, process, and reveal new insights from electronic health record data in ways that might help improve healthcare services or guide treatment. This study illuminates the potential of large language models (LLMs) in identifying key medical information, specifically clinical phenotypes, within the intricate notes of rare genetic disease patients. By comparing different LLM approaches and their efficacy in this precise task, the research underscores significant variations in their operational aspects. Crucially, the findings serve as a beacon for diverse fields, guiding the adaptation of advanced LLMs for specialized, real-world applications, thus bridging the gap between AI technology and domain-specific challenges.
DOI: 10.1186/s12859-019-3321-4
发表时间: 2019-12-30
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Cho, Hyejin;Lee, Hyunju
通讯作者: Lee, Hyunju
DOI: 10.1186/s13073-021-00909-8
发表时间: 2021-05-25
期刊: Genome medicine
影响因子: 12.3
作者:
Havrilla JM;Liu C;Dong X;Weng C;Wang K
通讯作者: Wang K
DOI: 10.1126/scitranslmed.aau9113
发表时间: 2020-05-20
影响因子: 17.1
作者:
通讯作者: --
DOI: 10.1007/s00439-017-1843-2
发表时间: 2017-11-01
期刊: HUMAN GENETICS
影响因子: 5.3
作者:
Anazi, Shams;Maddirevula, Sateesh;Alkuraya, Fowzan S.
通讯作者: Alkuraya, Fowzan S.
DOI: 10.1093/jamia/ocac219
发表时间: 2022-11-23
影响因子: 6.4
作者:
Chambon, Pierre J.;Wu, Christopher;Langlotz, Curtis P.
通讯作者: Langlotz, Curtis P.