Enhancing phenotype recognition in clinical notes using large language models: PhenoBCBERT and PhenoGPT.
Enhancing phenotype recognition in clinical notes using large language models: PhenoBCBERT and PhenoGPT.
复制标题
DOI:
10.1016/j.patter.2023.100887
复制
发表时间:
2024-01-12
期刊:
影响因子:
6.5
通讯作者:
Wang, Kai
中科院分区:
文献类型:
--
作者:
Yang, Jingye;Liu, Cong;Deng, Wendy;Wu, Da;Weng, Chunhua;Zhou, Yunyun;Wang, Kai
To enhance phenotype recognition in clinical notes of genetic diseases, we developed two models—PhenoBCBERT and PhenoGPT—for expanding the vocabularies of Human Phenotype Ontology (HPO) terms. While HPO offers a standardized vocabulary for phenotypes, existing tools often fail to capture the full scope of phenotypes due to limitations from traditional heuristic or rule-based approaches. Our models leverage large language models to automate the detection of phenotype terms, including those not in the current HPO. We compare these models with PhenoTagger, another HPO recognition tool, and found that our models identify a wider range of phenotype concepts, including previously uncharacterized ones. Our models also show strong performance in case studies on biomedical literature. We evaluate the strengths and weaknesses of BERT- and GPT-based models in aspects such as architecture and accuracy. Overall, our models enhance automated phenotype detection from clinical texts, improving downstream analyses on human diseases. This study uses LLMs to effectively detect abnormal phenotypes in clinical notes It evaluates various parameters of open-source and closed-source models We cover significant details on adapting general LLMs for domain-specific research Large language models (LLMs) are types of artificial intelligence models that leverage massive training datasets to achieve broad capabilities to “understand” and generate natural language content, for example, in the form of text. LLMs are being actively developed by both commercial companies and academic researchers to power many popular AI tools. Given their broad capabilities, there is widespread interest in how these models can be adapted to more specific technical tasks. One such area of interest is the development of LLMs that can understand, process, and reveal new insights from electronic health record data in ways that might help improve healthcare services or guide treatment. This study illuminates the potential of large language models (LLMs) in identifying key medical information, specifically clinical phenotypes, within the intricate notes of rare genetic disease patients. By comparing different LLM approaches and their efficacy in this precise task, the research underscores significant variations in their operational aspects. Crucially, the findings serve as a beacon for diverse fields, guiding the adaptation of advanced LLMs for specialized, real-world applications, thus bridging the gap between AI technology and domain-specific challenges.
登录
查看更多内容
影响因子:
3
作者:
Cho, Hyejin;Lee, Hyunju
通讯作者:
Lee, Hyunju
影响因子:
12.3
作者:
Havrilla JM;Liu C;Dong X;Weng C;Wang K
通讯作者:
Wang K
影响因子:
17.1
作者:
通讯作者:
--
影响因子:
5.3
作者:
Anazi, Shams;Maddirevula, Sateesh;Alkuraya, Fowzan S.
通讯作者:
Alkuraya, Fowzan S.
DOI:
10.1093/jamia/ocac219
发表时间:
2022-11-23
影响因子:
6.4
作者:
Chambon, Pierre J.;Wu, Christopher;Langlotz, Curtis P.
通讯作者:
Langlotz, Curtis P.