Evaluation of large language models for discovery of gene set function.

Evaluation of large language models for discovery of gene set function.
复制标题

DOI:
10.21203/rs.3.rs-3270331/v1
复制
发表时间:
2023-09-18
期刊:
Research square
影响因子:
--
通讯作者:
Pratt D
Pratt D
中科院分区:
其他
文献类型:
--
作者:
Hu M;Alkhairy S;Lee I;Pillich RT;Bachelder R;Ideker T;Pratt D

文献摘要

相似文献

基因集分析是功能基因组学的支柱,但它依赖于人工管理的基因功能数据库,这些数据库不完整,也不了解生物学背景。在这里,我们评估了OpenAI的GPT-4,一个大型语言模型(LLM),从其嵌入的生物医学知识中开发关于常见基因功能的假设的能力。我们创建了一个GPT-4管道来用名称标记基因集,这些名称总结了它们的共识功能,并通过分析文本和引用得到证实。在与基因本体论中的命名基因集进行基准比较时,GPT-4在50%的情况下生成了非常相似的名称,而在其余大多数情况下,它恢复了更一般概念的名称。在组学数据中发现的基因组中,GPT-4名称比基因集浓缩更具信息性,具有支持声明和引用,这些声明和引用在很大程度上得到了人类审查的验证。快速合成常见基因功能的能力使LLMS成为有价值的功能基因组学助手。
Gene set analysis is a mainstay of functional genomics, but it relies on manually curated databases of gene functions that are incomplete and unaware of biological context. Here we evaluate the ability of OpenAI’s GPT-4, a Large Language Model (LLM), to develop hypotheses about common gene functions from its embedded biomedical knowledge. We created a GPT-4 pipeline to label gene sets with names that summarize their consensus functions, substantiated by analysis text and citations. Benchmarking against named gene sets in the Gene Ontology, GPT-4 generated very similar names in 50% of cases, while in most remaining cases it recovered the name of a more general concept. In gene sets discovered in ‘omics data, GPT-4 names were more informative than gene set enrichment, with supporting statements and citations that largely verified in human review. The ability to rapidly synthesize common gene functions positions LLMs as valuable functional genomics assistants.