Harnessing large language models (LLMs) for candidate gene prioritization and selection.

Harnessing large language models (LLMs) for candidate gene prioritization and selection.
复制标题

利用大型语言模型(LLM)进行候选基因优先级排序和选择。

DOI:
10.1186/s12967-023-04576-8
复制
发表时间:
2023-10-16
影响因子:
7.4
通讯作者:
Chaussabel D
Chaussabel D
中科院分区:
医学2区
文献类型:
--
作者:
Toufiq M;Rinchai D;Bettacchioli E;Kabeer BSA;Khan T;Subba B;White O;Yurieva M;George J;Jourde-Chiche N;Chiche L;Palucka K;Chaussabel D

文献摘要

参考文献

相似文献

特征选择是将系统规模分子分析所带来的进步转化为可操作的临床见解的关键步骤。虽然数据驱动的方法通常用于选择候选基因,但知识驱动的方法必须应对有效筛选大量生物医学信息的挑战。这项工作旨在评估大型语言模型 (LLM) 在知识驱动的基因优先级排序和选择中的效用。在这个概念验证中,我们重点关注与红细胞特征相关的 11 个血液转录模块。我们评估了四位跨多个任务的领先法学硕士。接下来,我们建立了一个利用法学硕士的工作流程。步骤包括: (1) 选择 11 个模块之一; (2) 利用法学硕士鉴定组成基因之间的功能趋同; (3) 根据六个标准对候选基因进行评分,捕捉基因的生物学和临床相关性; (4) 对候选基因进行优先排序并总结理由; (5) 事实核查理由并确定支持性参考资料; (6) 根据经过验证的评分理由选择最佳候选基因; (7)考虑转录组分析数据以最终确定顶级候选基因的选择。在评估的四位法学硕士中,OpenAI 的 GPT-4 和 Anthropic 的 Claude 表现出了最佳性能,并被选用于实施候选基因优先级排序和选择工作流程。数据挖掘研讨会的参与者对 11 个红细胞模块中的每一个模块并行运行此工作流程。模块 M9.2 作为说明性用例。对构成该模块的 30 个候选基因进行了评估,得分最高的 5 个基因被确定为 BCL2L1、ALAS2、SLC4A1、CA1 和 FECH。研究人员仔细核实了总结的评分理由,之后提示法学硕士根据这些信息选择最佳候选人。 GPT-4最初选择了BCL2L1,而Claude选择了ALAS2。当提供来自三个参考数据集的转录分析数据作为附加上下文时,GPT-4 将其最初选择修改为 ALAS2,而 Claude 重申了其对该模块的原始选择。总而言之,我们的研究结果凸显了法学硕士以最少的人为干预对候选基因进行优先排序的能力。这表明该技术具有提高生产力的潜力,特别是对于需要利用广泛的生物医学知识的任务。在线版本包含可在 10.1186/s12967-023-04576-8 获取的补充材料。
Feature selection is a critical step for translating advances afforded by systems-scale molecular profiling into actionable clinical insights. While data-driven methods are commonly utilized for selecting candidate genes, knowledge-driven methods must contend with the challenge of efficiently sifting through extensive volumes of biomedical information. This work aimed to assess the utility of large language models (LLMs) for knowledge-driven gene prioritization and selection. In this proof of concept, we focused on 11 blood transcriptional modules associated with an Erythroid cells signature. We evaluated four leading LLMs across multiple tasks. Next, we established a workflow leveraging LLMs. The steps consisted of: (1) Selecting one of the 11 modules; (2) Identifying functional convergences among constituent genes using the LLMs; (3) Scoring candidate genes across six criteria capturing the gene’s biological and clinical relevance; (4) Prioritizing candidate genes and summarizing justifications; (5) Fact-checking justifications and identifying supporting references; (6) Selecting a top candidate gene based on validated scoring justifications; and (7) Factoring in transcriptome profiling data to finalize the selection of the top candidate gene. Of the four LLMs evaluated, OpenAI's GPT-4 and Anthropic's Claude demonstrated the best performance and were chosen for the implementation of the candidate gene prioritization and selection workflow. This workflow was run in parallel for each of the 11 erythroid cell modules by participants in a data mining workshop. Module M9.2 served as an illustrative use case. The 30 candidate genes forming this module were assessed, and the top five scoring genes were identified as BCL2L1, ALAS2, SLC4A1, CA1, and FECH. Researchers carefully fact-checked the summarized scoring justifications, after which the LLMs were prompted to select a top candidate based on this information. GPT-4 initially chose BCL2L1, while Claude selected ALAS2. When transcriptional profiling data from three reference datasets were provided for additional context, GPT-4 revised its initial choice to ALAS2, whereas Claude reaffirmed its original selection for this module. Taken together, our findings highlight the ability of LLMs to prioritize candidate genes with minimal human intervention. This suggests the potential of this technology to boost productivity, especially for tasks that require leveraging extensive biomedical knowledge. The online version contains supplementary material available at 10.1186/s12967-023-04576-8.
DOI: 10.1093/bioinformatics/btu638
发表时间: 2015-01-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Anders S;Pyl PT;Huber W
通讯作者: Huber W
DOI: 10.1016/j.yexcr.2012.01.021
发表时间: 2012-07-01
影响因子: 3.7
作者:
Ottina, Eleonora;Tischner, Denise;Herold, Marco J.;Villunger, Andreas
通讯作者: Villunger, Andreas
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1038/nrclinonc.2010.227
发表时间: 2011-03-01
影响因子: 78.8
作者:
Hood, Leroy;Friend, Stephen H.
通讯作者: Friend, Stephen H.