Tuning Large language models to read biological literature
Tuning Large language models to read biological literature
批准号:
BB/Y514032/1
负责人:
Antony McCabe
金额:
$23.78万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
In this application, we focus on two related bioinformatics challenges that require interpretation and knowledge extraction from biological and biomedical literature at great scale.First, gene/genome databases store information on gene function, which is ultimately derived from scientific experiments with results reported in publications. It is exceptionally time-consuming and expensive for human curators to read all relevant scientific literature, interpret what has reported about the function or localisation of gene products, and assign specific controlled vocabulary terms (e.g. Gene Ontology terms) or short free text descriptions (gene names or product descriptions.Second, there are enormous volumes of raw data sets accompanying scientific publications, which are deposited in archival databases from expensive omics experiments, including mass spectrometry (MS) proteomics. Our group and others develop and apply pipelines for re-analysing MS data for new purposes, including annotating genomes, discovery of post-translational modifications and building quantitative atlases of species or tissues amongst others. There is a major bottleneck interpreting the original experimental design, sample descriptions and software parameters, which are currently described in blocks of free text submitted to the archival repository or within Materials and Methods sections of accompanying articles. For both challenges, we believe that with the recent extraordinary improvements in large language models (LLMs), they can be retrained and harnessed for these tasks, to remove the bottleneck in knowledge extraction from literature. Our group has significant expertise in bioinformatics and machine learning, but limited expertise in natural language processing (NLP) to date. In this international partnering application, we are collaborating with a leading group in artificial intelligence and NLP from the University of Pennsylvania (UPenn). The UPenn team will help to guide us in the optimal approach for re-training open source LLMs, using training data that our team has amassed over many years. We will produce open source code for the two challenge areas, with a longer term plan to put these into production within the context of major international databases and consortia, within which we have leading roles.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:黄洛将
-
依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:黄洛将
-
依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
-
批准号:12074246
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2020
-
负责人:Yoshitomo Kamiya
-
依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
-
批准号:31972875
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:石江华
-
依托单位:
Large PB/PB小鼠 视网膜新生血管模型的研究
-
批准号:30971650
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2009
-
负责人:周旻
-
依托单位:
基因discs large在果蝇卵母细胞的后端定位及其体轴极性形成中的作用机制
-
批准号:30800648
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2008
-
负责人:于玲珠
-
依托单位:
LARGE基因对口腔癌细胞中α-DG糖基化及表达的分子调控
-
批准号:30772435
-
项目类别:面上项目
-
资助金额:29.0万元
-
批准年份:2007
-
负责人:尚政军
-
依托单位: