A novel machine learning framework for automated biomedical relation extraction from large-scale literature repositories

A novel machine learning framework for automated biomedical relation extraction from large-scale literature repositories
复制标题

一种新颖的机器学习框架,用于从大规模文献存储库中自动提取生物医学关系

DOI:
10.1038/s42256-020-0189-y
复制
发表时间:
2020-06-01
影响因子:
23.8
通讯作者:
Zeng, Jianyang
Zeng, Jianyang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hong, Lixiang;Lin, Jinjian;Zeng, Jianyang

文献摘要

被引文献

相似文献

关于生物医学实体(例如药物和靶标)之间关系的知识广泛分布在超过 3000 万篇研究文章中,并在生物医学科学的发展中始终发挥着重要作用。在这项工作中,我们提出了一种新颖的机器学习框架,名为 BERE,用于从大规模文献存储库中自动提取生物医学关系。 BERE 使用混合编码网络从语义和句法方面更好地表示每个句子,并采用特征聚合网络在考虑所有相关语句后进行预测。更重要的是,BERE 还可以通过远程监督技术在没有任何人工注释的情况下进行训练。通过广泛的测试,BERE在提取生物医学关系方面表现出了良好的性能,并且还可以找到现有数据库中未报道的有意义的关系,从而为指导湿实验室实验和推进生物知识发现过程提供有用的提示。
Knowledge about the relations between biomedical entities (such as drugs and targets) is widely distributed in more than 30 million research articles and consistently plays an important role in the development of biomedical science. In this work, we propose a novel machine learning framework, named BERE, for automatically extracting biomedical relations from large-scale literature repositories. BERE uses a hybrid encoding network to better represent each sentence from both semantic and syntactic aspects, and employs a feature aggregation network to make predictions after considering all relevant statements. More importantly, BERE can also be trained without any human annotation via a distant supervision technique. Through extensive tests, BERE has demonstrated promising performance in extracting biomedical relations, and can also find meaningful relations that were not reported in existing databases, thus providing useful hints to guide wet-lab experiments and advance the biological knowledge discovery process.