Fast and accurate semantic annotation of bioassays exploiting a hybrid of machine learning and user confirmation

Fast and accurate semantic annotation of bioassays exploiting a hybrid of machine learning and user confirmation
复制标题

DOI:
10.7717/peerj.524
复制
发表时间:
2014-08-14
期刊:
影响因子:
2.7
通讯作者:
Visser, Ubbo
Visser, Ubbo
中科院分区:
生物学3区
文献类型:
--
作者:
Clark, Alex M.;Bunin, Barry A.;Visser, Ubbo

文献摘要

被引文献

相似文献

生物信息学和计算机辅助药物设计依赖于大量用于生物测定的方案的管理,所述生物测定测量潜在药物实现治疗效果的能力。这些分析方案通常由科学家以纯文本的形式发布,需要更精确地注释以便对软件方法有用。我们已经开发了一种实用的方法来描述分析根据BioAssay本体(BAO)项目的语义定义,使用基于自然语言处理的机器学习的混合,以及旨在帮助科学家以最小的努力管理他们的数据的简化用户界面。我们开展这项工作的前提是,纯机器学习不够准确,并且期望科学家找到时间手动注释他们的协议是不现实的。通过结合这些方法,我们已经创建了一个有效的原型,可以非常快速地完成训练集域内的生物测定文本的注释。经过良好训练的注释需要单击用户批准,而来自训练集域之外的注释可以使用设计良好的用户界面的搜索功能来识别,并随后用于改进底层模型。通过大幅减少科学家注释其分析所需的时间,我们可以切实倡导语义注释成为出版过程的标准部分。一旦标记了公共生物测定数据的一小部分,生物信息学研究人员就可以开始构建复杂而有用的搜索和分析算法,为药物发现研究人员提供一套多样而强大的工具。
Bioinformatics and computer aided drug design rely on the curation of a large number of protocols for biological assays that measure the ability of potential drugs to achieve a therapeutic effect. These assay protocols are generally published by scientists in the form of plain text, which needs to be more precisely annotated in order to be useful to software methods. We have developed a pragmatic approach to describing assays according to the semantic definitions of the BioAssay Ontology (BAO) project, using a hybrid of machine learning based on natural language processing, and a simplified user interface designed to help scientists curate their data with minimum effort. We have carried out this work based on the premise that pure machine learning is insufficiently accurate, and that expecting scientists to find the time to annotate their protocols manually is unrealistic. By combining these approaches, we have created an effective prototype for which annotation of bioassay text within the domain of the training set can be accomplished very quickly. Well-trained annotations require single-click user approval, while annotations from outside the training set domain can be identified using the search feature of a well-designed user interface, and subsequently used to improve the underlying models. By drastically reducing the time required for scientists to annotate their assays, we can realistically advocate for semantic annotation to become a standard part of the publication process. Once even a small proportion of the public body of bioassay data is marked up, bioinformatics researchers can begin to construct sophisticated and useful searching and analysis algorithms that will provide a diverse and powerful set of tools for drug discovery researchers.