Text mining meets community curation: a newly designed curation platform to improve author experience and participation at WormBase

Text mining meets community curation: a newly designed curation platform to improve author experience and participation at WormBase
复制标题

DOI:
10.1093/database/baaa006
复制
发表时间:
2020-03-17
影响因子:
5.8
通讯作者:
Sternberg, Paul W.
Sternberg, Paul W.
中科院分区:
生物学4区
文献类型:
--
作者:
Arnaboldi, Valerio;Raciti, Daniela;Sternberg, Paul W.

文献摘要

被引文献

相似文献

生物学知识库依赖于研究文献的专家生物化,以保持以机器可读形式组织的最新数据集。要将信息输入知识库,策展人需要遵循三个步骤:(i)识别包含相关数据的论文,这一过程称为分类;(ii)识别命名实体;(iii)根据底层数据模型提取和策展数据。WormBase(WB)是秀丽隐杆线虫和其他线虫研究数据的权威存储库,它使用文本挖掘(TM)来半自动化其管理管道。此外,WB通过作者第一次通过(AFP)系统与其社区合作,帮助识别实体并对最近发表的论文中的数据类型进行分类。在本文中,我们提出了一个新的WB AFP系统,结合TM和AFP到一个单一的应用程序,以提高社区策展。该系统采用字符串搜索算法和统计方法(例如支持向量机(SVM))来提取生物实体并对数据类型进行分类,并将结果以网络形式呈现给作者,在那里他们验证提取的信息,而不是像以前那样重新输入。有了这个新系统,我们减轻了作者的负担,同时收到了关于我们TM工具性能的宝贵反馈。新的用户界面还链接到特定的结构化数据提交表单,例如表型或表达模式数据,使作者有机会贡献更详细的策展,可以在最少的策展人审查的情况下纳入WB。我们的方法是可推广的,可以应用于其他知识库,希望参与他们的用户社区协助策展。在新系统启动后的五个月里,答复率与法新社以前的版本相当,但收到的数据的质量和数量都有很大提高。
Biological knowledgebases rely on expert biocuration of the research literature to maintain up-to-date collections of data organized in machine-readable form. To enter information into knowledgebases, curators need to follow three steps: (i) identify papers containing relevant data, a process called triaging; (ii) recognize named entities; and (iii) extract and curate data in accordance with the underlying data models. WormBase (WB), the authoritative repository for research data on Caenorhabditis elegans and other nematodes, uses text mining (TM) to semi-automate its curation pipeline. In addition, WB engages its community, via an Author First Pass (AFP) system, to help recognize entities and classify data types in their recently published papers. In this paper, we present a new WB AFP system that combines TM and AFP into a single application to enhance community curation. The system employs string-searching algorithms and statistical methods (e.g. support vector machines (SVMs)) to extract biological entities and classify data types, and it presents the results to authors in a web form where they validate the extracted information, rather than enter it de novo as the previous form required. With this new system, we lessen the burden for authors, while at the same time receive valuable feedback on the performance of our TM tools. The new user interface also links out to specific structured data submission forms, e.g. for phenotype or expression pattern data, giving the authors the opportunity to contribute a more detailed curation that can be incorporated into WB with minimal curator review. Our approach is generalizable and could be applied to additional knowledgebases that would like to engage their user community in assisting with the curation. In the five months succeeding the launch of the new system, the response rate has been comparable with that of the previous AFP version, but the quality and quantity of the data received has greatly improved.