Identifying and classifying goals for scientific knowledge.

Identifying and classifying goals for scientific knowledge.
复制标题

DOI:
10.1093/bioadv/vbab012
复制
发表时间:
2021
期刊:
Bioinformatics advances
影响因子:
--
通讯作者:
Hunter LE
Hunter LE
中科院分区:
其他
文献类型:
--
作者:
Boguslav MR;Salem NM;White EK;Leach SM;Hunter LE

文献摘要

被引文献

相似文献

科学通过提出好的问题而进步,但生物医学文本挖掘的工作并没有太多关注它们。我们提出了一个新的想法,生物医学自然语言处理:识别和描述的问题,在生物医学文献。从形式上讲,这项任务是确定和描述无知的陈述,即科学知识缺失或不完整的陈述。此类技术的创建可能会产生许多重大影响,从博士生的培训到对出版物进行排名以及根据特定感兴趣的问题优先考虑资金。这里介绍的工作是为了实现这些目标的第一步。我们提出了一种新的无知分类法驱动的无知在研究中发挥的作用声明,确定未来的科学知识的具体目标。使用这种分类法和可靠的注释指南(注释者之间的一致性高于80%),我们创建了一个黄金标准的无知语料库,其中包含来自产前营养文献的60个全文文档,具有超过10000个注释,并使用它来训练分类器,达到超过0.80 F1分数。语料库和源代码可在https://github.com/UCDenver-ccp/Ignorance-Question-Work免费下载。源代码用Python实现。
Science progresses by posing good questions, yet work in biomedical text mining has not focused on them much. We propose a novel idea for biomedical natural language processing: identifying and characterizing the questions stated in the biomedical literature. Formally, the task is to identify and characterize statements of ignorance, statements where scientific knowledge is missing or incomplete. The creation of such technology could have many significant impacts, from the training of PhD students to ranking publications and prioritizing funding based on particular questions of interest. The work presented here is intended as the first step towards these goals. We present a novel ignorance taxonomy driven by the role statements of ignorance play in research, identifying specific goals for future scientific knowledge. Using this taxonomy and reliable annotation guidelines (inter-annotator agreement above 80%), we created a gold standard ignorance corpus of 60 full-text documents from the prenatal nutrition literature with over 10 000 annotations and used it to train classifiers that achieved over 0.80 F1 scores. Corpus and source code freely available for download at https://github.com/UCDenver-ccp/Ignorance-Question-Work. The source code is implemented in Python.