The first step in the development of Text Mining technology for Cancer Risk Assessment: identifying and organizing scientific evidence in risk assessment literature.

The first step in the development of Text Mining technology for Cancer Risk Assessment: identifying and organizing scientific evidence in risk assessment literature.
复制标题

DOI:
10.1186/1471-2105-10-303
复制
发表时间:
2009-09-22
期刊:
影响因子:
3
通讯作者:
Stenius U
Stenius U
中科院分区:
生物学4区
文献类型:
--
作者:
Korhonen A;Silins I;Sun L;Stenius U

文献摘要

参考文献

被引文献

相似文献

生物医学文本挖掘(TM)中最被忽视的领域之一是基于仔细评估的用户需求开发系统。我们最近调查了TM尚未解决的一项重要任务——癌症风险评估(CRA)的用户需求。在这里,我们向TM技术的发展迈出了第一步:在一个能够支持从生物医学文献中广泛收集数据的分类法中识别和组织CRA所需的科学证据。该分类基于专家对从相关PubMed期刊下载的1297篇摘要的注释。它将语料库中发现的1742个唯一关键字分类为48个类别,这些类别指定了CRA所需的核心证据。我们报告了注释者间协议测试和PubMed摘要自动分类到分类类的有希望的结果。在接近真实世界的CRA场景中还报告了一个简单的用户测试,该测试与其他评估一起演示了我们构建的资源在实践中是定义良好的、准确的和可应用的。我们提出了我们的注释准则和一个我们设计的PubMed摘要专家注释工具。本文还提供了一个对关键字和文档相关性进行标注的语料库,以及将关键字组织成定义CRA核心证据的类的分类法。评价结果表明,我们构建的材料为CRA文献的多维分类提供了良好的基础。它们可以支持当前的手动CRA,也可以促进基于TM的方法的开发。我们讨论了通过人工和机器学习方法进一步扩展分类法,以及开发满足CRA需求的TM技术所需的后续步骤。
One of the most neglected areas of biomedical Text Mining (TM) is the development of systems based on carefully assessed user needs. We have recently investigated the user needs of an important task yet to be tackled by TM -- Cancer Risk Assessment (CRA). Here we take the first step towards the development of TM technology for the task: identifying and organizing the scientific evidence required for CRA in a taxonomy which is capable of supporting extensive data gathering from biomedical literature. The taxonomy is based on expert annotation of 1297 abstracts downloaded from relevant PubMed journals. It classifies 1742 unique keywords found in the corpus to 48 classes which specify core evidence required for CRA. We report promising results with inter-annotator agreement tests and automatic classification of PubMed abstracts to taxonomy classes. A simple user test is also reported in a near real-world CRA scenario which demonstrates along with other evaluation that the resources we have built are well-defined, accurate, and applicable in practice. We present our annotation guidelines and a tool which we have designed for expert annotation of PubMed abstracts. A corpus annotated for keywords and document relevance is also presented, along with the taxonomy which organizes the keywords into classes defining core evidence for CRA. As demonstrated by the evaluation, the materials we have constructed provide a good basis for classification of CRA literature along multiple dimensions. They can support current manual CRA as well as facilitate the development of an approach based on TM. We discuss extending the taxonomy further via manual and machine learning approaches and the subsequent steps required to develop TM technology for the needs of CRA.
DOI: 10.1109/tnn.1997.641482
发表时间: 1997-01-01
影响因子: --
作者:
Cherkassky, V
通讯作者: Cherkassky, V
DOI: 10.1093/bioinformatics/btl350
发表时间: 2006-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Han, Bo;Obradovic, Zoran;Vucetic, Slobodan
通讯作者: Vucetic, Slobodan
DOI: 10.1177/001316446002000104
发表时间: 1960-01-01
影响因子: 2.7
作者:
COHEN, J
通讯作者: COHEN, J
DOI: 10.1186/1471-2105-9-193
发表时间: 2008-04-14
期刊: BMC bioinformatics
影响因子: 3
作者:
Karamanis N;Seal R;Lewin I;McQuilton P;Vlachos A;Gasperin C;Drysdale R;Briscoe T
通讯作者: Briscoe T
DOI: 10.1197/jamia.m2401
发表时间: 2008-01-01
影响因子: 6.4
作者:
Chen, Elizabeth S.;Hripcsak, George;Friedman, Carol
通讯作者: Friedman, Carol