Automated recognition of malignancy mentions in biomedical literature.

Automated recognition of malignancy mentions in biomedical literature.
复制标题

DOI:
10.1186/1471-2105-7-492
复制
发表时间:
2006-11-07
期刊:
影响因子:
3
通讯作者:
White PS
White PS
中科院分区:
生物学4区
文献类型:
--
作者:
Jin Y;McDonald RT;Lerman K;Mandel MA;Carroll S;Liberman MY;Pereira FC;Winters RS;White PS

文献摘要

参考文献

被引文献

相似文献

生物医学文本的快速扩散使得研究人员越来越难以识别,综合和利用他们感兴趣的领域的发达知识。自动信息提取程序可以帮助获取和管理这些知识。生物医学文本挖掘中以前的努力主要集中在命名实体识别定义明确的分子对象,如基因,但较少的工作已经执行,以确定疾病相关的对象和概念。此外,由于无法以最小化手动工作并仍然以高准确性执行的方式有效地扩展方法,因此前景受到了影响。在这里,我们应用了一种机器学习方法,这种方法以前成功地将分子实体识别到疾病概念中,以确定潜在的概率模型是否有效地推广到不相关的概念,并以最小的手动干预进行模型再训练。我们开发了一个命名实体识别器(MTag),一个实体标签识别恶性肿瘤的临床描述的文本。该应用程序使用机器学习技术条件随机场与其他特定领域的功能。MTag测试了1,010份与癌症基因组学相关的培训和432份评估文件。总的来说,我们的实验在评估集上获得了0.85的精确度,0.83的召回率和0.84的F-测量。与使用文本与肿瘤术语列表的字符串匹配的基线系统相比,MTag的召回率要高得多(92.1%对42.1%),并表现出学习新模式的能力。将MTag应用于所有MEDLINE摘要,识别出580,002个独特的恶性肿瘤和9,153,340个恶性肿瘤的总体提及。值得注意的是,添加广泛的恶性肿瘤词汇作为提取特征集对性能的影响最小。总之,这些结果表明,不同的生物医学实体类在自由文本中的识别可能是可实现的,具有高精度,只有适度的额外努力,为每个新的应用领域。
The rapid proliferation of biomedical text makes it increasingly difficult for researchers to identify, synthesize, and utilize developed knowledge in their fields of interest. Automated information extraction procedures can assist in the acquisition and management of this knowledge. Previous efforts in biomedical text mining have focused primarily upon named entity recognition of well-defined molecular objects such as genes, but less work has been performed to identify disease-related objects and concepts. Furthermore, promise has been tempered by an inability to efficiently scale approaches in ways that minimize manual efforts and still perform with high accuracy. Here, we have applied a machine-learning approach previously successful for identifying molecular entities to a disease concept to determine if the underlying probabilistic model effectively generalizes to unrelated concepts with minimal manual intervention for model retraining. We developed a named entity recognizer (MTag), an entity tagger for recognizing clinical descriptions of malignancy presented in text. The application uses the machine-learning technique Conditional Random Fields with additional domain-specific features. MTag was tested with 1,010 training and 432 evaluation documents pertaining to cancer genomics. Overall, our experiments resulted in 0.85 precision, 0.83 recall, and 0.84 F-measure on the evaluation set. Compared with a baseline system using string matching of text with a neoplasm term list, MTag performed with a much higher recall rate (92.1% vs. 42.1% recall) and demonstrated the ability to learn new patterns. Application of MTag to all MEDLINE abstracts yielded the identification of 580,002 unique and 9,153,340 overall mentions of malignancy. Significantly, addition of an extensive lexicon of malignancy mentions as a feature set for extraction had minimal impact in performance. Together, these results suggest that the identification of disparate biomedical entity classes in free text may be achievable with high accuracy and only moderate additional effort for each new application domain.
DOI: 10.1186/1471-2407-4-88
发表时间: 2004-11-30
期刊: BMC cancer
影响因子: 3.8
作者:
Berman JJ
通讯作者: Berman JJ
DOI: 10.1186/1471-2105-6-s1-s10
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Tamames J
通讯作者: Tamames J
DOI: 10.1016/j.jbi.2004.08.008
发表时间: 2004-12-01
影响因子: 4.5
作者:
Collier, N;Takeuchi, K
通讯作者: Takeuchi, K
DOI: 10.1016/s1386-5056(02)00053-9
发表时间: 2002-12-04
影响因子: 4.9
作者:
Hahn, U;Romacker, M;Schulz, S
通讯作者: Schulz, S
DOI: 10.1038/sj.ejhg.5201585
发表时间: 2006-05-01
影响因子: 5.2
作者:
van Driel, MA;Bruggeman, J;Leunissen, JA
通讯作者: Leunissen, JA