Improved Chemical Text Mining of Patents with Infinite Dictionaries and Automatic Spelling Correction

Improved Chemical Text Mining of Patents with Infinite Dictionaries and Automatic Spelling Correction
复制标题

DOI:
10.1021/ci200463r
复制
发表时间:
2012-01-01
影响因子:
5.6
通讯作者:
Muresan, Sorel
Muresan, Sorel
中科院分区:
化学2区
文献类型:
--
作者:
Sayle, Roger;Xie, Paul Hongxing;Muresan, Sorel

文献摘要

被引文献

相似文献

制药领域专利的文本挖掘提出了许多其他文本挖掘领域未遇到的独特挑战。与生物信息学等领域不同,在生物信息学中,感兴趣的术语数量是可枚举的,并且本质上是静态的,系统的化学命名法可以描述无限数量的分子。因此,通常用于基因名称、疾病、物种等的基于词典和本体的技术在专利中搜索新型治疗化合物时效用有限。此外,类似 IUPAC 名称的长度和构成使它们更容易受到印刷问题的影响:OCR 失败、人为拼写错误以及连字和断行问题。这项工作描述了一种名为 CaffeineFix 的新技术,旨在有效识别自由文本中的化学名称,即使存在印刷错误。生成更正的化学名称作为名称到结构软件的输入。这形成了一个预处理过程,独立于所使用的名称到结构软件,并且在我们的研究中被证明可以极大地改善化学文本挖掘的结果。
The text mining of patents of pharmaceutical interest poses a number of unique challenges not encountered in other fields of text mining. Unlike fields, such as bioinformatics, where the number of terms of interest is enumerable, and essentially static, systematic chemical nomenclature can describe an infinite number of molecules. Hence, the dictionary- and ontology-based techniques that are commonly used for gene names, diseases, species, etc., have limited utility when searching for novel therapeutic compounds in patents. Additionally, the length and the composition of IUPAC-like names make them more susceptible to typographic problems: OCR failures, human spelling errors, and hyphenation and line breaking issues. This work describes a novel technique, called CaffeineFix, designed to efficiently identify chemical names in free text, even in the presence of typographical errors. Corrected chemical names are generated as input for name-to-structure software. This forms a preprocessing pass, independent of the name-to-structure software used, and is shown to greatly improve the results of chemical text mining in our study.