HypertenGene: extracting key hypertension genes from biomedical literature with position and automatically-generated template features.

HypertenGene: extracting key hypertension genes from biomedical literature with position and automatically-generated template features.
复制标题

DOI:
10.1186/1471-2105-10-s15-s9
复制
发表时间:
2009-12-03
期刊:
影响因子:
3
通讯作者:
Hsu WL
Hsu WL
中科院分区:
生物学4区
文献类型:
--
作者:
Tsai RT;Lai PT;Dai HJ;Huang CH;Bow YY;Chang YC;Pan WH;Hsu WL

文献摘要

被引文献

相似文献

导致高血压的遗传因素已被广泛研究,并发表了大量关于该主题的研究论文。高血压相关基因的定位是高血压研究者的主要研究任务之一。然而,用现有的工具收集这些信息并不容易:(1)搜索文章通常会返回太多的点击量,无法浏览。(2)搜索结果没有突出显示摘要中发现的高血压相关基因。(3)尽管一些文本挖掘服务在摘要中标记了基因名称,但论文中研究的关键基因仍然无法与其他基因区分开来。为了方便高血压研究人员的信息收集过程,一个解决方案是在每个摘要中提取关键的高血压相关基因。该系统的构建主要包括三个方面的工作:(1)基因和高血压命名实体识别,(2)片段分类,(3)基因与高血压关系提取。我们首先比较检索性能所取得的分别添加模板功能和位置功能的基线系统。然后,研究两者的结合。我们发现,使用位置特征几乎可以使基线系统的原始AUC分数(0.8140vs.0.4936)加倍。然而,添加模板功能只会导致边际改善(0.0197)。包括两者将AUC提高到0.8184,表明这两组特征是互补的,并且没有重叠的效果。然后,我们检查了在不同领域-糖尿病中的性能,结果显示令人满意的AUC为0.83。我们的方法成功地利用模板特征来识别真正的高血压相关基因提及和位置特征,以区分关键基因从其他相关基因。模板由生物学家自动生成和检查,以最大限度地减少劳动力成本。我们的方法集成了机器学习模型和模式匹配的优势。据我们所知,这是第一次对高血压相关基因的提取进行系统的研究,也是第一次尝试建立基于GAD数据库的高血压基因关系语料库。此外,我们的论文提出并测试了用于提取高血压关键基因的新特征,如相对位置,部分和模板特征,这些特征也可以应用于其他疾病的关键基因提取。
The genetic factors leading to hypertension have been extensively studied, and large numbers of research papers have been published on the subject. One of hypertension researchers' primary research tasks is to locate key hypertension-related genes in abstracts. However, gathering such information with existing tools is not easy: (1) Searching for articles often returns far too many hits to browse through. (2) The search results do not highlight the hypertension-related genes discovered in the abstract. (3) Even though some text mining services mark up gene names in the abstract, the key genes investigated in a paper are still not distinguished from other genes. To facilitate the information gathering process for hypertension researchers, one solution would be to extract the key hypertension-related genes in each abstract. Three major tasks are involved in the construction of this system: (1) gene and hypertension named entity recognition, (2) section categorization, and (3) gene-hypertension relation extraction. We first compare the retrieval performance achieved by individually adding template features and position features to the baseline system. Then, the combination of both is examined. We found that using position features can almost double the original AUC score (0.8140vs.0.4936) of the baseline system. However, adding template features only results in marginal improvement (0.0197). Including both improves AUC to 0.8184, indicating that these two sets of features are complementary, and do not have overlapping effects. We then examine the performance in a different domain--diabetes, and the result shows a satisfactory AUC of 0.83. Our approach successfully exploits template features to recognize true hypertension-related gene mentions and position features to distinguish key genes from other related genes. Templates are automatically generated and checked by biologists to minimize labor costs. Our approach integrates the advantages of machine learning models and pattern matching. To the best of our knowledge, this the first systematic study of extracting hypertension-related genes and the first attempt to create a hypertension-gene relation corpus based on the GAD database. Furthermore, our paper proposes and tests novel features for extracting key hypertension genes, such as relative position, section, and template features, which could also be applied to key-gene extraction for other diseases.