课题基金 / 基金详情

EAPSI: Investigating a Novel Approach for Biological Named Entity Recognition in Text Mining

EAPSI: Investigating a Novel Approach for Biological Named Entity Recognition in Text Mining
EAPSI:研究文本挖掘中生物命名实体识别的新方法
批准号:
1614261
负责人:
Dally Shvets
金额:
$0.54万
依托单位:
依托单位国家:
美国
项目类别:
Fellowship Award
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-15 至 2017-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
目前,由于生物文本的特定领域术语,使用计算资源从生物文献中提取信息面临着许多挑战。为了从海量的现有生物文献中阅读和获取信息,需要大量的资金、时间和人力。研究文献和数据的数量一直在增加,从生物学文本中收集信息的可扩展解决方案是必要的。该项目旨在开发一种从生物文献中提取特定注释的新技术。在这个项目中,PI将前往台湾台北的中央研究院与徐文连博士合作,他在自然语言处理和生物文献文本挖掘领域的专业知识是实施这一方法所必需的。DNA甲基化被视为癌症诊断和治疗的潜在生物标记物。最近的研究已经确定了基因甲基化异常与癌症发生之间的关系。在以前的生物文献挖掘方法中,人工精选和基于机器学习的方法相结合被用于提取基因甲基化与癌症的关系。这个项目将利用一种新的基于统计模式的方法,重点是确定为疾病术语和身体互动所描述的深层和重要的语言模式。这种方法具有独特的灵活匹配机制,可以捕获文本中现有的疾病术语或蛋白质名称,而不考虑它们以前在训练数据中的存在。这个项目解决了在短时间内阅读和分析大量生物文献的资源限制。东亚和太平洋暑期学院计划下的这个奖项支持一名美国研究生的暑期研究,由NSF和台湾科技部共同资助。
英文摘要
Currently, using computational resources to extract information from biological literature faces many challenges due to the domain specific terminology of biological texts. In order to read and obtain information from the vast amount of available biological literature, a large amount of funds, time and manpower are required. The amount of research literature and data is always increasing, and a scalable solution to gleaning information from biological text is necessary. This project aims to develop a new technique for extracting specific annotations from biological literature. For this project, the PI will travel to Academia Sinica in Taipei, Taiwan to work with Dr. WenLian Hsu, whose expertise in the areas of natural language processing and text mining of biological literature is necessary for the implementation of this approach.DNA methylation is regarded as a potential biomarker in the diagnosis and treatment of cancer. Recent studies have identified relationships between aberrant gene methylation and cancer development. In previous approaches of mining biological literature, a combination of manual curation and machine learning-based approaches have been used to extract gene methylation cancer relations. This project will utilize a novel statistical pattern based approach which focuses on identifying deep and important linguistic patterns described for disease terms and physical interactions. This approach has a unique flexible matching mechanism that captures existing disease terms or protein names within the text, regardless of their previous presence in the training data. This project provides a solution to the resource limitation of reading and analyzing large amounts of biological literature in a short amount of time.This award under the East Asia and Pacific Summer Institutes program supports summer research by a U.S. graduate student and is jointly funded by NSF and the Ministry of Science and Technology of Taiwan.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金