课题基金 / 基金详情

Accelerating Curation of GWAS Catalog by Automatic Text Mining

Accelerating Curation of GWAS Catalog by Automatic Text Mining
通过自动文本挖掘加速 GWAS 目录的管理
批准号:
8791780
负责人:
Chunnan Hsu
金额:
$21.74万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-09-24 至 2015-06-30

项目摘要

项目成果

Chunnan Hsu的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):全基因组关联研究(Gwas)是一种通过高通量扫描大规模样本基因组的标记来检测与特定疾病或特征相关的遗传变异的方法。在不到十年的时间里,GWAS的研究已经成功地发现并复制了许多新的疾病位点。已发现的遗传关联导致了诊断、治疗和预防疾病的更好策略的发展。全球气候变化网络的数量正在迅速增长。需要一种允许研究人员容易地查询和搜索先前结果的数据库。一个精心策划的数据库也为相关遗传位点的概述和总结研究提供了资源,并可能有助于提出多效性基因。这样的数据库是由美国国家人类基因组研究所(NHGRI)创建和维护的,名为“已发表的全基因组联合研究目录”(Catalog Of Gwas)。该目录导致了GWAS对先前结果的有趣描述,NHGRI继续定期更新和整理该目录。然而,这是通过从已发表的GWAS文章中手动提取信息来执行的。因此,与全球气候变化网络所有出版物的数量相比,覆盖率很低,不可能跟上新出版物的速度。该项目的目标是开发一种新的工具,自动从研究文章中提取信息,以供管理全球气候变化网络的目录之用。我们的建议是使用目前可从NHGRI获得的精选数据作为训练样本,并应用新的机器学习算法来训练信息抽取器,以实现准确的自动抽取。鉴于我们最近在将机器学习应用于生物文本挖掘方面取得的成功,我们相信这将导致一个有用的工具,以提高馆长的生产力并解决覆盖面问题。我们的第一个具体目标是开发一个准确的信息萃取器。我们的第二个具体目标是开发一种简单易用的馆藏工具,让馆长能够有效地检查和纠正自动信息提取中的错误,从而使他们的馆藏生产率提高18倍。然后,我们将使该工具适用于使用来自下一代测序的数据来提取和整理报告关联研究的研究论文。目前,研究设计和使用NGS数据报告GWAS结果没有标准化。这些结果还没有被认为包括在目录中。然而,我们预计这些限制将被克服,方法将很快趋同。我们将密切监测进展情况,并对工具进行调整,以便纳入NGS数据。最后,我们将把软件分发到公共领域,以便志愿者或感兴趣的各方可以在本地创建自己的目录。我们的目标是与研究社区共享开发的软件,以推动该领域的发展。在这个项目中开发的新算法和整个开发周期,从设计到部署,也将有助于生物文本挖掘的最新技术。
英文摘要
DESCRIPTION (provided by applicant): A genome-wide association study (GWAS) is an approach to detecting genetic variations associated with particular diseases or traits by scanning markers across the genomes of a large-scale sample of subjects in a high-throughput manner. In less than a decade, GWAS studies have been successfully producing discovery and replication of many new disease loci. Discovered genetic associations have led to development of better strategies to diagnose, treat and prevent diseases. The number of GWAS is growing rapidly. There is a need for a database that allows researchers to easily query and search for previous results. A well-curated database also provides a resource for overview and summarization investigations of associated genetic sites and may help suggest pleiotropic genes. Such a database has been created and maintained by the National Human Genome Research Institute (NHGRI), called "A Catalog of Published Genome-Wide Association Studies" (Catalog of GWAS). The catalog has led to interesting characterization of previous results in GWAS and NHGRI has continued to update and curate the catalog regularly. However, this is performed by manually extracting information from published GWAS articles. As a result, the coverage is low compared to the volume of all GWAS publications and would be impossible to catch up the pace of new publications. The goal of this project is to develop a new tool to automatically extract the information from research articles for the curation of the catalog of GWAS. Our proposal is to use the curated data currently available from NHGRI as the training examples and apply novel machine-learning algorithms to train an information extractor to allow accurate automatic extraction. Given our recent success in applying machine learning to biological text mining, we are confident that this will lead to a useful tool to improve the productivity of curators and solve the coverage problem. Our first specific aim is to develop an accurate information extractor. Our second specific aim is to develop an easy-to-use curation tool for curators to efficiently check and correct errors from automatic information extraction so that their curation productivity can be improved by 18 folds. Then we will adapt the tool to extraction and curation of research papers reporting association studies using data from next generation sequencing. Currently, study design and the reporting of GWAS results using NGS data are not standardized. These results have not been considered to be included in the catalog yet. However, we expect that the limitations will be overcome and the methodology will converge soon. We will closely monitor the progress and adapt the tool to allow for inclusion of the NGS data. Finally, we will distribute the software to the public domain so that volunteers or interested parties can create their own catalog locally. It is our goal to share the developed software with the research community to advance the field. The new algorithms developed in this project and the entire development cycle, from design to deployment, will also contribute to the state-of-the- arts of biological text mining.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Accelerating Curation of GWAS Catalog by Automatic Text Mining
Accelerating Curation of GWAS Catalog by Automatic Text Mining
Accelerating Curation of GWAS Catalog by Automatic Text Mining
海外基金