课题基金 / 基金详情

项目摘要

项目成果

Chunnan Hsu的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):全基因组关联研究(GWAS)是一种通过高通量方式扫描大规模受试者基因组中的标记来检测与特定疾病或性状相关的遗传变异的方法。在不到十年的时间里,GWAS研究已经成功地发现和复制了许多新的疾病位点。发现的遗传关联导致了更好的诊断、治疗和预防疾病策略的发展。GWAS的数量正在迅速增长。有必要建立一个数据库,使研究人员能够方便地查询和搜索以前的结果。一个精心策划的数据库也为相关遗传位点的概述和总结调查提供了资源,并可能有助于发现多效基因。美国国家人类基因组研究所(NHGRI)已经建立并维护了这样一个数据库,名为《全基因组关联研究出版目录》(Catalog of GWAS)。该目录对GWAS以前的结果进行了有趣的描述,NHGRI继续定期更新和管理该目录。但是,这是通过手动从已发布的GWAS文章中提取信息来执行的。因此,与所有GWAS出版物的数量相比,其覆盖率较低,不可能赶上新出版物的步伐。本课题的目标是开发一种新的工具来自动地从研究文章中提取信息,用于GWAS目录的管理。我们的建议是使用NHGRI目前提供的整理数据作为训练示例,并应用新颖的机器学习算法来训练信息提取器,以实现准确的自动提取。鉴于我们最近在将机器学习应用于生物文本挖掘方面的成功,我们有信心这将导致一个有用的工具来提高策展人的生产力并解决覆盖问题。我们的第一个具体目标是开发一个准确的信息提取器。我们的第二个具体目标是开发一个易于使用的策展工具,供策展人有效地检查和纠正自动信息提取中的错误,从而使他们的策展效率提高18倍。然后,我们将调整工具,以提取和管理研究论文报告关联研究使用下一代测序的数据。目前,使用NGS数据的研究设计和GWAS结果报告尚未标准化。这些结果还没有被考虑包括在目录中。然而,我们期望这些限制将被克服,方法将很快趋于一致。我们将密切监测进展情况,并调整该工具,以便纳入国家地质勘探局的数据。最后,我们将把软件发布到公共领域,以便志愿者或感兴趣的团体可以在本地创建他们自己的目录。我们的目标是与研究界分享开发的软件,以推动该领域的发展。在这个项目中开发的新算法和整个开发周期,从设计到部署,也将有助于生物文本挖掘的最新技术。
英文摘要
DESCRIPTION (provided by applicant): A genome-wide association study (GWAS) is an approach to detecting genetic variations associated with particular diseases or traits by scanning markers across the genomes of a large-scale sample of subjects in a high-throughput manner. In less than a decade, GWAS studies have been successfully producing discovery and replication of many new disease loci. Discovered genetic associations have led to development of better strategies to diagnose, treat and prevent diseases. The number of GWAS is growing rapidly. There is a need for a database that allows researchers to easily query and search for previous results. A well-curated database also provides a resource for overview and summarization investigations of associated genetic sites and may help suggest pleiotropic genes. Such a database has been created and maintained by the National Human Genome Research Institute (NHGRI), called "A Catalog of Published Genome-Wide Association Studies" (Catalog of GWAS). The catalog has led to interesting characterization of previous results in GWAS and NHGRI has continued to update and curate the catalog regularly. However, this is performed by manually extracting information from published GWAS articles. As a result, the coverage is low compared to the volume of all GWAS publications and would be impossible to catch up the pace of new publications. The goal of this project is to develop a new tool to automatically extract the information from research articles for the curation of the catalog of GWAS. Our proposal is to use the curated data currently available from NHGRI as the training examples and apply novel machine-learning algorithms to train an information extractor to allow accurate automatic extraction. Given our recent success in applying machine learning to biological text mining, we are confident that this will lead to a useful tool to improve the productivity of curators and solve the coverage problem. Our first specific aim is to develop an accurate information extractor. Our second specific aim is to develop an easy-to-use curation tool for curators to efficiently check and correct errors from automatic information extraction so that their curation productivity can be improved by 18 folds. Then we will adapt the tool to extraction and curation of research papers reporting association studies using data from next generation sequencing. Currently, study design and the reporting of GWAS results using NGS data are not standardized. These results have not been considered to be included in the catalog yet. However, we expect that the limitations will be overcome and the methodology will converge soon. We will closely monitor the progress and adapt the tool to allow for inclusion of the NGS data. Finally, we will distribute the software to the public domain so that volunteers or interested parties can create their own catalog locally. It is our goal to share the developed software with the research community to advance the field. The new algorithms developed in this project and the entire development cycle, from design to deployment, will also contribute to the state-of-the- arts of biological text mining.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1186/s12859-015-0844-1
发表时间: 2016-01-11
期刊: BMC bioinformatics
影响因子: 3
作者: [Jain S, Tumkur KR, Kuo TT, Bhargava S, Lin G, Hsu CN]
通讯作者: Hsu CN
Accelerating Curation of GWAS Catalog by Automatic Text Mining
Accelerating Curation of GWAS Catalog by Automatic Text Mining
Accelerating Curation of GWAS Catalog by Automatic Text Mining
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: