Accelerating Curation of GWAS Catalog by Automatic Text Mining
Accelerating Curation of GWAS Catalog by Automatic Text Mining
批准号:
9052542
负责人:
Chunnan Hsu
金额:
$13.1万
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-01 至 2016-09-30
关键词:
AlgorithmsAreaBiologicalCatalogingCatalogsCommunitiesComputer softwareDataDatabasesDevelopmentDiagnosisDiseaseEnsureFutureGenesGeneticGenetic VariationGenomeGoalsHuman GeneticsInvestigationLeadLinkMachine LearningManualsMethodologyMonitorNational Human Genome Research InstitutePaperProductivityPublic DomainsPublicationsPublishingReportingResearchResearch DesignResearch PersonnelResourcesSamplingScanningSiteStandardizationSurveysTechnologyTextTrainingUpdateWorkloadbasedesigngenetic associationgenome wide association studyimprovedinterestnext generation sequencingnovelpreventpublic health relevancesoftware developmentsuccesstext searchingtooltraituser-friendlyvolunteer
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): A genome-wide association study (GWAS) is an approach to detecting genetic variations associated with particular diseases or traits by scanning markers across the genomes of a large-scale sample of subjects in a high-throughput manner. In less than a decade, GWAS studies have been successfully producing discovery and replication of many new disease loci. Discovered genetic associations have led to development of better strategies to diagnose, treat and prevent diseases. The number of GWAS is growing rapidly. There is a need for a database that allows researchers to easily query and search for previous results. A well-curated database also provides a resource for overview and summarization investigations of associated genetic sites and may help suggest pleiotropic genes. Such a database has been created and maintained by the National Human Genome Research Institute (NHGRI), called "A Catalog of Published Genome-Wide Association Studies" (Catalog of GWAS). The catalog has led to interesting characterization of previous results in GWAS and NHGRI has continued to update and curate the catalog regularly. However, this is performed by manually extracting information from published GWAS articles. As a result, the coverage is low compared to the volume of all GWAS publications and would be impossible to catch up the pace of new publications. The goal of this project is to develop a new tool to automatically extract the information from research articles for the curation of the catalog of GWAS. Our proposal is to use the curated data currently available from NHGRI as the training examples and apply novel machine-learning algorithms to train an information extractor to allow accurate automatic extraction. Given our recent success in applying machine learning to biological text mining, we are confident that this will lead to a useful tool to improve the productivity of curators and solve the coverage problem. Our first specific aim is to develop an accurate information extractor. Our second specific aim is to develop an easy-to-use curation tool for curators to efficiently check and correct errors from automatic information extraction so that their curation productivity can be improved by 18 folds. Then we will adapt the tool to extraction and curation of research papers reporting association studies using data from next generation sequencing. Currently, study design and the reporting of GWAS results using NGS data are not standardized. These results have not been considered to be included in the catalog yet. However, we expect that the limitations will be overcome and the methodology will converge soon. We will closely monitor the progress and adapt the tool to allow for inclusion of the NGS data. Finally, we will distribute the software to the public domain so that volunteers or interested parties can create their own catalog locally. It is our goal to share the developed software with the research community to advance the field. The new algorithms developed in this project and the entire development cycle, from design to deployment, will also contribute to the state-of-the- arts of biological text mining.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1186/s12859-015-0844-1
发表时间:
2016-01-11
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Jain S, Tumkur KR, Kuo TT, Bhargava S, Lin G, Hsu CN]
通讯作者:
Hsu CN
Accelerating Curation of GWAS Catalog by Automatic Text Mining
-
批准号:8348769
-
项目类别:
-
资助金额:$24.79万
-
财政年份:2012
-
负责人:Chunnan Hsu
-
依托单位:
Accelerating Curation of GWAS Catalog by Automatic Text Mining
-
批准号:8549925
-
项目类别:
-
资助金额:$5.42万
-
财政年份:2012
-
负责人:Chunnan Hsu
-
依托单位:
Accelerating Curation of GWAS Catalog by Automatic Text Mining
-
批准号:8791780
-
项目类别:
-
资助金额:$21.74万
-
财政年份:2012
-
负责人:Chunnan Hsu
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: