课题基金 / 基金详情

Transforming dbGaP genetic and genomic data to FAIR-ready by artificial intelligence and machine learning algorithms

Transforming dbGaP genetic and genomic data to FAIR-ready by artificial intelligence and machine learning algorithms
通过人工智能和机器学习算法将 dbGaP 遗传和基因组数据转变为 FAIR-ready
批准号:
10842954
负责人:
Zhongming Zhao
金额:
$30.61万
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
未结题
起止时间:
2017-09-14 至 2025-05-31
关键词:

项目摘要

项目成果

Zhongming Zhao的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
dbGaP is a repository for NIH funded projects and it contains many genetic and genomic data. However, data there are not ready for AI and machine learning applications. This application proposes methods to address this issue. We have two aims: 1). Develop and standardize procedures to transfer genetic and genomic data into image like objects and tokenized custom vocabulary so that the data can be utilized by advanced AI algorithms such CNN, autoencoder and transformer. To transform genetic data into image, we recode allele dosage value as pixel intensity and arrange a collection of genetic markers such as SNPs and CNVs into an artificial image object so that it can be analyzed by CNN algorithms. Genetic markers can also be used to define haplotypes, which can be tokenized into custom vocabularies for use in NLP models. 2). Use Alzheimer's disease and schizophrenia as case studies to demonstrate the utilities of transformed data for the discovery and identification of risk variants/genes for both conditions. We plan to impute genetically controlled gene expression using brain specific eQTLs and individual genotypes for an AD dataset, and transform the expression data into image objects for analyses by CNN model with self attention mechanism. For schizophrenia, we plan to use k- mer tokenizer to break haplotypes into a collection of small haplotype blocks and treat them as tokens for analyses by NLP models. We use both CNN and NLP models as screen tools to select promising candidates using the attention weights, and then directly test these candidates for their association with AD/schizophrenia using logistic regression. Due to the selection effect, we can dramatically reduce the number of testing, significantly increase our statistical power to detect risk variants/genes to AD/schizophrenia.
期刊论文(42)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1002/advs.202004958
发表时间: 2021-05
期刊: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
影响因子: --
作者: [Xu H, Jia P, Zhao Z]
通讯作者: Zhao Z
DOI: 10.1093/nar/gkx850
发表时间: 2018-01-04
期刊: Nucleic acids research
影响因子: 14.9
作者: [Kim P, Park A, Han G, Sun H, Jia P, Zhao Z]
通讯作者: Zhao Z
DOI: 10.1371/journal.pone.0291631
发表时间: 2023
期刊: PLOS ONE
影响因子: 3.7
作者: [Fathima, Afifa Salsabil, Basha, Syed Muzamil, Ahmed, Syed Thouheed, Mathivanan, Sandeep Kumar, Rajendran, Sukumar, Mallik, Saurav, Zhao, Zhongming]
通讯作者: Zhao, Zhongming
DOI: 10.1093/nar/gkaa945
发表时间: 2021-01-08
期刊: Nucleic acids research
影响因子: 14.9
作者: [Hu R, Xu H, Jia P, Zhao Z]
通讯作者: Zhao Z
21
    Constructing A Transcriptomic Atlas of Retrotransposon in Alzheimer's Disease
    Deep learning methods to predict the function of genetic variants in orofacial clefts
    Predicting Phenotype by Deep Learning Heterogeneous Multi-Omics Data
    Predicting Phenotype by Deep Learning Heterogeneous Multi-Omics Data
    海外基金