Developing novel deep-learning based methods for deciphering non-coding gene regulatory code
Developing novel deep-learning based methods for deciphering non-coding gene regulatory code
批准号:
10615784
负责人:
RAMANA V DAVULURI
金额:
$33.08万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-08-01 至 2025-04-30
关键词:
AddressBenchmarkingBindingBioinformaticsBiological AssayBipolar DisorderCRISPR/Cas technologyChIP-seqClinVarCodeCommunitiesComplexComputer Vision SystemsConsumptionDNADNA SequenceDNA Sequence AnalysisDataData SetDatabasesDevelopmentDiseaseDistantFamily memberFutureGene Expression RegulationGenesGeneticGenetic CodeGenetic DatabasesGenomeGoalsHumanHuman Cell LineHuman GenomeLabelLanguageLuciferasesMalignant NeoplasmsMethodsModelingMusNatural Language ProcessingNeural Network SimulationOrganismParkinson DiseasePerformanceProtein IsoformsProteinsPublic HealthRNA SplicingRegulator GenesRegulatory ElementReporterResearchResearch PersonnelResourcesSchizophreniaScientistSemanticsSiteSource CodeSpecificityTechniquesTimeTissuesTrainingTranslatingUnited States National Library of MedicineUntranslated RNAVariantVisualizationautism spectrum disordercandidate validationcell typedatabase of Genotypes and PhenotypesdbSNPdeep learningdeep neural networkgenetic variantgenome editinghuman DNAinsightlearning strategyneuropsychiatric disordernovelprediction algorithmpredictive modelingpromoterpublic health relevancetooltranscriptometransfer learningweb server
中文摘要
点击翻译按钮获取中文摘要
英文摘要
SUMMARY
This project will contribute novel pre-trained DNA Bidirectional Encoder Representations from Transformers,
called DNABERT, and associated deep-learning tools to decipher the language of non-coding DNA and facilitate
integration of gene regulatory information from rapidly accumulating sequence data with NLM’s genetic
databases (for example, dbSNP, dbGaP and ClinVar), which serve both scientists and the public health by
helping identify the genetic components of disease. While the genetic code explaining how DNA is translated
into proteins is universal, the regulatory code that determines when and how the genes are expressed varies
across different cell-types and organisms. Non-coding DNA is highly complex due to the existence of polysemy
and distant semantic relationship, from a language modeling perspective. Recently, deep learning methods have
been used in unraveling the gene regulatory code, but failed to globally and robustly model such language
features in the genome, especially in data-scarce scenarios. To address this challenge, we propose DNABERT
to model DNA as a language, by adapting the idea of Bidirectional Encoder Representations from Transformers
(BERT). Based on recent observations in natural language processing research, we hypothesize that pre-trained
transformer-based neural network model offer a promising, and yet not fully explored, deep learning approach
for a variety of sequence prediction tasks in the analysis of non-coding DNA. Our preliminary results showed
that DNABERT on the human genome achieved state-of-the-art performance on promoter and splice-site
prediction tasks, after easy fine-tuning on small task-specific data (Ji, Y. et al. 2020). The goal of our proposed
research is to develop DNABERT for a variety of sequence prediction tasks, and benchmark with existing state-
of-the-art deep-learning based methods. Specific aims are (1) develop novel deep-learning methods by adapting
BERT; (2) apply the proposed deep-learning methods to specifically target non-coding DNA sequence analyses
and predictions; and (3) predict and validate functional non-coding genetic variants by applying DNABERT
prediction models. A major contribution of the proposed research is development of pre-trained DNABERT model
and prediction algorithms, which present new powerful methods for analyses and predictions of DNA sequences.
Since the pre-training of DNABERT is resource-intensive, we will provide the source code and pre-trained model
at Github for future academic research. We will also develop an integrated web server to (1) deploy DNABERT
model, (2) database to store the identified sequence features and predictions, and (3) tutorials to help users to
apply DNABERT to their specific research problems. We anticipate that DNABERT can bring new advancements
and insights to the bioinformatics community by bringing advanced language modeling perspective to gene
regulation analyses.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI:
10.48550/arxiv.2402.08777
发表时间:
2024-02
期刊:
ArXiv
影响因子:
--
作者:
[Zhihan Zhou;Weimin Wu;Harrison Ho;Jiayi Wang;Lizhen Shi;R. Davuluri;Zhong Wang;Han Liu]
通讯作者:
Zhihan Zhou;Weimin Wu;Harrison Ho;Jiayi Wang;Lizhen Shi;R. Davuluri;Zhong Wang;Han Liu
DOI:
10.1093/bioadv/vbad075
发表时间:
2023
期刊:
BIOINFORMATICS ADVANCES
影响因子:
--
作者:
[Ji, Yanrong, Dutta, Pratik, Davuluri, Ramana]
通讯作者:
Davuluri, Ramana
DOI:
10.1371/journal.pone.0243127
发表时间:
2021
期刊:
PloS one
影响因子:
3.7
作者:
[Taha K, Davuluri R, Yoo P, Spencer J]
通讯作者:
Spencer J
Developing novel deep-learning based methods for deciphering non-coding gene regulatory code
-
批准号:10451673
-
项目类别:
-
资助金额:$33.07万
-
财政年份:2021
-
负责人:RAMANA V DAVULURI
-
依托单位:
Informatics Platform for Mammalian Gene Regulation at Isoform-level
-
批准号:10273985
-
项目类别:
-
资助金额:$34.36万
-
财政年份:2020
-
负责人:RAMANA V DAVULURI
-
依托单位:
Informatics Platform for Mammalian Gene Regulation at Isoform-level
-
批准号:9922347
-
项目类别:
-
资助金额:$0.0万
-
财政年份:2013
-
负责人:RAMANA V DAVULURI
-
依托单位:
Informatics Platform for Mammalian Gene Regulation at Isoform-level
-
批准号:8843951
-
项目类别:
-
资助金额:$33.72万
-
财政年份:2013
-
负责人:RAMANA V DAVULURI
-
依托单位:
Informatics platform for mammalian gene regulation at isoform-level
-
批准号:8658144
-
项目类别:
-
资助金额:$33.72万
-
财政年份:2013
-
负责人:RAMANA V DAVULURI
-
依托单位:
Bioinformatics Facility
-
批准号:7945001
-
项目类别:
-
资助金额:$20.72万
-
财政年份:2009
-
负责人:RAMANA V DAVULURI
-
依托单位:
Genomewide discovery & analysis of alternative promoters
-
批准号:7678211
-
项目类别:
-
资助金额:$28.89万
-
财政年份:2006
-
负责人:RAMANA V DAVULURI
-
依托单位:
Genomewide discovery & analysis of alternative promoters
-
批准号:7226994
-
项目类别:
-
资助金额:$31.56万
-
财政年份:2006
-
负责人:RAMANA V DAVULURI
-
依托单位:
Genomewide discovery & analysis of alternative promoters
-
批准号:7033451
-
项目类别:
-
资助金额:$32.5万
-
财政年份:2006
-
负责人:RAMANA V DAVULURI
-
依托单位:
Genomewide discovery & analysis of alternative promoters
-
批准号:7371108
-
项目类别:
-
资助金额:$2.07万
-
财政年份:2006
-
负责人:RAMANA V DAVULURI
-
依托单位:
Genomewide discovery & analysis of alternative promoters
-
批准号:7580978
-
项目类别:
-
资助金额:$35.76万
-
财政年份:2006
-
负责人:RAMANA V DAVULURI
-
依托单位:
Core--Data Management and Computation Modeling
-
批准号:6993688
-
项目类别:
-
资助金额:$10.18万
-
财政年份:2004
-
负责人:RAMANA V DAVULURI
-
依托单位:
Bioinformatics Facility
-
批准号:8378483
-
项目类别:
-
资助金额:$19.86万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Core--Data Management and Computation Modeling
-
批准号:7681682
-
项目类别:
-
资助金额:$15.48万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Core--Data Management and Computation Modeling
-
批准号:7287751
-
项目类别:
-
资助金额:$10.8万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Core--Data Management and Computation Modeling
-
批准号:7557490
-
项目类别:
-
资助金额:$16.94万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Bioinformatics Facility
-
批准号:8102106
-
项目类别:
-
资助金额:$21.73万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Bioinformatics Facility
-
批准号:8233468
-
项目类别:
-
资助金额:$19.94万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Core--Data Management and Computation Modeling
-
批准号:7123769
-
项目类别:
-
资助金额:$10.49万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
Bioinformatics Facility
-
批准号:8461261
-
项目类别:
-
资助金额:$18.22万
-
财政年份:--
-
负责人:RAMANA V DAVULURI
-
依托单位:
国内基金
海外基金
企业绩效评价的DEA-Benchmarking方法及动态博弈研究
-
批准号:70571028
-
项目类别:面上项目
-
资助金额:16.5万元
-
批准年份:2005
-
负责人:杨印生
-
依托单位: