SOFTWARE FOR LARGE-SCALE INFERENCE OF THE GENETICS OF LIFESTYLE MEASURES, BIOMARKERS, AND COMMON AND RARE DISEASES
SOFTWARE FOR LARGE-SCALE INFERENCE OF THE GENETICS OF LIFESTYLE MEASURES, BIOMARKERS, AND COMMON AND RARE DISEASES
批准号:
10440494
负责人:
Anshul Kundaje
金额:
$39.25万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-06 至 2023-12-31
关键词:
AddressAffectBayesian ModelingBiological MarkersCodeCommunitiesComputer softwareDNA sequencingDataDatabase Management SystemsDiseaseDisease OutcomeEnsureEnvironmental ExposureEnvironmental Risk FactorFrequenciesFundingFutureGenerationsGenesGeneticGenetic DiseasesGenetic ProcessesGenetic ResearchGenetic studyGenomicsGenotypeGoalsHealthHourHuman GeneticsInflammatory Bowel DiseasesInternationalJointsKnowledgeLaboratoriesLettersLife StyleMeasurementMeasuresMedicalMedical GeneticsMethodsMissionNational Human Genome Research InstituteNegative FindingPerformancePersonsPhenotypePolygenic TraitsPopulationPopulation GeneticsPrecision Medicine InitiativePrivacyProgramming LanguagesQuality ControlRare DiseasesResearchResearch DesignResearch PersonnelResourcesSecureStatistical AlgorithmStatistical MethodsStatistical ModelsStreamTimeTrans-Omics for Precision MedicineUnited StatesUnited States National Institutes of HealthUntranslated RNAVariantVisualizationalgorithm developmentbiobankcloud basedcost effectivedata disseminationdata integrationdata sharingdata visualizationdesigndisease phenotypedisorder riskepigenomicsexome sequencingexperienceflexibilitygenetic analysisgenetic associationgenetic variantgenome sequencinggenomic datahuman diseaseimprovedinsightlarge scale datalarge-scale databasemethod developmentnovelphenotypic datapleiotropismprogramssharing platformsoftware developmenttherapeutically effectivetoolusabilityweb platform
中文摘要
世界各地大规模的种群生物库,疾病集中在NHGRI基因组测序计划
(GSP)和美国的我们所有人精确医学倡议项目将产生大量基因组
结合疾病结果和其他健康衡量标准的数据集。这些基因组研究将
确定与健康和疾病相关的基因组变异。然而,他们在所有背景下的联系
如果对数据进行单独分析,确定的可能关联仍不清楚。有一个不断增长的
认识到大多数性状都是多基因的。此外,人们越来越认识到,多效性是普遍存在的。
出于隐私方面的考虑,共享所有可能的基因和表型数据是一件具有挑战性的事情。方法可以
对汇总级数据进行推断,例如p值、效果大小估计和频率,将有助于我们
了解人类疾病和健康的遗传学。在这里,我们建议开发软件,用于
生活方式指标、生物标志物以及常见和罕见的遗传学的大规模推断
疾病。实现这一目标需要医学和人口遗传学、统计学方法方面的专业知识
开发,以及大型数据库管理方面的专业知识。该项目有三个主要目标。
首先,我们将创建全球生物库引擎:一个强大的、交互式的Web平台,用于推理
生活方式、生物标志物、常见病和罕见病的遗传学研究。我们将通过以下方式扩展功能
实施用于标记变体和表型的质量控制可视化和方法。我们将添加
研究设计工具,使用经验数据来估计统计能力,并创建灵活的框架
联合分析多种表型,同时控制假阳性和假阴性的统计模型
调查结果。其次,我们将提高全球生物库引擎的性能、可扩展性和可访问性
方便今后人群生物库和有针对性的常见病和罕见病。我们将创建一个托管的、
安全、经济高效的基于云的社区资源,并设计一个数据库系统,以减少
加载遗传关联研究的时间从几小时到几分钟,并允许传输统计数据
算法直接转化为遗传数据。最后,我们将改进基因组解释、可视化和数据
共享,通过实施新颖的分析大幅提高翻译发现的速度
方法:研究方法。我们将支持新的变体标注方法,并整合编码和非编码信息,
包括来自大规模表观基因组学研究的数据,用于变异和基因水平的推断。我们将实施
用概率编程语言实现的新的贝叶斯统计模型,稀疏规范
相关分析和截断奇异值分解。皮里瓦斯和他的团队有足够的
NIH资助的财团的经验,他们致力于NIH及其资助的整体使命
调查人员将发现新的知识,这些知识将为每个人带来更好的健康。
英文摘要
Large-scale population biobanks around the world, the disease focused NHGRI Genome Sequencing Program
(GSP), and the United States’ All of Us Precision Medicine Initiative project will generate massive genomic
datasets combined with disease outcomes, and other health measurements. These genomic studies will
identify genomic variants relevant to health and disease. However, their association in the context of all
possible associations identified will remain unclear if the data are separately analyzed. There is a growing
recognition that most traits are polygenic. In addition, it is increasingly appreciated that pleiotropy is pervasive.
Due to privacy concerns, it is challenging to share all possible genotype and phenotype data. Methods that can
perform inference on summary level data, e.g. p-values, effect size estimates, and frequency, will facilitate our
understanding of the genetics of human diseases and health. Here, we propose to develop software for
large-scale inference of the genetics of lifestyle measures, biomarkers, and common and rare
diseases. Achieving this goal requires expertise in medical and population genetics, statistical methods
development, and expertise in management of large-scale databases. The project has three main objectives.
First, we will create Global Biobank Engine: a powerful, interactive web platform for inference of the
genetics of lifestyle measures, biomarkers, common and rare diseases. We will expand the features by
implementing quality control visualizations and methods for flagging variants and phenotypes. We will add
tools for study design that use empirical data to estimate statistical power, and create a flexible framework for
statistical models that jointly analyze multiple phenotypes while controlling for false positive and negative
findings. Secondly, we will improve Global Biobank Engine performance, scalability, and accessibility
to facilitate future population biobanks and targeted common and rare disease. We will create a hosted,
secure, and cost-effective cloud-based community resource, and design a database system that reduces the
loading time for genetic association studies from hours to minutes and allows for streaming of statistical
algorithms directly to genetic data. Lastly, we will improve genomic interpretation, visualization, and data
sharing to dramatically increase the rate of translational discoveries by implementing novel analysis
methods. We will support new variant annotation methods and integrate coding and non-coding information,
including data from large-scale epigenomics studies, for variant and gene level inference. We will implement
new Bayesian statistical models implemented in probabilistic programming languages, sparse canonical
correlation analysis, and truncated singular value decomposition. PI Rivas and his team have ample
experience with NIH-funded consortia, and they are dedicated to the overall mission of NIH and its funded
investigators to uncover new knowledge that will lead to better health for everyone.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
Integrative machine learning approaches for predicting disease risk using multi-omics data from the UK Biobank.
使用英国生物银行的多组学数据预测疾病风险的综合机器学习方法。
DOI:
10.1101/2024.04.16.589819
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Aguilar,Oscar, Chang,Cheng, Bismuth,Elsa, Rivas,ManuelA]
通讯作者:
Rivas,ManuelA
SALAI-Net: species-agnostic local ancestry inference network.
SALAI-Net:与物种无关的本地祖先推理网络。
DOI:
10.1093/bioinformatics/btac464
发表时间:
2022
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[OriolSabat,Benet, MasMontserrat,Daniel, Giro-I-Nieto,Xavier, Ioannidis,AlexanderG]
通讯作者:
Ioannidis,AlexanderG
DOI:
10.1101/2023.10.12.561949
发表时间:
2023-10-17
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Bonet D, Levin M, Montserrat DM, Ioannidis AG]
通讯作者:
Ioannidis AG
Multi-Omics DACC: The Data Analysis and Coordination Center for the collaborative multi-omics for health and disease initiative
-
批准号:10744561
-
项目类别:
-
资助金额:$311.62万
-
财政年份:2023
-
负责人:Anshul Kundaje
-
依托单位:
A Comprehensive Genomic Community Resource of Transcriptional Regulation
-
批准号:10411262
-
项目类别:
-
资助金额:$83.35万
-
财政年份:2022
-
负责人:Anshul Kundaje
-
依托单位:
A Comprehensive Genomic Community Resource of Transcriptional Regulation
-
批准号:10842047
-
项目类别:
-
资助金额:$20.28万
-
财政年份:2022
-
负责人:Anshul Kundaje
-
依托单位:
A Comprehensive Genomic Community Resource of Transcriptional Regulation
-
批准号:10625529
-
项目类别:
-
资助金额:$80.94万
-
财政年份:2022
-
负责人:Anshul Kundaje
-
依托单位:
Identifying causal genetic variants and molecular mechanisms impacting mental health
-
批准号:10571911
-
项目类别:
-
资助金额:$61.6万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Identifying causal genetic variants and molecular mechanisms impacting mental health
-
批准号:10380573
-
项目类别:
-
资助金额:$61.62万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Predicting context-specific molecular and phenotypic effects of genetic variation through the lens of the cis-regulatory code
-
批准号:10659170
-
项目类别:
-
资助金额:$72.74万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Predicting context-specific molecular and phenotypic effects of genetic variation through the lens of the cis-regulatory code
-
批准号:10297562
-
项目类别:
-
资助金额:$35.22万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Predicting context-specific molecular and phenotypic effects of genetic variation through the lens of the cis-regulatory code
-
批准号:10474459
-
项目类别:
-
资助金额:$72.74万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Multi-omic functional assessment of novel AD variants using high-throughput and single-cell technologies
-
批准号:10684210
-
项目类别:
-
资助金额:$166.0万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Multi-omic functional assessment of novel AD variants using high-throughput and single-cell technologies
-
批准号:10217784
-
项目类别:
-
资助金额:$169.82万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Multi-omic functional assessment of novel AD variants using high-throughput and single-cell technologies
-
批准号:10436207
-
项目类别:
-
资助金额:$166.92万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
Identifying causal genetic variants and molecular mechanisms impacting mental health
-
批准号:10116649
-
项目类别:
-
资助金额:$61.8万
-
财政年份:2021
-
负责人:Anshul Kundaje
-
依托单位:
SOFTWARE FOR LARGE-SCALE INFERENCE OF THE GENETICS OF LIFESTYLE MEASURES, BIOMARKERS, AND COMMON AND RARE DISEASES
-
批准号:10251897
-
项目类别:
-
资助金额:$39.25万
-
财政年份:2018
-
负责人:Anshul Kundaje
-
依托单位:
Deep learning frameworks for regulatory genomics.
-
批准号:9169521
-
项目类别:
-
资助金额:$235.5万
-
财政年份:2016
-
负责人:Anshul Kundaje
-
依托单位:
海外基金