Scaling up computational genomics with tree sequences
Scaling up computational genomics with tree sequences
批准号:
10471496
负责人:
PETER Lochhead RALPH
金额:
$55.68万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-24 至 2023-08-31
关键词:
AddressAffectAlgorithmic SoftwareAlgorithmsArchitectureAreaBase SequenceCollectionCommunitiesComplexComputer softwareComputing MethodologiesDataData CompressionData SetDevelopmentDiseaseEcologyEnsureEpidemiologyEtiologyEvolutionGenealogical TreeGenealogyGenerationsGeneticGenetic ProcessesGenetic RecombinationGenetic VariationGenomeGenomicsGenotypeGoalsHaplotypesHealthHealth BenefitHumanHuman GeneticsHuman GenomeIndividualInternetLibrariesMapsMethodsModelingModernizationMutationPerformancePhasePhenotypePopulationPopulation GeneticsPopulation SizesPositioning AttributeProcessProductionRecording of previous eventsRecordsResearchRunningSample SizeSamplingStatistical Data InterpretationStructureTestingTimeTrainingTreesTsunamiValidationVariantWorkalgorithm developmentbasecomputer frameworkcostdata formatdata reusedata structuredeep learningdesignexperiencefrontiergenome-widegenomic datahuman diseaseimprovedinteroperabilitylearning strategymembermulticore processornext generationnovelnovel strategiesopen sourceoperationscale upsequence learningsimulationstatisticsstructural genomicssuccesssupervised learningwhole genome
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary/Abstract
Increasing sample size is a tremendously important factor in building our understanding of the genetics of
human disease. As we discover that more and more diseases have a complex web of genetic causation, we
need larger and larger genetic datasets to disentangle them, and to ultimately produce successful therapies.
Driven in part by this need, the community is now assembling vast collections of human genome sequences,
and millions of samples will soon be commonplace. Nonhuman datasets, with applications in epidemiology,
ecology, and evolution, will not be far behind. There is a profound problem, however: our computational
methods for storing, processing, simulating, and analyzing genomic data are lagging far behind our ability to
collect such data. The algorithms and data structures underlying today's computational methods were designed
for thousands of samples, not millions, and we are in danger of being overwhelmed by the impending tsunami
of data. Without a fundamental change in how we store and process genomic data, we will either not fully tap
the potential of the data we collect, or the computational costs will be astronomical – or both.
Our proposal addresses this critical need by focusing on a new data structure: the succinct tree sequence.
This data structure (the “tree sequence”, for brevity) encodes genetic variation data using the population ge-
netics processes that produced the data itself – by representing variation among contemporary samples via
mutations on the branches of the underlying genealogical trees. This yields extraordinary levels of data com-
pression, with file sizes hundreds of times smaller than current community standards. Since the tree sequence
was introduced in 2016 it has led to performance increases of 2–4 orders of magnitude in the diverse applica-
tions of genome simulation, calculation of statistics, and ancestry inference. Such sudden leaps in computa-
tional performance are vanishingly rare, and only possible through deep algorithmic advances.
Our research plan builds on the extraordinary successes of tree sequence methods so far, scaling up three
crucial layers of computational genomics: analysis, simulation, and inference. First, we will continue our
development of highly efficient tree-sequence-based methods for fundamental operations in statistical and
population genetics. Second, we will scale up genome simulations by integrating tree sequence methods into
complex forward-time simulations, utilizing modern, multicore processors. Third, we will combine efficient
genome simulations with cutting-edge deep-learning methods to improve existing inference methods, both
of tree sequences from genomic data, and of population parameters from novel tree-sequence encodings of
genotype data. Together, we aim to revolutionize the way we work with population genetic variation data, and
how we use it to understand human health and evolutionary processes.
Our experienced, interdisciplinary team is committed to producing rigorously tested and validated software
and accessible, interoperable, and reusable data formats through inclusive and open development.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1214/24-ejp1075
发表时间:
2024-01-01
期刊:
ELECTRONIC JOURNAL OF PROBABILITY
影响因子:
1.4
作者:
[Etheridge,Alison M., Kurtz,Thomas G., Lung,Terence Tsui Ho]
通讯作者:
Lung,Terence Tsui Ho
tstrait: a quantitative trait simulator for ancestral recombination graphs.
tstrait:祖先重组图的数量性状模拟器。
DOI:
10.1101/2024.03.13.584790
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Tagami,Daiki, Bisschop,Gertjan, Kelleher,Jerome]
通讯作者:
Kelleher,Jerome
Estimating evolutionary and demographic parameters via ARG-derived IBD.
通过 ARG 衍生的 IBD 估计进化和人口统计参数。
DOI:
10.1101/2024.03.07.583855
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Huang,Zhendong, Kelleher,Jerome, Chan,Yao-Ban, Balding,DavidJ]
通讯作者:
Balding,DavidJ
DOI:
10.1093/bioadv/vbad163
发表时间:
2023
期刊:
Bioinformatics advances
影响因子:
--
作者:
[]
通讯作者:
Genetic architecture, spatial heterogeneity, and the coevolutionary arms race between newts and snakes.
遗传结构、空间异质性以及蝾螈和蛇之间的共同进化军备竞赛。
DOI:
10.1101/2023.12.07.570693
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Caudill,Victoria, Ralph,PeterL]
通讯作者:
Ralph,PeterL
Scaling up computational genomics with tree sequences
-
批准号:10585745
-
项目类别:
-
资助金额:$60.57万
-
财政年份:2023
-
负责人:PETER Lochhead RALPH
-
依托单位:
Geographic models of selective sweeps
-
批准号:8370584
-
项目类别:
-
资助金额:$2.49万
-
财政年份:2011
-
负责人:PETER Lochhead RALPH
-
依托单位:
Geographic models of selective sweeps
-
批准号:8198779
-
项目类别:
-
资助金额:$5.13万
-
财政年份:2011
-
负责人:PETER Lochhead RALPH
-
依托单位:
海外基金