Leveraging long-range haplotypes in sequencing data to advance large scale genetic studies
Leveraging long-range haplotypes in sequencing data to advance large scale genetic studies
批准号:
10653188
负责人:
Sebastian Zoellner
金额:
$36.51万
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2024-06-30
关键词:
AccelerationAddressAffectAgeAlgorithmsAwarenessBase SequenceBiologyComplexComputer softwareComputing MethodologiesDataData SetDevelopmentDiseaseDisease modelEtiologyExonsGene FrequencyGenesGeneticGenetic CodeGenetic VariationGenetic studyGenomeGenomicsGenotypeHaplotypesHeterozygoteHumanHuman Gene MappingHuman GeneticsHuman Genome ProjectIndividualLengthMethodsMinorModelingPatternPhasePhenotypePopulationPopulation GeneticsPropertyRegulatory ElementResearchResourcesSample SizeSamplingSignal TransductionSoftware ToolsStatistical MethodsStatistical ModelsStructureTechnologyTestingTrans-Omics for Precision MedicineUntranslated RNAVariantcausal variantcomputer sciencedesigndisease classificationdisorder riskempowermentexomegenome sequencinggenome wide association studyhuman diseaseidentity by descentimprovedinsightlarge datasetslarge scale datanovelnovel therapeuticsprogramsrare variantrisk predictionsoftware developmentstatisticssuccesstargeted treatmenttooltraituser-friendlyvariant of interest
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The Human Genome Project and subsequent projects such as 1000 Genomes, Genome Sequencing Program
(GSP), and Trans-Omics Precision Medicine (TOPMed) are providing powerful resources for studying the
genetic basis of human diseases. Combining these resources and technologies with the development of new
statistical and computational methods have in the last decade led to identification of thousands of loci
associated with disease-related phenotypes, primarily through array-based genome-wide association studies
(GWAS), empowered by genotype imputation from sequence-based haplotype panel. However, serious
problems remain when analyzing these data: (1) As short read sequencing data only provides unphased
genotype data, methods for statistical phasing are used to allow advanced analyses and to generate reference
haplotypes for genotype imputation. However, current methods to phase sequence data result in several
thousand switch errors per genome. These phasing errors in turn limit the accuracy of genotype imputation and
hamper our ability to study haplotype-aware disease models such as compound heterozygotes. (2) Due to the
abundance of rare variants, it is necessary to identify high-interest variants to obtain powerful test statistics.
Within exons, the genetic code provides some of the necessary information, but for most the genome we have
very little information that allows us to prioritize variants. (3) While samples sequenced from diverse and
admixed populations are becoming more common, few methods are designed to make use of the unique
properties of such data. For example, the distribution of local ancestry in admixed samples generate unique
haplotype structure that can be informative about the underlying phasing. Here we propose a set of novel
methods that will address these challenges: recognizing that in very large datasets most sequences will have a
recent common ancestor with at least one other sequence and that these closely related sequences will share
long segments (>1 cM) identical by descent (IBD). These IBD segments provides information about the
phasing of the underlying variants similar to large sibships. Moreover, the length of the IBD segment provides
information about the age of variants located on the IBD segment. As young variants are more likely to be
under selection, IBD length can be used to prioritize functional noncoding variants. We also aim to leverage the
long-distance correlation of genotypes in admixed samples to identify phasing errors in admixed samples. As
phasing errors also change the local ancestry of a sample in individuals of heterozygous ancestry, identifying
these breaks allows identifying and correcting phasing errors. We will develop statistical models that leverage
these conceptual ideas and implement these methods in algorithms efficient enough to be applied to sample
sizes >100,000. We will use our algorithms to annotate and re-phase existing large sequencing datasets and
thus improve commonly used imputation reference panels. All software developed in this proposal will be
publicly released in user-friendly, well-documented packages.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
FICTURE: Scalable segmentation-free analysis of submicron resolution spatial transcriptomics.
图:亚微米分辨率空间转录组学的可扩展无分割分析。
DOI:
10.1101/2023.11.04.565621
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Si,Yichen, Lee,ChangHee, Hwang,Yongha, Yun,JeongH, Cheng,Weiqiu, Cho,Chun-Seok, Quiros,Miguel, Nusrat,Asma, Zhang,Weizhou, Jun,Goo, Zöllner,Sebastian, Lee,JunHee, Kang,HyunMin]
通讯作者:
Kang,HyunMin
Seq-Scope Protocol: Repurposing Illumina Sequencing Flow Cells for High-Resolution Spatial Transcriptomics.
Seq-Scope 协议:重新利用 Illumina 测序流动槽实现高分辨率空间转录组学。
DOI:
10.1101/2024.03.29.587285
发表时间:
2024
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
作者:
[Kim,Yongsung, Cheng,Weiqiu, Cho,Chun-Seok, Hwang,Yongha, Si,Yichen, Park,Anna, Schrank,Mitchell, Hsu,Jer-En, Xi,Jingyue, Kim,Myungjin, Pedersen,Ellen, Koues,OliviaI, Wilson,Thomas, Jun,Goo, Kang,HyunMin, Lee,JunHee]
通讯作者:
Lee,JunHee
SiftCell: A robust framework to detect and isolate cell-containing droplets from single-cell RNA sequence reads.
SiftCell:一个强大的框架,用于从单细胞 RNA 序列读取中检测和分离含有细胞的液滴。
DOI:
10.1016/j.cels.2023.06.002
发表时间:
2023
期刊:
Cell systems
影响因子:
9.3
作者:
[Xi,Jingyue, Park,SungRye, Lee,JunHee, Kang,HyunMin]
通讯作者:
Kang,HyunMin
Leveraging long-range haplotypes in sequencing data to advance large scale genetic studies
-
批准号:10477336
-
项目类别:
-
资助金额:$36.2万
-
财政年份:2020
-
负责人:Sebastian Zoellner
-
依托单位:
Leveraging long-range haplotypes in sequencing data to advance large scale genetic studies
-
批准号:10251017
-
项目类别:
-
资助金额:$35.9万
-
财政年份:2020
-
负责人:Sebastian Zoellner
-
依托单位:
Computational Statistic Approaches to Gene-Environment Interaction
-
批准号:7348103
-
项目类别:
-
资助金额:$37.22万
-
财政年份:2007
-
负责人:Sebastian Zoellner
-
依托单位:
Computational Statistic Approaches to Gene-Environment Interaction
-
批准号:7666932
-
项目类别:
-
资助金额:$37.22万
-
财政年份:2007
-
负责人:Sebastian Zoellner
-
依托单位:
海外基金