课题基金 / 基金详情

项目摘要

项目成果

Shaojie Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Abstract In the past few years, large genotyped cohorts are getting larger. We are getting closer to the era where genotype information of a large portion of the population is available. Informatics methods are critically needed for translating the new information into insights for human genetics. Powered by informatics innovations, the landscape of IBD segment detection has been transformed in the past 3 years. In 2019, we published RaPID, the first IBD segment calling method efficient enough for biobank-scale cohorts. Afterwards, a generation of methods has been developed to offer solutions for IBD segment calling. In addition, we delivered new algorithms and methods that enriched the PBWT data structure. Also, the impacts of calling out IBD segments in large cohorts are demonstrated by the powering of precision characterization of diversity in general population cohorts, the studies of population history and human behavior, IBD-based relatedness estimate, and IBD-mapping. However, current success in identifying IBD segments from biobank-scale cohorts is only the beginning. More informatics method developments are needed to fully unleash the power of genotype information. First, current methods are mainly for longer IBD segments (greater than 3 or 5 centimorgans (cM)), and the detection power for shorter segments are insufficient. Also the accuracy is not uniformly high across all genomic regions and all populations. Second, current methods are mainly for IBD segments shared between a pair of haplotypes. With large sample sizes, multi-way IBDs are omnipresent but under-studied. Third, methods for identifying IBD segments between a query haplotype and reference panels (1-vs-n) are needed. For a small sample or even individuals, 1-vs-n query against a panel will enable powerful interpretation leveraging the rich information in the reference panel. However, current IBD segment detection methods are mainly a batch calling mode that conducts n-vs-n comparisons and thus are not flexible enough to address such needs. In this competitive renewal project, we propose to further develop efficient, accurate, and flexible algorithms for IBD segment detection for large biobank-scale data. We will improve IBD segment calling across the genome, across length-spectrum, and across ethnicities; we will develop methods for multi-way IBD cluster detection; and we will develop reference- based IBD calling and threading methods. These new informatics methods will enable the community to better leverage the genetic relationships in large genotyped cohorts for genetic discovery.
期刊论文(17)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1371/journal.pgen.1009315
发表时间: 2021-01
期刊: PLoS genetics
影响因子: 4.5
作者: [Naseri A, Shi J, Lin X, Zhang S, Zhi D]
通讯作者: Zhi D
DOI: 10.1093/bioinformatics/btac734
发表时间: 2023-01-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [Wang V, Naseri A, Zhang S, Zhi D]
通讯作者: Zhi D
DOI: 10.1093/bioinformatics/btad312
发表时间: 2023-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: []
通讯作者:
DOI: 10.1093/bioadv/vbac045
发表时间: 2022
期刊: Bioinformatics advances
影响因子: --
作者: []
通讯作者:
12
    Genome Informatics For Biobank-scale Data
    Scalable methods for identity by descent
    Identification, Discovery, and Public Archiving of RNA Structural Motifs
    • 批准号:
      8348532
    • 项目类别:
    • 资助金额:
      $16.92万
    • 财政年份:
      2012
    • 负责人:
      Shaojie Zhang
    • 依托单位:
    Identification, Discovery, and Public Archiving of RNA Structural Motifs
    • 批准号:
      8723857
    • 项目类别:
    • 资助金额:
      $16.88万
    • 财政年份:
      2012
    • 负责人:
      Shaojie Zhang
    • 依托单位:
    海外基金