课题基金 / 基金详情

Optimized workflows for structural variant analysis of the Kids First genomes using short and long reads

Optimized workflows for structural variant analysis of the Kids First genomes using short and long reads
使用短读长和长读长对 Kids First 基因组进行结构变异分析的优化工作流程
批准号:
10432507
负责人:
MICHAEL SCHATZ
金额:
$15.63万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-04-01 至 2024-03-31

项目摘要

项目成果

MICHAEL SCHATZ的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 加布里埃拉·米勒儿童第一儿科研究计划的总体目标是减轻 通过促进合作研究来揭示儿童癌症和结构性出生缺陷的病因 这些疾病。该计划最近增加了一个项目,即儿童第一长读试点项目,该项目利用 长读测序技术,进一步解析患者的基因组。这些技术已经在 通过允许人类基因组的完整端粒到端粒(T2T)重建来改变基因组学 这是第一次,通过允许发现结构变体和其他复杂的变体, 以前使用短读取序列无法访问。 在这里,我们将通过开发和应用优化的云规模来增强Kids First数据集的效用 使用新的T2T-CHM13人类基因组分析短读和长读数据集的工作流程。在 T2T联盟,我们领导了研究CHM13基因组如何影响变体呼叫的工作,以及 我发现T2T参考普遍改进了使用短和长的遗传变异分析 阅读排序。在这里,我们将开发优化的工作流,用于使用 使用GATK的T2T-CHM13参考基因组用于SNV和小INDELS,而Lement2用于短读SV 发现号。接下来,我们将为长读结构变异检测开发优化的工作流。短读数 面临检测多种类型的突变(例如,SVS、重复扩展等)的挑战,但无法解决许多问题 基因组的重复区域,包括许多医学上相关的基因。长篇阅读显示出很好的效果 承诺应对这些挑战并发现新的疾病关联,因为其可映射性增加, 不同的分辨率,以及分相功能。为了首先为儿童启用这些技术,我们将开发 优化的工作流,用于使用Jasmine准确识别和比较长时间读取样本中的SVS,如 以及在带有段落的短读数据集中通过长读发现的SVS的基因分型。这将使我们能够 分析在数量大得多的短读取数据集中通过长读取发现的变体并对其进行优先排序。 然后,我们将这些工作流应用到Kids First数据资源,以开发改进的变体调用和 改进了对这些珍贵样本的变异分析。这将导致发现数以千计的SVS ,并将减少错误变体的数量,否则会混淆任何 下游分析。我们还将开发新的统计和机器学习方法,以确定 最有可能与所研究的疾病相关的变异,利用系谱信息和 可用的基因组注释,以支持我们确定这些基因的驱动突变的总体目标 疾病。所有的工作流程和软件开发都将以开源方式发布,以供在CAVATICA、 所有儿童优先研究人员使用的基于云的分析平台,确保可伸缩性和可重复性。
英文摘要
Project Summary The overall goal of the Gabriella Miller Kids First Pediatric Research Program is to alleviate suffering from childhood cancer and structural birth defects by fostering collaborative research to uncover the etiology of these diseases. A recent addition to the program is the Kids First Long Read Pilot Projects, which are leveraging long-read sequencing technologies to further resolve the patients’ genomes. Already these technologies are transforming genomics by allowing complete telomere-to-telomere (T2T) reconstructions of human genomes for the first time, and by allowing the discovery of structural variants and other complex variants that were previously inaccessible using short read sequencing. Here we will enhance the utility of the Kids First data sets by developing and applying optimized cloud-scale workflows for analyzing short and long read datasets with the new T2T-CHM13 human genome. Within the T2T consortium, we have led the effort to characterize how the CHM13 genome influences variant calling, and have found the T2T reference universally improves the analysis of genetic variation using both short and long read sequencing. Here we will develop optimized workflows for analyzing short read datasets with the T2T-CHM13 reference genome using GATK for SNVs and small indels, and Parliament2 for short-read SV discovery. Next we will develop optimized workflows for Long Read Structural Variant Detection. Short-reads are challenged to detect many classes of mutations (e.g. SVs, repeat expansions, etc), and cannot resolve many repetitive regions of the genome, including within many medically relevant genes. Long-reads show great promise to address these challenges and discover new disease associations due to its increased mappability, variant resolution, and phasing capabilities. To enable these technologies for Kids First, we will develop optimized workflows for accurately identifying and comparing SVs across long read samples with Jasmine, as well as genotyping SVs discovered by long reads within short read datasets with Paragraph. This will enable us to analyze and prioritize variants found by long reads within the much larger numbers of short read datasets. We will then apply these workflows to the Kids First data resource to develop improved variant calls and improved variant analysis of these precious samples. This will lead to the discovery of thousands of SVs that were previously missed, and will reduce the number of false variants that would otherwise confuse any downstream analysis. We will also develop new statistical and machine learning approaches for prioritizing the variants that are most likely to be related to the studied diseases, leveraging the pedigree information and genome annotations available, in support of our overall goal of identifying the driver mutations for these diseases. All workflows and software developments will be released open source for use in CAVATICA, the cloud-based analysis platform used by all Kids First researchers, ensuring scalability and reproducibility.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EXPANDING THE GENOMIC DATA SCIENCE COMMUNITY NETWORK FOR NHGRI.
  • 批准号:
    10944109
  • 项目类别:
  • 资助金额:
    $74.91万
  • 财政年份:
    2023
  • 负责人:
    MICHAEL SCHATZ
  • 依托单位:
Optimized workflows for structural variant analysis of the Kids First genomes using short and long reads
  • 批准号:
    10602532
  • 项目类别:
  • 资助金额:
    $15.6万
  • 财政年份:
    2022
  • 负责人:
    MICHAEL SCHATZ
  • 依托单位:
Integrative genomic and epigenomic analysis of cancer using long read sequencing
  • 批准号:
    10396074
  • 项目类别:
  • 资助金额:
    $35.67万
  • 财政年份:
    2021
  • 负责人:
    MICHAEL SCHATZ
  • 依托单位:
Integrative genomic and epigenomic analysis of cancer using long read sequencing
  • 批准号:
    10599150
  • 项目类别:
  • 资助金额:
    $35.4万
  • 财政年份:
    2021
  • 负责人:
    MICHAEL SCHATZ
  • 依托单位:
海外基金