Computational methods for variant calling and haplotyping using long-read sequencing technologies
Computational methods for variant calling and haplotyping using long-read sequencing technologies
批准号:
10058104
负责人:
Vikas Bansal
金额:
$38.17万
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2024-06-30
关键词:
AddressAlgorithmsBar CodesBenchmarkingBiological SciencesCatalogsCharacteristicsChromosome MappingClinicalCodeComplexComputing MethodologiesDNADataData SetDetectionDiploidyDiseaseEvolutionExhibitsGene DuplicationGenerationsGenesGenetic DiseasesGenetic VariationGenetic studyGenomeGenomicsGenotypeGoalsHaplotypesHereditary Nonpolyposis Colorectal NeoplasmsHuman GeneticsHuman GenomeIndividualLibrariesLinkMalignant NeoplasmsMapsMassive Parallel SequencingMedicalMedical GeneticsMendelian disorderMethodsModelingMutationPMS2 genePhasePopulationPopulation GeneticsPopulation StudyPreparationRepetitive SequenceSequence HomologsSequence HomologySingle Nucleotide PolymorphismSoftware ToolsTechnologyVariantbasecomputerized toolsdisease-causing mutationgenetic analysisgenome sequencinghearing impairmenthuman diseasehuman genome sequencinghuman reference genomeimprovedinnovationinsertion/deletion mutationmultiple datasetsnanoporenext generation sequencingpreventpublic health relevancesingle moleculetoolvariant detectionvirtualwhole genome
中文摘要
项目摘要/摘要
在这个项目中,我们建议开发全基因组单倍型的计算方法和工具
使用太平洋生物科学公司和牛津大学等长读测序技术的小型变异叫声
纳米孔和链接读取技术。单倍型信息是解释遗传变异的关键
在个体基因组、疾病图谱、临床基因组学和其他几种人类遗传分析中
变种。在使用短阅读测序的人类基因组中缺乏阶段或单倍型信息是一种
识别与复合杂合突变有关的疾病的主要障碍。600多个基因
序列同源性高的重叠片段复制和100多个此类基因的变异
与罕见的孟德尔疾病和包括癌症在内的复杂疾病有关。无法检测到
利用短读测序技术在基因组重复区域中获得高精度的变异
降低了在医学遗传学研究中识别引起疾病的突变的能力。在目标1中,我们将开发
一种基于长阅读的二倍体基因分型通用计算方法,可实现准确的单倍体分型
对于使用长阅读和连锁阅读以及准确的小片段的单核苷酸变体和短indels
使用短信技术的变体呼叫。在目标2中,我们将开发敏感映射的计算方法
在人类基因组的重复区域中的短信读取和准确的变体调用,目前
排除在参考人类基因组的基准小变体呼叫集之外。最后,在目标3中,我们将
利用AIMS 1和AIMS 2中的方法对使用短信测序的多个基因组执行变体调用
对变体PSV进行分类并利用此目录改进读取映射和变体调用的技术
基因组重复区域短读测序的准确性。我们将在
强大且计算高效的软件工具,并使用公开可用的Long-
读取不同祖先的多个人类基因组的序列数据集。
英文摘要
Project Summary/Abstract
In this project, we propose to develop computational methods and tools for whole-genome haplotyping and
small variant calling using long-read sequencing technologies such as Pacific Biosciences and Oxford
Nanopore and linked-read technologies. Haplotype information is crucial for interpretation of genetic variation
in individual genomes, disease mapping, clinical genomics and several other analysis of human genetic
variation. The lack of phase or haplotype information in human genomes sequenced using short reads is a
major barrier in identifying disease associations with compound heterozygous mutations. More than 600 genes
overlap segmental duplications with high sequence identity and variants in more than 100 such genes have
been associated with rare Mendelian disorders and complex diseases including cancer. The inability to detect
variants with high accuracy in duplicated regions of the genome using short-read sequencing technologies
reduces the ability to identify disease causing mutations in medical genetics studies. In Aim 1, we will develop
a general computational method for long-read based diploid genotyping that will enable accurate haplotyping
for single nucleotide variants and short indels using long-read and linked-reads as well as accurate small
variant calling using SMS technologies. In Aim 2, we will develop computational methods for sensitive mapping
of SMS reads and accurate variant calling in repetitive regions of the human genome that are currently
excluded from benchmark small variant call sets for reference human genomes. Finally, in Aim 3, we will
leverage the methods from Aims 1 and 2 to perform variant calling on multiple genomes sequenced using SMS
technologies to catalog variant PSVs and leverage this catalog to improve read mapping and variant calling
accuracy of short-read sequencing in repetitive regions of the genome. We will implement the methods in
robust and computationally efficient software tools and benchmark their accuracy using publicly available long-
read sequence datasets for multiple human genomes of diverse ancestries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational methods for variant calling and haplotyping using long-read sequencing technologies
-
批准号:10657420
-
项目类别:
-
资助金额:$38.1万
-
财政年份:2020
-
负责人:Vikas Bansal
-
依托单位:
Computational methods for variant calling and haplotyping using long-read sequencing technologies
-
批准号:10441522
-
项目类别:
-
资助金额:$38.16万
-
财政年份:2020
-
负责人:Vikas Bansal
-
依托单位:
Computational methods for variant calling and haplotyping using long-read sequencing technologies
-
批准号:10247821
-
项目类别:
-
资助金额:$38.23万
-
财政年份:2020
-
负责人:Vikas Bansal
-
依托单位:
Methods for detecting short indels from high-throughput sequence data
-
批准号:8706938
-
项目类别:
-
资助金额:$19.38万
-
财政年份:2013
-
负责人:Vikas Bansal
-
依托单位:
Methods for detecting short indels from high-throughput sequence data
-
批准号:8572023
-
项目类别:
-
资助金额:$28.26万
-
财政年份:2013
-
负责人:Vikas Bansal
-
依托单位:
海外基金