Computational methods for variant calling and haplotyping using long-read sequencing technologies
Computational methods for variant calling and haplotyping using long-read sequencing technologies
批准号:
10657420
负责人:
Vikas Bansal
金额:
$38.1万
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2024-06-30
关键词:
AddressAlgorithmsBar CodesBenchmarkingBiological SciencesCatalogsCharacteristicsChromosome MappingClinicalCodeComplexComputing MethodologiesDNADataData SetDetectionDiploidyDiseaseEvolutionExclusionExhibitsGene DuplicationGenerationsGenesGenetic DiseasesGenetic VariationGenetic studyGenomeGenomicsGenotypeGoalsHaplotypesHereditary Nonpolyposis Colorectal NeoplasmsHeterozygoteHuman GeneticsHuman GenomeIndividualLibrariesLinkMalignant NeoplasmsMapsMassive Parallel SequencingMedicalMedical GeneticsMethodsModelingMutationOverlapping GenesPMS2 genePhasePopulationPopulation GeneticsPopulation StudyPreparationRepetitive SequenceSequence HomologsSequence HomologySingle Nucleotide PolymorphismSoftware ToolsTechnologyVariantcomputerized toolsdisease-causing mutationgenetic analysisgenome sequencinghearing impairmenthuman diseasehuman genome sequencinghuman reference genomeimprovedinnovationinsertion/deletion mutationmultiple datasetsnanoporenext generation sequencingpreventpublic health relevancerare mendelian disordersingle moleculetoolvariant detectionvirtualwhole genome
中文摘要
项目概要/摘要
在这个项目中,我们建议开发用于全基因组单体型分析的计算方法和工具,
使用长读序测序技术(如Pacific Biosciences和Oxford)的小变异识别
纳米孔和链接读取技术。单倍型信息是解释遗传变异的关键
在个体基因组、疾病图谱、临床基因组学和人类遗传学的其他几种分析中,
变化量在使用短读段测序的人类基因组中缺乏相位或单体型信息是一个潜在的问题。
在确定疾病与复合杂合突变的关系方面存在主要障碍。600多个基因
在100多个这样的基因中具有高序列同一性的重叠片段重复和变体,
与罕见的孟德尔疾病和包括癌症在内的复杂疾病有关。无法检测
使用短读测序技术在基因组重复区域中具有高准确性的变体
降低了在医学遗传学研究中识别致病突变的能力。在目标1中,我们将开发
一种基于长读段的二倍体基因分型的通用计算方法,
对于单核苷酸变体和短插入缺失,使用长读段和连接读段以及精确的小插入缺失,
使用SMS技术的变体呼叫。在目标2中,我们将开发敏感映射的计算方法
的SMS读取和准确的变异呼叫在人类基因组的重复区域,目前是
排除在参考人类基因组的基准小变异调用集之外。在目标3中,我们将
利用目标1和2中的方法,对使用SMS测序的多个基因组进行变异识别
技术来对变体PSV进行编目,并利用该编目来改进读段映射和变体调用
在基因组的重复区域中短读测序的准确性。我们将在
强大和计算效率高的软件工具,并使用公开的长期可用的基准,
读取不同祖先的多个人类基因组的序列数据集。
英文摘要
Project Summary/Abstract
In this project, we propose to develop computational methods and tools for whole-genome haplotyping and
small variant calling using long-read sequencing technologies such as Pacific Biosciences and Oxford
Nanopore and linked-read technologies. Haplotype information is crucial for interpretation of genetic variation
in individual genomes, disease mapping, clinical genomics and several other analysis of human genetic
variation. The lack of phase or haplotype information in human genomes sequenced using short reads is a
major barrier in identifying disease associations with compound heterozygous mutations. More than 600 genes
overlap segmental duplications with high sequence identity and variants in more than 100 such genes have
been associated with rare Mendelian disorders and complex diseases including cancer. The inability to detect
variants with high accuracy in duplicated regions of the genome using short-read sequencing technologies
reduces the ability to identify disease causing mutations in medical genetics studies. In Aim 1, we will develop
a general computational method for long-read based diploid genotyping that will enable accurate haplotyping
for single nucleotide variants and short indels using long-read and linked-reads as well as accurate small
variant calling using SMS technologies. In Aim 2, we will develop computational methods for sensitive mapping
of SMS reads and accurate variant calling in repetitive regions of the human genome that are currently
excluded from benchmark small variant call sets for reference human genomes. Finally, in Aim 3, we will
leverage the methods from Aims 1 and 2 to perform variant calling on multiple genomes sequenced using SMS
technologies to catalog variant PSVs and leverage this catalog to improve read mapping and variant calling
accuracy of short-read sequencing in repetitive regions of the genome. We will implement the methods in
robust and computationally efficient software tools and benchmark their accuracy using publicly available long-
read sequence datasets for multiple human genomes of diverse ancestries.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
HapCUT2: A Method for Phasing Genomes Using Experimental Sequence Data.
HapCUT2:一种使用实验序列数据对基因组进行定相的方法。
DOI:
10.1007/978-1-0716-2819-5_9
发表时间:
2023
期刊:
Methods in molecular biology (Clifton, N.J.)
影响因子:
--
作者:
[Bansal,Vikas]
通讯作者:
Bansal,Vikas
DOI:
10.1016/j.xgen.2022.100128
发表时间:
2022-05
期刊:
CELL GENOMICS
影响因子:
--
作者:
[Wagner, Justin, Olson, Nathan D., Harris, Lindsay, Khan, Ziad, Farek, Jesse, Mahmoud, Medhat, Stankovic, Ana, Kovacevic, Vladimir, Yoo, Byunggil, Miller, Neil, Rosenfeld, Jeffrey A., Ni, Bohan, Zarate, Samantha, Kirsche, Melanie, Aganezov, Sergey, Schatz, Michael C., Narzisi, Giuseppe, Byrska-Bishop, Marta, Clarke, Wayne, Evani, Uday S., Markello, Charles, Shafin, Kishwar, Zhou, Xin, Sidow, Arend, Bansal, Vikas, Ebert, Peter, Marschall, Tobias, Lansdorp, Peter, Hanlon, Vincent, Mattsson, Carl-Adam, Barrio, Alvaro Martinez, Fiddes, Ian T., Xiao, Chunlin, Fungtammasan, Arkarachai, Chin, Chen-Shan, Wenger, Aaron M., Rowell, William J., Sedlazeck, Fritz J., Carroll, Andrew, Salit, Marc, Zook, Justin M.]
通讯作者:
Zook, Justin M.
Computational methods for variant calling and haplotyping using long-read sequencing technologies
-
批准号:10441522
-
项目类别:
-
资助金额:$38.16万
-
财政年份:2020
-
负责人:Vikas Bansal
-
依托单位:
Computational methods for variant calling and haplotyping using long-read sequencing technologies
-
批准号:10058104
-
项目类别:
-
资助金额:$38.17万
-
财政年份:2020
-
负责人:Vikas Bansal
-
依托单位:
Computational methods for variant calling and haplotyping using long-read sequencing technologies
-
批准号:10247821
-
项目类别:
-
资助金额:$38.23万
-
财政年份:2020
-
负责人:Vikas Bansal
-
依托单位:
Methods for detecting short indels from high-throughput sequence data
-
批准号:8706938
-
项目类别:
-
资助金额:$19.38万
-
财政年份:2013
-
负责人:Vikas Bansal
-
依托单位:
Methods for detecting short indels from high-throughput sequence data
-
批准号:8572023
-
项目类别:
-
资助金额:$28.26万
-
财政年份:2013
-
负责人:Vikas Bansal
-
依托单位:
海外基金