Exhaustive Analysis of Microsatellite Loci in the 1000 Genomes Project
Exhaustive Analysis of Microsatellite Loci in the 1000 Genomes Project
批准号:
7882989
负责人:
HAROLD R GARNER
金额:
$26.4万
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-06-26 至 2012-04-30
关键词:
AreaBiological MarkersCharacteristicsDNADNA SequenceDataDiseaseElementsExhibitsFamilyForensic MedicineGenesGenetic PolymorphismGenomeGenomicsGoalsHumanHuman GenomeIndividualInstructionInternetLaboratoriesLengthMeasurementMeasuresMetadataMethodsMicrosatellite RepeatsModelingOntologyPaternity testingPrincipal InvestigatorProcessReagentRepetitive SequenceResearchResourcesRoleSingle Nucleotide PolymorphismTechniquesVariantcohortgenome sequencinggenome-widepressuretrenduser-friendly
中文摘要
描述(申请人提供):重复DNA,微卫星,一类表现出比单核苷酸多态高10,000倍的突变的基因组变异,由于在含有微卫星的基因座上缺乏数据而受到阻碍。这是到目前为止,随着1000基因组计划数据的出现。我们假设,一旦深入分析这些高变数基因座,就会对它们作为新的生物标志物和功能元件在基因组中的价值和作用产生新的认识。对1000个基因组计划队列中这些基因座变异性的基线测量将提供在计算和实验室中开发这些基因座所需的重要信息。这项拟议研究的主要目标是利用1000基因组计划提供的2,500套基因组序列来完成对大约700,000个微卫星座位的详尽分析和解释,以测量它们的大小、纯度和基序依赖的分布,然后用元数据(基因本体论、保守性等)覆盖这些数据,以创建一个资源,在那里我们和其他人可以探索微卫星多态在人类变异和疾病中的重要但被低估的作用。我们已经证明了所需的技术和有效的初步结果证实了可行性、价值和潜力。具体目标1)将所有1000个基因组计划的序列数据与包含座位的微卫星进行比对,以测量这些重复区域中的等位基因分布、多态率、特征和序列质量;检查和表征基序长度和家族组(AAT、AAAT、AATT等)。以寻找选择压力、偏差和全基因组趋势的证据;2)将分布与作为特定序列基序、基序大小、副本和纯度(是否存在任何SNPs)的函数的多态倾向估计模型进行比较,从而确定我们怀疑的任何一般复制或纠错机制偏差;3)用本体、保守和其他位置数据注释每个基因座,以确定任何过程、功能或疾病倾向的相关性;以及4)创建一个网络资源来发布我们的研究结果和其他试剂,以便其他人可以研究单个基因座或整个基因组的微卫星序列变异性。相关性(见说明):人类基因组包含超过500,000个区域,这些区域带有重复的DNA序列(例如CACACACACA),称为微卫星。它们变化无常,导致许多疾病,被用于法医/亲子鉴定,并可能改变我们的许多特征,但它们没有得到充分的研究和认识。1000基因组计划的数据使他们能够通过我们的方法进行全面的分析。
英文摘要
DESCRIPTION (provided by applicant): The study of repetitive DNA, microsatellites, a class of genomic variation which exhibits a 10,000 fold higher mutability than single nucleotide polymorphisms has been hampered by the lack of data at microsatellite- containing loci. That is, until now, with the emergence of data from the 1000 Genomes Project. We hypothesize that these hypervariable loci, once analyzed in depth will yield a new appreciation for their value and role in the genome as new biomarkers and functional elements. Baseline measurements of the variability at these loci in the substantial 1000 Genomes Project cohort will provide important information required to exploit these loci, both computationally and in the laboratory. The primary goal of the proposed research is to complete an exhaustive analysis and interpretation of the ~700,000 microsatellite loci using the -2,500 sets of genome sequence becoming available from the 1000 Genomes Project to measure their size, purity and motif dependent distributions and then overlay those data with metadata (gene ontologies, conservation and more) to create a resource where we an others can explore the significant, yet underappreciated role of microsatellite polymorphism in human variation and disease. We have demonstrated the techniques required and impactful preliminary results confirm feasibility and value and potential. Specific aims 1) align all 1000 Genomes Project sequence data to the microsatellite containing loci to measure the allelic distribution, polymorphism rate, characteristics, quality of the sequence in these repetitive regions; inspect and characterize groups of motif lengths and families (AAT,AAAT,AATT, etc.) to look for evidence for selection pressure, bias and genome wide trends; 2) compare the distributions with models for estimating polymorphism propensity as a function of specific sequence motifs, motif size, copies and purity (are there any SNPs), thus identifying any general replication or error correction mechanism bias, which we suspect; 3) annotate each locus with ontology, conservation and other positional data to identify any process, functional or disease propensity correlations; and 4) create a web resource to distribute our findings and other reagents derived from this study so others can investigate microsatellite sequence variability at individual loci or across the genome. RELEVANCE (See instructions): The human genome contains over 500,000 areas with repeated DNA sequence (e.g. CACACACACA) called microsatellites. They are extremely variable, cause numerous diseases, are used in forensics/ paternity testing and may alter many of our characteristics, but they are understudied and under- appreciated. The 1000 Genome Project data enables their thorough analysis en masse by our methods.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Exhaustive Analysis of Microsatellite Loci in the 1000 Genomes Project
-
批准号:8099068
-
项目类别:
-
资助金额:$26.5万
-
财政年份:2010
-
负责人:HAROLD R GARNER
-
依托单位:
Duplicate Article/Plagiarism Discovery
-
批准号:7911433
-
项目类别:
-
资助金额:$2.87万
-
财政年份:2009
-
负责人:HAROLD R GARNER
-
依托单位:
Computational Biology Core
-
批准号:7676466
-
项目类别:
-
资助金额:$43.32万
-
财政年份:2009
-
负责人:HAROLD R GARNER
-
依托单位:
Duplicate Article/Plagiarism Discovery
-
批准号:7850279
-
项目类别:
-
资助金额:$1.34万
-
财政年份:2009
-
负责人:HAROLD R GARNER
-
依托单位:
RCE Communications Center
-
批准号:7649861
-
项目类别:
-
资助金额:$18.69万
-
财政年份:2008
-
负责人:HAROLD R GARNER
-
依托单位:
Computational Biology Core
-
批准号:7649744
-
项目类别:
-
资助金额:$36.04万
-
财政年份:2008
-
负责人:HAROLD R GARNER
-
依托单位:
CD: Bioinformatics Core
-
批准号:7507395
-
项目类别:
-
资助金额:$9.1万
-
财政年份:2008
-
负责人:HAROLD R GARNER
-
依托单位:
Duplicate Article/Plagiarism Discovery
-
批准号:8121295
-
项目类别:
-
资助金额:$8.49万
-
财政年份:2007
-
负责人:HAROLD R GARNER
-
依托单位:
Duplicate Article/Plagiarism Discovery
-
批准号:7286877
-
项目类别:
-
资助金额:$28.6万
-
财政年份:2007
-
负责人:HAROLD R GARNER
-
依托单位:
Duplicate Article/Plagiarism Discovery
-
批准号:8121296
-
项目类别:
-
资助金额:$9.93万
-
财政年份:2007
-
负责人:HAROLD R GARNER
-
依托单位:
The Role of Microsatellite Instability in Cancer
-
批准号:6687904
-
项目类别:
-
资助金额:$25.94万
-
财政年份:2003
-
负责人:HAROLD R GARNER
-
依托单位:
The Role of Microsatellite Instability in Cancer
-
批准号:6901860
-
项目类别:
-
资助金额:$25.94万
-
财政年份:2003
-
负责人:HAROLD R GARNER
-
依托单位:
The Role of Microsatellite Instability in Cancer
-
批准号:6765911
-
项目类别:
-
资助金额:$25.94万
-
财政年份:2003
-
负责人:HAROLD R GARNER
-
依托单位:
CORE--AUTOMATION AND INSTRUMENTATION
-
批准号:6344947
-
项目类别:
-
资助金额:$126.32万
-
财政年份:2000
-
负责人:HAROLD R GARNER
-
依托单位:
GENOMICS AND PROTEOMICS OF CELL INJURY AND INFLAMMATION
-
批准号:6652508
-
项目类别:
-
资助金额:$348.97万
-
财政年份:2000
-
负责人:HAROLD R GARNER
-
依托单位:
SOFTWARE AND INSTRUMENTATION TO IDENTIFY CANCER GENES
-
批准号:6513581
-
项目类别:
-
资助金额:$54.53万
-
财政年份:1999
-
负责人:HAROLD R GARNER
-
依托单位:
SOFTWARE AND INSTRUMENTATION TO IDENTIFY CANCER GENES
-
批准号:6262509
-
项目类别:
-
资助金额:$62.58万
-
财政年份:1999
-
负责人:HAROLD R GARNER
-
依托单位:
SOFTWARE AND INSTRUMENTATION TO IDENTIFY CANCER GENES
-
批准号:6377237
-
项目类别:
-
资助金额:$62.87万
-
财政年份:1999
-
负责人:HAROLD R GARNER
-
依托单位:
SOFTWARE AND INSTRUMENTATION TO IDENTIFY CANCER GENES
-
批准号:2859814
-
项目类别:
-
资助金额:$14.68万
-
财政年份:1999
-
负责人:HAROLD R GARNER
-
依托单位:
CORE--AUTOMATION AND INSTRUMENTATION
-
批准号:6109077
-
项目类别:
-
资助金额:$126.32万
-
财政年份:1998
-
负责人:HAROLD R GARNER
-
依托单位:
海外基金