Composition Patterns in Nucleotide Sequences
Composition Patterns in Nucleotide Sequences
批准号:
0073081
负责人:
Gary Benson
金额:
$28.93万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2000
资助国家:
美国
项目状态:
已结题
起止时间:
2000-08-15 至 2004-03-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PI: Gary BensonProposal Number: 0073081Institution: Mt. Sinai School of MedicineProject SummaryThis project is an investigation of computational problems that arise for a new type of discrete pattern in DNA sequences, the composition pattern.Composition is a vector quantity describing the frequency of occurrence of each alphabet letter in a particular string. Let S be a string over E.Then, C (S) =(f1; f2; pH j _ j) is the composition of S, wherefore 2 _, fi is the fraction of the characters in S that are i. A composition pattern is a string P = r1r2 _ _ _ rp, where RI represents a composition region i.e. a substring of homogenous composition which differs from that of its surrounding regions. Note that the order of letters in RI is irrelevant, as it has no effect on the composition of RI. To date, algorithms which characterize DNA functional sites have concentrated primarily on identifying what this proposal terms position-specific patterns, such as the consensus sequence or themore flexible, but less specific weight matrix based pattern profile. Unfortunately, position-specific patterns are usually not selective enough to distinguish actual occurrences of a feature from false positives. Too often, when these patterns are used to search for unknown matches, one to several orders of magnitude more false positives than true positives are obtained. The composition pattern is a new approach which embraces an important physical property, the potential for variation in structural conformation (shape) of the DNA double helix, yet does so in the context of a type of discrete pattern which has apparently not been previously explored by the algorithmic community. DNA crystallization studies support the idea that certain dinucleotides base-steps confer specific types of flexibility. Further evidence is provided by studies of intrinsically curved and `kinkable' DNA. Based on these observations, it is suggested here that for conformational flexibility, the order of nucleotides in a sequence may be less important than the effect, which certain nucleotide or dinucleotide base-step biases impart on the sequence as a whole. In support of this assertion is an accumulating body of evidence of important DNA features whose unifying characteristic is composition bias rather than position-specific information. This research project encompasses algorithm development for three related problem areas which form the basis for understanding the functional importance of composition variation in nucleotide sequences and the detection of composition patterns. These areas are:Pattern matching. A composition pattern and sequence are given. Find all occurrences of the pattern in the sequence. Occurrences may be exact or approximate.Pattern detection. A sequence or set of sequences is given. Find all recurring composition patterns. Occurrences may be exact or approximate. The patterns are not specified or only partially specifiedSequence segmentation. A sequence is given. Partition it into statistically distinct regions of homogenous composition. These problems have theoretical interest in their own right, independent of biology and have received almost no attention from the algorithmic community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REU Site: Bioinformatics Research and Interdisciplinary Training Experience in Analysis and Interpretation of Information-Rich Biological Data Sets (REU-BRITE)
-
批准号:1949968
-
项目类别:Standard Grant
-
资助金额:$40.27万
-
财政年份:2020
-
负责人:Gary Benson
-
依托单位:
REU Site: Bioinformatics Research and Interdisciplinary Training Experience in Analysis and Interpretation of Information-Rich Biological Data Sets (REU-BRITE)
-
批准号:1559829
-
项目类别:Continuing Grant
-
资助金额:$29.23万
-
财政年份:2016
-
负责人:Gary Benson
-
依托单位:
III: Small: Bit-Parallel Algorithms for Sequence Alignment and Applications in Detecting Human Genetic Variation and Bacterial Strain Typing
-
批准号:1423022
-
项目类别:Continuing Grant
-
资助金额:$56.08万
-
财政年份:2014
-
负责人:Gary Benson
-
依托单位:
III:Small:Algorithms for Tandem Repeat Variant Discovery Using Next Generation Sequencing Data
-
批准号:1017621
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2010
-
负责人:Gary Benson
-
依托单位:
IGERT: Integrating Computational Science into Research in Biological Networks
-
批准号:0654108
-
项目类别:Continuing Grant
-
资助金额:$189.99万
-
财政年份:2007
-
负责人:Gary Benson
-
依托单位:
SEI(BIO): DNA Inverted Repeats: Sensitive Detection Methods and Research Database
-
批准号:0612153
-
项目类别:Standard Grant
-
资助金额:$66.9万
-
财政年份:2006
-
负责人:Gary Benson
-
依托单位:
Composition Patterns in Nucleotide Sequences
-
批准号:0413463
-
项目类别:Standard Grant
-
资助金额:$8.66万
-
财政年份:2003
-
负责人:Gary Benson
-
依托单位:
TRDB: A Multi-genome Database of Tandem Repeats
-
批准号:0413462
-
项目类别:Continuing Grant
-
资助金额:$70.5万
-
财政年份:2003
-
负责人:Gary Benson
-
依托单位:
TRDB: A Multi-genome Database of Tandem Repeats
-
批准号:0090789
-
项目类别:Continuing Grant
-
资助金额:$117.89万
-
财政年份:2001
-
负责人:Gary Benson
-
依托单位:
CAREER: Tandem Repeats: Sequence Comparison and Search Algorithms
-
批准号:9623532
-
项目类别:Continuing Grant
-
资助金额:$20.5万
-
财政年份:1996
-
负责人:Gary Benson
-
依托单位:
海外基金