AMC-SS: Markovian Embeddings for the Analysis and Computation of Patterns in non-Markovian Random Sequences
AMC-SS: Markovian Embeddings for the Analysis and Computation of Patterns in non-Markovian Random Sequences
批准号:
0805950
负责人:
Manuel Lladser
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2013-06-30
中文摘要
PI的目标是开发新的工具,用于系统分析具有任意相关结构的随机序列中的模式。该项目考虑了与非马尔科夫序列中可能的非规则模式相关的理论和计算方面的问题。PI以前的工作已经证明了最优马尔可夫结构的存在和唯一性,以跟踪随机序列中与给定模式匹配的数量。在理论层面上,该项目的目标是识别嵌入中由于小扰动而导致非遍历行为的遍历马尔可夫结构,并使用这些遍历结构来表征与非马尔可夫序列中的模式匹配的数目的渐近分布。还比较了非马尔可夫序列和原始序列的最优马尔可夫嵌入的熵。在计算层面上,PI的目标是应用生成函数技术结合马尔可夫嵌入来表征具有多个模式的匹配数目的联合渐近分布。PI还旨在量化在用中等大小的序列中的模式与二项分布的模式近似匹配数量时的误差。该项目的第二个目标是开发一种用于分支过程的符号规范的方法,该方法将以概率方式表示非线性oD.E.S的解,并探索这些方法与非马尔科夫序列中的模式分析之间的新联系。随机序列中的模式分析位于几个新兴领域的核心,如计算生物学、安全系统、语音识别和文本挖掘。随机字符序列被用来模拟从书面文本到基因组序列再到审计文件的各种现象。在这些情况下,特殊模式,即过多或过少的词语,可能会提供大量的洞察力或知识。例如,一个广泛使用的启发式方法是,DNA中过度表达的模式可能是基因表达的关键,而表达不足的模式可能会干扰这一过程。作为另一个例子,黑客在审计文件中留下了入侵安全数据库的痕迹,并可能使用不寻常的模式来警告潜在的安全漏洞。然而,如果没有一个真正的文本统计模型,人们就无法评估一个模式有多特殊。流行的模型不能适应大多数文本类型中存在的长范围相关性。例如,出现在书面文本或审计文件中的字符遵循语法规则,在RNA序列中,由碱基配对诱导的回文结构传达了全基因组范围的相关性。不幸的是,现有的评估模式特殊程度的技术不能系统地处理更现实的模型。此外,PI最近表明,广泛使用的范式,即与长文本中的模式匹配的数量近似为高斯分布,并不一定适用于存在长范围相关性的情况。由于这些考虑,PI的目标是开发新的定性和定量工具,以便在更现实的文本统计模式下系统地处理高度复杂的模式的发生。
英文摘要
The PI aims to develop new tools for the systematic analysis of patterns in random sequences with an arbitrary correlation structure. The project considers theoretical and computational aspects associated with possibly non-regular patterns in non-Markovian sequences. Previous work by the PI has shown the existence and uniqueness of optimal Markovian structures to keep track of the number of matches with a given pattern in a random sequence. At a theoretical level, the project aims to identify ergodic Markovian structures in embeddings that result in non-ergodic behavior due to small perturbations, and to use these ergodic structures to characterize the asymptotic distribution of the number of matches with a pattern in a non-Markovian sequence. It also compares the entropy of the optimal Markovian embedding of a non-Markovian sequence with that of the original sequence. At a computational level, the PI aims to apply generating function techniques in conjunction with Markovian embeddings to characterize the joint asymptotic distribution of the number of matches with several patterns. The PI also aims to quantify the error in approximating the number of matches with a pattern in a sequence of a moderate size with that of a Binomial distribution. A secondary goal of the project is to develop a method for the symbolic specification of branching processes that will represent the solutions of non-linear o.d.e.'s probabilistically and to explore new connections between these and the analysis of patterns in non-Markovian sequences.The analysis of patterns in random sequences lies at the core of several emerging fields such as computational biology, security systems, speech recognition and text mining. Random sequences of characters are used to model diverse phenomena ranging from written text to genomic sequences to audit files. In these, exceptional patterns, i.e., over- or under-represented words, may provide a great deal of insight or knowledge. For instance, a widely used heuristic is that overrepresented patterns in DNA may be key for gene expression whereas underrepresented patterns may interfere with this process. As another example, hackers leave traces of their intrusions into secured databases in audit files and unusual patterns may be used to warn of potential security breaches. However, one cannot assess how truly exceptional a pattern is without a bona fide statistical model of the text in which it is immersed. The prevalent models cannot accommodate the long-range correlations present in most types of text. For example, the characters occurring in written text or audit files follow syntax rules and, in RNA sequences, palindromic structures induced by base-pairing convey genome-wide correlations. Unfortunately, the available techniques to assess how exceptional a pattern is, cannot systematically handle the more realistic models. Furthermore, the PI has recently shown that the widely used paradigm that the number of matches with a pattern in a long text is approximately Gaussian distributed does not necessarily apply when long-range correlations are present. Due to these considerations the PI aims to develop new qualitative and quantitative tools to address systematically the occurrence of highly complex patterns under more realistic statistical models of text.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: F: Metric-space Positioning Systems for Symbolic Data Science
-
批准号:1836914
-
项目类别:Standard Grant
-
资助金额:$61.06万
-
财政年份:2018
-
负责人:Manuel Lladser
-
依托单位:
国内基金
海外基金
登录
查看更多内容
沙眼衣原体感染入侵新机制:通过T3SS效应蛋白CT622与ARP2相互作用介导宿主细胞质膜重塑
-
批准号:2026JJ50576
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:雷文波
-
依托单位:
叶绿体衰老过程中淀粉合成酶 SS4 的蛋白稳态调控及其对淀粉代谢的作用
-
批准号:2026JJ60385
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:李晓平
-
依托单位:
溶血素衍生肽SS-7的杀菌功能及作用机制研究
-
批准号:JCZRQNB202600167
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
中医证型与RA-SS风险的关联机制:一项人群队列与代谢组学研究
-
批准号:JCZRLH202601121
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
T6SS新型辅助蛋白TagP在鲍曼不动杆菌致病中的作用机制研究
-
批准号:2026JJ81603
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:李艳冰
-
依托单位:
SS31肽通过AMPK/SIRT3通路调控氧化磷酸化在脓毒症心肌病中的作用及机制研究
-
批准号:JCZRLH202501261
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位:
NrtR通过调控T6SS参与溶藻弧菌竞争定
植的分子机制
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2025
-
负责人:蔡双虎
-
依托单位:
沙眼衣原体II型分泌系统(T2SS)次要假菌毛的鉴定及其分子组装机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:张时恒
-
依托单位:
SMARCB1乙酰化修饰调控驱动基因SS18-
SSX1增强子活性影响滑膜肉瘤侵袭转移
的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2025
-
负责人:齐妍
-
依托单位:
铜绿假单胞菌T6SS调控因子TsrF作用机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:15.0万元
-
批准年份:2024
-
负责人:曲久鑫
-
依托单位: