课题基金 / 基金详情

AMC-SS: Markovian Embeddings for the Analysis and Computation of Patterns in non-Markovian Random Sequences

AMC-SS: Markovian Embeddings for the Analysis and Computation of Patterns in non-Markovian Random Sequences
AMC-SS:用于非马尔可夫随机序列中模式分析和计算的马尔可夫嵌入
批准号:
0805950
负责人:
Manuel Lladser
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2013-06-30

项目摘要

项目成果

Manuel Lladser的其他基金

相似基金

相关文献

中文摘要
翻译
PI的目标是开发新的工具,用于系统分析具有任意相关结构的随机序列中的模式。该项目考虑了与非马尔科夫序列中可能的非规则模式相关的理论和计算方面的问题。PI以前的工作已经证明了最优马尔可夫结构的存在和唯一性,以跟踪随机序列中与给定模式匹配的数量。在理论层面上,该项目的目标是识别嵌入中由于小扰动而导致非遍历行为的遍历马尔可夫结构,并使用这些遍历结构来表征与非马尔可夫序列中的模式匹配的数目的渐近分布。还比较了非马尔可夫序列和原始序列的最优马尔可夫嵌入的熵。在计算层面上,PI的目标是应用生成函数技术结合马尔可夫嵌入来表征具有多个模式的匹配数目的联合渐近分布。PI还旨在量化在用中等大小的序列中的模式与二项分布的模式近似匹配数量时的误差。该项目的第二个目标是开发一种用于分支过程的符号规范的方法,该方法将以概率方式表示非线性oD.E.S的解,并探索这些方法与非马尔科夫序列中的模式分析之间的新联系。随机序列中的模式分析位于几个新兴领域的核心,如计算生物学、安全系统、语音识别和文本挖掘。随机字符序列被用来模拟从书面文本到基因组序列再到审计文件的各种现象。在这些情况下,特殊模式,即过多或过少的词语,可能会提供大量的洞察力或知识。例如,一个广泛使用的启发式方法是,DNA中过度表达的模式可能是基因表达的关键,而表达不足的模式可能会干扰这一过程。作为另一个例子,黑客在审计文件中留下了入侵安全数据库的痕迹,并可能使用不寻常的模式来警告潜在的安全漏洞。然而,如果没有一个真正的文本统计模型,人们就无法评估一个模式有多特殊。流行的模型不能适应大多数文本类型中存在的长范围相关性。例如,出现在书面文本或审计文件中的字符遵循语法规则,在RNA序列中,由碱基配对诱导的回文结构传达了全基因组范围的相关性。不幸的是,现有的评估模式特殊程度的技术不能系统地处理更现实的模型。此外,PI最近表明,广泛使用的范式,即与长文本中的模式匹配的数量近似为高斯分布,并不一定适用于存在长范围相关性的情况。由于这些考虑,PI的目标是开发新的定性和定量工具,以便在更现实的文本统计模式下系统地处理高度复杂的模式的发生。
英文摘要
The PI aims to develop new tools for the systematic analysis of patterns in random sequences with an arbitrary correlation structure. The project considers theoretical and computational aspects associated with possibly non-regular patterns in non-Markovian sequences. Previous work by the PI has shown the existence and uniqueness of optimal Markovian structures to keep track of the number of matches with a given pattern in a random sequence. At a theoretical level, the project aims to identify ergodic Markovian structures in embeddings that result in non-ergodic behavior due to small perturbations, and to use these ergodic structures to characterize the asymptotic distribution of the number of matches with a pattern in a non-Markovian sequence. It also compares the entropy of the optimal Markovian embedding of a non-Markovian sequence with that of the original sequence. At a computational level, the PI aims to apply generating function techniques in conjunction with Markovian embeddings to characterize the joint asymptotic distribution of the number of matches with several patterns. The PI also aims to quantify the error in approximating the number of matches with a pattern in a sequence of a moderate size with that of a Binomial distribution. A secondary goal of the project is to develop a method for the symbolic specification of branching processes that will represent the solutions of non-linear o.d.e.'s probabilistically and to explore new connections between these and the analysis of patterns in non-Markovian sequences.The analysis of patterns in random sequences lies at the core of several emerging fields such as computational biology, security systems, speech recognition and text mining. Random sequences of characters are used to model diverse phenomena ranging from written text to genomic sequences to audit files. In these, exceptional patterns, i.e., over- or under-represented words, may provide a great deal of insight or knowledge. For instance, a widely used heuristic is that overrepresented patterns in DNA may be key for gene expression whereas underrepresented patterns may interfere with this process. As another example, hackers leave traces of their intrusions into secured databases in audit files and unusual patterns may be used to warn of potential security breaches. However, one cannot assess how truly exceptional a pattern is without a bona fide statistical model of the text in which it is immersed. The prevalent models cannot accommodate the long-range correlations present in most types of text. For example, the characters occurring in written text or audit files follow syntax rules and, in RNA sequences, palindromic structures induced by base-pairing convey genome-wide correlations. Unfortunately, the available techniques to assess how exceptional a pattern is, cannot systematically handle the more realistic models. Furthermore, the PI has recently shown that the widely used paradigm that the number of matches with a pattern in a long text is approximately Gaussian distributed does not necessarily apply when long-range correlations are present. Due to these considerations the PI aims to develop new qualitative and quantitative tools to address systematically the occurrence of highly complex patterns under more realistic statistical models of text.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: F: Metric-space Positioning Systems for Symbolic Data Science
  • 批准号:
    1836914
  • 项目类别:
    Standard Grant
  • 资助金额:
    $61.06万
  • 财政年份:
    2018
  • 负责人:
    Manuel Lladser
  • 依托单位:
国内基金
海外基金
沙眼衣原体感染入侵新机制:通过T3SS效应蛋白CT622与ARP2相互作用介导宿主细胞质膜重塑
  • 批准号:
    2026JJ50576
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    雷文波
  • 依托单位:
叶绿体衰老过程中淀粉合成酶 SS4 的蛋白稳态调控及其对淀粉代谢的作用
  • 批准号:
    2026JJ60385
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    李晓平
  • 依托单位:
溶血素衍生肽SS-7的杀菌功能及作用机制研究
  • 批准号:
    JCZRQNB202600167
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
中医证型与RA-SS风险的关联机制:一项人群队列与代谢组学研究
  • 批准号:
    JCZRLH202601121
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位: