课题基金 / 基金详情

AMC-SS: Markovian Embeddings for the Analysis and Computation of Patterns in non-Markovian Random Sequences

AMC-SS: Markovian Embeddings for the Analysis and Computation of Patterns in non-Markovian Random Sequences
AMC-SS:用于非马尔可夫随机序列中模式分析和计算的马尔可夫嵌入
批准号:
0805950
负责人:
Manuel Lladser
金额:
$0.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2013-06-30

项目摘要

项目成果

Manuel Lladser的其他基金

相似基金

相关文献

中文摘要
翻译
PI旨在开发新的工具,用于系统分析具有任意相关结构的随机序列中的模式。该项目考虑了与非马尔可夫序列中可能的非规则模式相关的理论和计算方面。PI以前的工作已经证明了最优马尔可夫结构的存在性和唯一性,以跟踪随机序列中与给定模式匹配的数量。在理论层面上,该项目旨在识别嵌入中的遍历马尔可夫结构,这些嵌入由于小扰动而导致非遍历行为,并使用这些遍历结构来表征非马尔可夫序列中模式匹配数量的渐近分布。并将非马尔可夫序列的最优马尔可夫嵌入的熵与原始序列的熵进行了比较。在计算层面上,PI的目标是应用生成函数技术结合马尔可夫嵌入来表征具有几种模式的匹配数量的联合渐近分布。PI还旨在量化在中等大小序列中与二项分布模式匹配的近似数量的误差。该项目的第二个目标是开发一种方法,用于表示非线性o.d.e.的解决方案的分支过程的符号规范。随机序列模式分析是计算生物学、安全系统、语音识别和文本挖掘等新兴领域的核心内容。字符的随机序列被用来模拟从书面文本到基因组序列再到审计文件的各种现象。在这些特殊的模式中,即,过多或过少的词,可以提供大量的洞察力或知识。例如,一个广泛使用的启发式方法是,DNA中代表性过高的模式可能是基因表达的关键,而代表性不足的模式可能会干扰这一过程。作为另一个例子,黑客在审计文件中留下他们入侵安全数据库的痕迹,并且可以使用不寻常的模式来警告潜在的安全漏洞。然而,如果没有一个真正的文本统计模型,人们就无法评估一个模式有多特殊。流行的模型不能适应大多数类型的文本中存在的长程相关性。例如,出现在书面文本或审计文件中的字符遵循语法规则,而在RNA序列中,由碱基配对诱导的回文结构传达了全基因组的相关性。不幸的是,现有的技术来评估一个模式是多么的特殊,不能系统地处理更现实的模型。此外,PI最近已经表明,广泛使用的范例,即在长文本中与模式匹配的数量近似为高斯分布,并不一定适用于存在长程相关性时。由于这些考虑,PI旨在开发新的定性和定量工具,以系统地解决在更现实的文本统计模型下出现的高度复杂的模式。
英文摘要
The PI aims to develop new tools for the systematic analysis of patterns in random sequences with an arbitrary correlation structure. The project considers theoretical and computational aspects associated with possibly non-regular patterns in non-Markovian sequences. Previous work by the PI has shown the existence and uniqueness of optimal Markovian structures to keep track of the number of matches with a given pattern in a random sequence. At a theoretical level, the project aims to identify ergodic Markovian structures in embeddings that result in non-ergodic behavior due to small perturbations, and to use these ergodic structures to characterize the asymptotic distribution of the number of matches with a pattern in a non-Markovian sequence. It also compares the entropy of the optimal Markovian embedding of a non-Markovian sequence with that of the original sequence. At a computational level, the PI aims to apply generating function techniques in conjunction with Markovian embeddings to characterize the joint asymptotic distribution of the number of matches with several patterns. The PI also aims to quantify the error in approximating the number of matches with a pattern in a sequence of a moderate size with that of a Binomial distribution. A secondary goal of the project is to develop a method for the symbolic specification of branching processes that will represent the solutions of non-linear o.d.e.'s probabilistically and to explore new connections between these and the analysis of patterns in non-Markovian sequences.The analysis of patterns in random sequences lies at the core of several emerging fields such as computational biology, security systems, speech recognition and text mining. Random sequences of characters are used to model diverse phenomena ranging from written text to genomic sequences to audit files. In these, exceptional patterns, i.e., over- or under-represented words, may provide a great deal of insight or knowledge. For instance, a widely used heuristic is that overrepresented patterns in DNA may be key for gene expression whereas underrepresented patterns may interfere with this process. As another example, hackers leave traces of their intrusions into secured databases in audit files and unusual patterns may be used to warn of potential security breaches. However, one cannot assess how truly exceptional a pattern is without a bona fide statistical model of the text in which it is immersed. The prevalent models cannot accommodate the long-range correlations present in most types of text. For example, the characters occurring in written text or audit files follow syntax rules and, in RNA sequences, palindromic structures induced by base-pairing convey genome-wide correlations. Unfortunately, the available techniques to assess how exceptional a pattern is, cannot systematically handle the more realistic models. Furthermore, the PI has recently shown that the widely used paradigm that the number of matches with a pattern in a long text is approximately Gaussian distributed does not necessarily apply when long-range correlations are present. Due to these considerations the PI aims to develop new qualitative and quantitative tools to address systematically the occurrence of highly complex patterns under more realistic statistical models of text.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
BIGDATA: F: Metric-space Positioning Systems for Symbolic Data Science
  • 批准号:
    1836914
  • 项目类别:
    Standard Grant
  • 资助金额:
    $61.06万
  • 财政年份:
    2018
  • 负责人:
    Manuel Lladser
  • 依托单位:
国内基金
海外基金
沙眼衣原体感染入侵新机制:通过T3SS效应蛋白CT622与ARP2相互作用介导宿主细胞质膜重塑
  • 批准号:
    2026JJ50576
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    雷文波
  • 依托单位:
叶绿体衰老过程中淀粉合成酶 SS4 的蛋白稳态调控及其对淀粉代谢的作用
  • 批准号:
    2026JJ60385
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    李晓平
  • 依托单位:
溶血素衍生肽SS-7的杀菌功能及作用机制研究
  • 批准号:
    JCZRQNB202600167
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
中医证型与RA-SS风险的关联机制:一项人群队列与代谢组学研究
  • 批准号:
    JCZRLH202601121
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位: