课题基金 / 基金详情

Functional data analysis methods for genomics and financial data

Functional data analysis methods for genomics and financial data
基因组学和金融数据的功能数据分析方法
批准号:
RGPIN-2020-05657
负责人:
Cremona, Marzia
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Cremona, Marzia的其他基金

相似基金

相关文献

中文摘要
翻译
最近,许多领域的数据量和复杂性都在迅速增加。高维数据在科学、工程和现代工业世界中已经变得无处不在。有效的数据分析和解释仍然是学术界和工业界在许多研究领域推进知识的瓶颈。因此,迫切需要专门为分析高维数据量身定做的新颖的可靠的统计技术。 我的研究将通过开发功能数据的统计工具来满足这一需求,即在连续统中变化的数据,可以用曲线表示。函数数据分析通常被用作一种完全非参数的方法来建模随时间变化的数据,例如时间序列和纵向数据。一个不太受欢迎但非常有希望的应用领域是所谓的“全基因组学”科学(基因组学、表观基因组学),在这些科学中,现代高通量测序技术产生的高维数据可以表示为基因组上的曲线。 我的研究计划的长期目标是开发新的统计方法来分析函数数据,并提供这种方法的计算效率实现。特别是,在接下来的几年里,我的研究将集中在发现功能主题的问题上,即典型的“形状”,可能会沿着一组曲线重复多次,并跨越一组曲线,捕捉到重要的局部特征。我最近开发了一种概率K-均值局部比对(概率KMA),这是一种能够在一组曲线中识别K个候选功能基序的聚类方法。我的研究计划打算建立在这一最新发展的基础上,并将追求几个方法论和应用方向。我的第一个重点将是开发一种基于双聚类的新的基序发现技术。我还将扩展这些方法来检测单个曲线中的基序以及其实例具有相似形状但不同长度的基序,并将对所发现的基序的统计意义进行严格评估,这对于区分真实基序和随机出现在曲线背景中的基序至关重要。然后,我将把开发的方法应用于激励他们的现实世界问题,特别是对“Omics”数据和资产价格时间序列的分析。 拟议的研究将产生用户友好和快速的软件,将向公众免费提供,并将能够从曲线中提取相关信息。这项研究的多学科性质将转化为参与我研究的学生的广泛培训机会。我的项目在几个STEM学科的边界上提供的培训,将有助于数据科学家的教育,他们对公司和大学都变得至关重要。
英文摘要
Many fields have recently seen a rapid increase in data volume and complexity. High-dimensional data have become pervasive in the sciences, in engineering, and in the modern industrial world. Effective data analysis and interpretation still represent the bottleneck in advancing knowledge in many areas of research, in both academia and industry. Hence, there is an urgency for novel statistically sound techniques, specifically tailored for the analysis of high-dimensional data. My research will address this need by developing statistical tools for functional data, i.e. data that vary over a continuum and can be represented as curves. Functional data analysis is often employed as a fully nonparametric approach for the modeling of data varying over time, such as time series and longitudinal data. A less popular but very promising application area is represented by the so-called “Omics” sciences (genomics, epigenomics) in which the modern high-throughput sequencing technologies produce high-dimensional data that can be represented as curves over the genome. The long-term objective of my research program is to develop novel statistical methods to analyze functional data and to provide computationally efficient implementations of such methods. In particular, my research over the next few years will focus on the problem of discovering functional motifs, i.e. typical “shapes” that may recur several times along and across a set of curves, capturing important local characteristics. I recently developed probabilistic K-mean with local alignment (probKMA), a clustering method able to identify K candidate functional motifs in a set of curves. My research program intends to build upon this recent development and will pursue several methodological and applied directions. My first focus will be to develop a new motif discovery technique based on biclustering. I will also extend these methods to detect motifs in a single curve as well as motifs whose instances have similar shapes but different lengths, and I will develop a rigorous assessment of the statistical significance of motifs found, that is critical to distinguish between real motifs and motifs that are randomly present in the background of curves. Afterward, I will apply the developed methods to the real-world problems that motivate them in particular to the analysis of “Omics” data and time series of asset prices. The proposed research will result in user-friendly and fast software that will be freely available to the general public and will enable the extraction of relevant information from curves. The multidisciplinary nature of this research will translate into a broad training opportunity for students involved in my research. The training provided by my program, at the boundaries of several STEM disciplines, will contribute to the education of data scientists, who are becoming vital for both companies and universities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Functional data analysis methods for genomics and financial data
  • 批准号:
    RGPIN-2020-05657
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2022
  • 负责人:
    Cremona, Marzia
  • 依托单位:
Functional data analysis methods for genomics and financial data
  • 批准号:
    RGPIN-2020-05657
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2021
  • 负责人:
    Cremona, Marzia
  • 依托单位:
Functional data analysis methods for genomics and financial data
  • 批准号:
    DGECR-2020-00353
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2020
  • 负责人:
    Cremona, Marzia
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
  • 批准号:
    72101261
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    孙韬
  • 依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位: