课题基金 / 基金详情

Functional data analysis methods for genomics and financial data

Functional data analysis methods for genomics and financial data
基因组学和金融数据的功能数据分析方法
批准号:
RGPIN-2020-05657
负责人:
Cremona, Marzia
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Cremona, Marzia的其他基金

相似基金

相关文献

中文摘要
翻译
最近,许多领域的数据量和复杂性都在迅速增加。高维数据在科学、工程和现代工业世界中已经变得无处不在。在学术界和工业界的许多研究领域,有效的数据分析和解释仍然是推进知识的瓶颈。因此,迫切需要新的统计可靠的技术,专门为高维数据的分析量身定制。我的研究将通过开发功能数据的统计工具来解决这一需求,即在连续统上变化的数据,可以用曲线表示。功能数据分析通常被用作对随时间变化的数据(如时间序列和纵向数据)建模的完全非参数方法。一个不太受欢迎但非常有前途的应用领域是所谓的“组学”科学(基因组学、表观基因组学)。现代高通量测序技术产生高维数据,可以用基因组曲线表示。我的研究计划的长期目标是开发新的统计方法来分析功能数据,并提供这些方法的计算效率实现。特别是,在接下来的几年里,我的研究将集中在发现功能母题的问题上,即典型的“形状”,可能沿着一组曲线反复出现几次,捕捉重要的局部特征。我最近开发了基于局部对齐的概率K-均值(probKMA),这是一种能够在一组曲线中识别K个候选功能基元的聚类方法。我的研究计划打算以这一最新发展为基础,并将追求几个方法和应用方向。我的第一个重点将是开发一种新的基于双聚类的motif发现技术。我还将扩展这些方法来检测单个曲线中的母题以及具有相似形状但长度不同的母题实例,并且我将对所发现的母题的统计显著性进行严格评估,这对于区分真实母题和随机出现在曲线背景中的母题至关重要。之后,我将把这些发展起来的方法应用到现实世界的问题中,尤其是对“组学”数据和资产价格的时间序列的分析。拟议的研究将产生方便用户和快速的软件,这些软件将免费提供给一般公众,并将能够从曲线中提取有关信息。这项研究的多学科性质将为参与我研究的学生提供广泛的培训机会。我的项目在几个STEM学科的边界上提供的培训,将有助于数据科学家的教育,他们对公司和大学都变得至关重要。
英文摘要
Many fields have recently seen a rapid increase in data volume and complexity. High-dimensional data have become pervasive in the sciences, in engineering, and in the modern industrial world. Effective data analysis and interpretation still represent the bottleneck in advancing knowledge in many areas of research, in both academia and industry. Hence, there is an urgency for novel statistically sound techniques, specifically tailored for the analysis of high-dimensional data. My research will address this need by developing statistical tools for functional data, i.e. data that vary over a continuum and can be represented as curves. Functional data analysis is often employed as a fully nonparametric approach for the modeling of data varying over time, such as time series and longitudinal data. A less popular but very promising application area is represented by the so-called "Omics" sciences (genomics, epigenomics.) - in which the modern high-throughput sequencing technologies produce high-dimensional data that can be represented as curves over the genome. The long-term objective of my research program is to develop novel statistical methods to analyze functional data and to provide computationally efficient implementations of such methods. In particular, my research over the next few years will focus on the problem of discovering functional motifs, i.e. typical "shapes" that may recur several times along and across a set of curves, capturing important local characteristics. I recently developed probabilistic K-mean with local alignment (probKMA), a clustering method able to identify K candidate functional motifs in a set of curves. My research program intends to build upon this recent development and will pursue several methodological and applied directions. My first focus will be to develop a new motif discovery technique based on biclustering. I will also extend these methods to detect motifs in a single curve as well as motifs whose instances have similar shapes but different lengths, and I will develop a rigorous assessment of the statistical significance of motifs found, that is critical to distinguish between real motifs and motifs that are randomly present in the background of curves. Afterward, I will apply the developed methods to the real-world problems that motivate them - in particular to the analysis of "Omics" data and time series of asset prices. The proposed research will result in user-friendly and fast software that will be freely available to the general public and will enable the extraction of relevant information from curves. The multidisciplinary nature of this research will translate into a broad training opportunity for students involved in my research. The training provided by my program, at the boundaries of several STEM disciplines, will contribute to the education of data scientists, who are becoming vital for both companies and universities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Functional data analysis methods for genomics and financial data
  • 批准号:
    RGPIN-2020-05657
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2021
  • 负责人:
    Cremona, Marzia
  • 依托单位:
Functional data analysis methods for genomics and financial data
  • 批准号:
    RGPIN-2020-05657
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2020
  • 负责人:
    Cremona, Marzia
  • 依托单位:
Functional data analysis methods for genomics and financial data
  • 批准号:
    DGECR-2020-00353
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2020
  • 负责人:
    Cremona, Marzia
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
  • 批准号:
    72101261
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    孙韬
  • 依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位: