课题基金 / 基金详情

Collaborative Research: Selection Methods for Algebraic Design of Experiments

Collaborative Research: Selection Methods for Algebraic Design of Experiments
协作研究:实验代数设计的选择方法
批准号:
1720341
负责人:
Elena Dimitrova
金额:
$10.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-15 至 2019-07-31

项目摘要

项目成果

Elena Dimitrova的其他基金

相似基金

相关文献

中文摘要
翻译
数据科学已经成为一个重要的领域,它根据从医疗保健和住房等不同部门收集的数据做出决策。尽管由于手机应用程序、商户会员卡和社交媒体账户的出现,数据非常丰富,但更多的数据能否转化为更多的知识,这仍然是一个问题。此外,收集和存储可能会出现问题,特别是当数据敏感时,就像临床试验和基因实验经常出现的情况一样。选择信息丰富的数据对于创建能够可靠地预测未来实验结果的模型至关重要。关于必要数据量的结果很少发表,目前也没有指导方针来生成能够明确识别预测模型的特定数据集。作为发展完整理论的第一步,pi将专注于由有限值非线性多项式函数描述的模型。(例如,WedMD的症状检查器中的内部“函数”根据用户输入的症状返回医疗条件。)他们将构建具有单个相关多项式模型的最小数据集,并研究这些数据集的属性。从这些计算实验中,他们将建立适当的理论,设计算法,并生成代码,这些代码可以稍后开发成带有图形用户界面的软件。研究生将以适当的水平参与项目的每个组成部分。这样的经历将为他们提供硕士或博士论文的可能主题,并很可能激发他们在STEM学科的长期职业参与。理论结果将通过确定选择数据集以唯一识别模型的标准来推进实验设计,网络推理和有限动力系统领域。这些算法将作为实验人员确定识别感兴趣的网络结构所需的数据的指南。这种知识有可能大大减少由于数据太多而信息太少而造成的资源浪费。虽然现在是大数据时代,但更多的数据是否会转化为更多的知识,这仍然是一个问题。特别是在收集数据昂贵或耗时的情况下,如临床试验和生物分子实验经常出现的情况,选择信息丰富的数据的问题对于创建相关模型至关重要。有限状态多元多项式函数已被成功地用于从离散数据中建立复杂网络模型;然而,关于这些模型所需的数据量的结果很少,大多数只适用于布尔模型。目前还不知道哪些数据点可以明确地识别这种离散模型,因此,没有方法可以生成明确识别模型的特定数据集。pi将通过开发适当的理论、设计算法和生成可稍后内置到软件中的代码来解决数据的最小化和特异性问题,从而唯一地识别离散多项式模型。研究生将以适当的水平参与项目的每个组成部分。该项目将解决网络推理中的一些重要计算问题,并通过消除处理非线性多元多项式时产生的计算伪影的影响来改进实验设计和模型选择。理论结果将通过建立标准来选择数据集以唯一识别模型,从而推进实验设计和网络推理领域。所提出的工作还将增加多项式动力系统作为复杂网络模型的效用,通过建立最少量的数据来进行唯一模型识别。这些算法将作为实验人员确定识别感兴趣的网络结构所需的数据的指南。这种知识有可能大大减少所进行的实验数量,并消除产生没有什么内在价值的数据。
英文摘要
Data science has emerged as an important field for making decisions based on data collected from sectors as varied as healthcare and housing. Though data are plentiful, thanks to phone apps, merchant loyalty cards, and social media accounts, there is still a question of whether more data translates to more knowledge. Furthermore collection and storage can be problematic especially when data are sensitive, as it is often the case with clinical trials and genetic experiments. The problem of selecting information-rich data becomes crucial for creating models that can reliably predict the outcome of future experiments. Few results have been published on the amount of necessary data, and currently there are no guidelines for generating specific data sets which would unambiguously identify a predictive model. As a first step towards developing a complete theory, the PIs will focus on models described by finite-valued nonlinear polynomial functions. (For example, the internal "function" in WedMD's Symptom Checker returns medical conditions according to symptoms input by the user.) They will construct the smallest data sets that have a single associated polynomial model and study properties of such data sets. From these computational experiments, they will build the appropriate theory, design algorithms, and generate code that can be later developed into software complete with a graphical user interface. Graduate students will participate at the appropriate level of each component of the project. Such an experience will provide them possible topics for an MS or PhD dissertation and will very likely inspire a career-long involvement in the STEM disciplines. The theoretical results will advance the fields of design of experiments, network inference, and finite dynamical systems through the determination of criteria for selecting data sets to uniquely identify models. The algorithms will serve as a guide for experimentalists in determining the data that are needed to identify the structure of a network of interest. Such knowledge has the potential to drastically reduce wasted resources that arise from too much data with too little information.While this is the age of big data, there is still a question of whether more data translates to more knowledge. Particularly when collecting data is expensive or time consuming, as it is often the case with clinical trials and biomolecular experiments, the problem of selecting information-rich data becomes crucial for creating relevant models. Finite-state multivariate polynomial functions have successfully been used to model complex networks from discretized data; however, few results have been published on the amount of data necessary for such models, with the majority applying to Boolean models only. It is still unknown which data points explicitly identify such discrete models, and as a consequence, there are no methods for generating the specific data sets which would unambiguously identify the model. The PIs will address the issue of the minimality and specificity of data to uniquely identify discrete polynomial models by developing the appropriate theory, designing algorithms, and generating code that can be later built into software. Graduate students will participate at the appropriate level of each component of the project. This project will resolve some important computational issues in network inference and will improve experimental design and model selection by eliminating the effect of computational artifacts that arise when working with nonlinear multivariate polynomials. The theoretical results will advance the fields of design of experiments and network inference through the establishment of criteria to select data sets to uniquely identify models. The proposed work will also increase the utility of polynomial dynamical systems as models of complex networks by establishing the minimal amount of the data for unique model identification. The algorithms will serve as a guide for experimentalists in determining the data that are needed to identify the structure of a network of interest. Such knowledge has the potential to drastically reduce the number of experiments performed and to eliminate the generation of data with little intrinsic value.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Selection Methods for Algebraic Design of Experiments
Collaborative research: Data selection for unique model identification
  • 批准号:
    1419038
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2015
  • 负责人:
    Elena Dimitrova
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)