Collaborative Research: Selection Methods for Algebraic Design of Experiments
Collaborative Research: Selection Methods for Algebraic Design of Experiments
批准号:
1720335
负责人:
Brandilyn Stigler
金额:
$10.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-15 至 2021-10-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Data science has emerged as an important field for making decisions based on data collected from sectors as varied as healthcare and housing. Though data are plentiful, thanks to phone apps, merchant loyalty cards, and social media accounts, there is still a question of whether more data translates to more knowledge. Furthermore collection and storage can be problematic especially when data are sensitive, as it is often the case with clinical trials and genetic experiments. The problem of selecting information-rich data becomes crucial for creating models that can reliably predict the outcome of future experiments. Few results have been published on the amount of necessary data, and currently there are no guidelines for generating specific data sets which would unambiguously identify a predictive model. As a first step towards developing a complete theory, the PIs will focus on models described by finite-valued nonlinear polynomial functions. (For example, the internal 'function' in WedMD's Symptom Checker returns medical conditions according to symptoms input by the user.) They will construct the smallest data sets that have a single associated polynomial model and study properties of such data sets. From these computational experiments, they will build the appropriate theory, design algorithms, and generate code that can be later developed into software complete with a graphical user interface. Graduate students will participate at the appropriate level of each component of the project. Such an experience will provide them possible topics for an MS or PhD dissertation and will very likely inspire a career-long involvement in the STEM disciplines. The theoretical results will advance the fields of design of experiments, network inference, and finite dynamical systems through the determination of criteria for selecting data sets to uniquely identify models. The algorithms will serve as a guide for experimentalists in determining the data that are needed to identify the structure of a network of interest. Such knowledge has the potential to drastically reduce wasted resources that arise from too much data with too little information.While this is the age of big data, there is still a question of whether more data translates to more knowledge. Particularly when collecting data is expensive or time consuming, as it is often the case with clinical trials and biomolecular experiments, the problem of selecting information-rich data becomes crucial for creating relevant models. Finite-state multivariate polynomial functions have successfully been used to model complex networks from discretized data; however, few results have been published on the amount of data necessary for such models, with the majority applying to Boolean models only. It is still unknown which data points explicitly identify such discrete models, and as a consequence, there are no methods for generating the specific data sets which would unambiguously identify the model. The PIs will address the issue of the minimality and specificity of data to uniquely identify discrete polynomial models by developing the appropriate theory, designing algorithms, and generating code that can be later built into software. Graduate students will participate at the appropriate level of each component of the project. This project will resolve some important computational issues in network inference and will improve experimental design and model selection by eliminating the effect of computational artifacts that arise when working with nonlinear multivariate polynomials. The theoretical results will advance the fields of design of experiments and network inference through the establishment of criteria to select data sets to uniquely identify models. The proposed work will also increase the utility of polynomial dynamical systems as models of complex networks by establishing the minimal amount of the data for unique model identification. The algorithms will serve as a guide for experimentalists in determining the data that are needed to identify the structure of a network of interest. Such knowledge has the potential to drastically reduce the number of experiments performed and to eliminate the generation of data with little intrinsic value.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
Small Gröbner fans of ideals of points
点理想的格罗布纳小粉丝
DOI:
10.1142/s0219498820500875
发表时间:
2020
期刊:
Journal of Algebra and Its Applications
影响因子:
0.8
作者:
[Dimitrova, Elena, He, Qijun, Robbiano, Lorenzo, Stigler, Brandilyn]
通讯作者:
Stigler, Brandilyn
Algebraic model selection and experimental design in biological data science
生物数据科学中的代数模型选择和实验设计
DOI:
10.1016/j.aam.2021.102282
发表时间:
2022
期刊:
Advances in Applied Mathematics
影响因子:
1.1
作者:
[Dimitrova, Elena, Hu, Jingzhen, Liang, Qingzhong, Stigler, Brandilyn, Zhang, Anyu]
通讯作者:
Zhang, Anyu
The Number of Gröbner Bases in Finite Fields
有限域中格罗布纳碱基的数量
DOI:
10.1007/978-3-030-42687-3_9
发表时间:
2020
期刊:
Association for Women in Mathematics series
影响因子:
--
作者:
[Stigler, Brandilyn, Zhang, Anyu]
通讯作者:
Zhang, Anyu
Collaborative Research: Data selection for unique model identification
-
批准号:1419023
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2015
-
负责人:Brandilyn Stigler
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: