课题基金 / 基金详情

项目摘要

项目成果

Alison Motsinger-Reif的其他基金

相似基金

相关文献

中文摘要
翻译
在处理高维生物数据时,在进行适当的变量和统计推断方面存在许多挑战。我今年的方法开发工作重点是改进全基因组关联研究、高通量代谢组学数据和综合基因组学研究的计算和统计方法。 与我最近毕业的学生陶江博士一起,我们基于有置信度的伪零假设比例的估计下界,得到了稀疏群Lasso中正则化参数的一个新的上界。这一界限是通过应用来自单标记/变量分析的依赖或独立p值的经验分布来估计的,其中使用了二级显著性检验,即较高的批评统计量。套索中调谐参数的上界对应于假零假设比例的下界来确定。因此,调谐范围较窄,因为的上限较低。非零估计的最终判决将包含更多的变量,使得改进的GWA的权值高于或等于原始稀疏群Lasso。我们使用模拟实验和糖尿病临床试验中控制心血管风险的行动在血脂性状遗传学中的实际数据应用来证明我们的方法的性能。开发了一个R包。 目前仿冒方法的应用使用线性回归模型,并且只对模型函数中存在的变量进行变量选择。在与陶江博士和李元元博士的一个项目中,我们扩展了仿冒品用于机器学习的增强树,这是成功的,并被广泛应用于不需要模型函数先验知识的问题中。我们开发了一种新的策略,在没有先验模型拓扑知识的情况下,使用Boost树模型的仿冒方法进行变量选择。通过使用基于树的模型,我们将当前的仿冒方法扩展到无模型变量选择。我们测试并比较了这些方法和原始仿冒方法对I类错误和功率的控制能力。在模拟测试中,我们比较了树模型的重要性测试统计量的性质和性能。 几十年来,联合用药一直是癌症治疗的主流,并已被证明可以减少宿主的毒性,防止获得性耐药的发展。因此,开发预测药物协同作用的计算方法并指导实验设计,以发现合理的治疗组合是至关重要的。与我的学生马军一起,我们开发了一种新的深度学习方法,通过整合细胞系的基因表达谱和化学结构数据来预测协同药物组合。具体地说,我们使用主成分分析对化学描述符数据和基因表达数据进行降维。然后,我们通过神经网络传播低维数据,以预测药物协同作用的值。降维的使用在不损失精度的情况下大大减少了计算时间。 此外,我最近毕业的博士生马军博士,我们致力于开发一种方法,解决非线性剂量-反应关系中的挑战。非线性剂量-反应关系广泛存在于细胞、生化和生理过程中,这些过程受到不同程度的生物、化学或辐射应激的影响。非线性剂量-反应关系广泛存在于细胞、生化和生理过程中,这些过程受到不同程度的生物、化学或辐射应激的影响。因此,我们建议使用EA对一系列潜在反应模型的函数形式进行剂量-反应建模。这种新方法不仅可以拟合最常用的非线性剂量-反应模型(如指数模型、3参数、4参数和5参数Logistic模型),而且可以在不作模型假设的情况下选择最好的模型,这在高通量曲线拟合的情况下尤其有用。开发了实现该方法的R包。 建立在检测基因-环境相互作用的方法上的正在进行的项目目前正在进行中,使用差异QTL来确定检测基因-基因相互作用的单核苷酸多态的优先顺序。
英文摘要
There are a number of challenges in conducting proper variable and statistical inference in working across high dimensional biological data. My methods develop work this year has focused improving computational and statistical approaches for genome-wide association studies, high throughput metabolomics data, and integrative genomics studies. With my recently graduated student Dr. Tao Jiang, we developed a new upper bound of the regularization parameter in sparse group Lasso based on an estimated lower bound of the proportion of false null hypotheses with confidence. The bound is estimated by applying the empirical distribution of dependent or independent p-values from single marker/variable analysis, where a second-level significance testing, the higher criticism statistic, is used. An upper bound of the tuning parameter in Lasso is decided corresponding to the lower bound of the proportion of false null hypotheses. Thus, the tuning range is narrow since the upper bound of is lower. The final decision of non-zero estimates will contain more variables so that the power of modified GWAS is higher than or equal to the original sparse group Lasso. We demonstrate the performance of our method using both simulation experiments and a real data application in lipid trait genetics from the Action to Control Cardiovascular Risk in Diabetes clinical trial. An R package was developed. Current applications of knockoff methods use linear regression models and conduct variable selection only for variables existing in model functions. In a project with Dr. Tao Jiang, and with Dr. Yuanyuan Li, we extended the use of knockoffs for machine learning with boosted trees, which are successful and widely used in problems where no prior knowledge of model function is required. We developed a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. We extended the current knockoff method to model-free variable selection through the use of tree-based models. We tested and compared these methods with the original knockoff method regarding their ability to control type I errors and power. In simulation tests, we compared the properties and performance of importance test statistics of tree models. Combination drug therapy has been a mainstay of cancer treatment for decades and has been shown to reduce host toxicity and prevent the development of acquired drug resistance. Therefore, it is crucial to develop computational approaches to predict drug synergy and guide experimental design for the discovery of rational combinations for therapy. With my student Jun Ma, we developed a new deep learning approach to predict synergistic drug combinations by integrating gene expression profiles from cell lines and chemical structure data. Specifically, we use principal component analysis to reduce the dimensionality of the chemical descriptor data and gene expression data. We then propagate the low-dimensional data through a neural network to predict drug synergy values. The use of dimension reduction dramatically decreases the computation time, without losing accuracy. Additionally, my recently graduated PhD student Dr. Jun Ma we worked on developing an approach that addresses challenges in nonlinear dose-response relationships. Nonlinear dose-response relationships exist extensively in the cellular, biochemical, and physiologic processes that are affected by varying levels of biological, chemical, or radiation stress. Nonlinear dose-response relationships exist extensively in the cellular, biochemical, and physiologic processes that are affected by varying levels of biological, chemical, or radiation stress. Therefore, we propose the use of an EA for dose-response modeling for a range of potential response model functional forms. This new method can not only fit the most commonly used nonlinear dose-response models (eg, exponential models and 3-, 4-, and 5-parameter logistic models) but also select the best model if no model assumption is made, which is especially useful in the case of high-throughput curve fitting. An R package to implement the method was developed. Ongoing projects building onto methods for detecting gene-environment interactions are currently ongoing, using variance QTLs to prioritize single nucleotide polymorphisms for detecting gene-gene interactions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
海外基金