课题基金 / 基金详情

Methods development for "Omics" data

Methods development for "Omics" data
“组学”数据的方法开发
批准号:
10008737
负责人:
Alison Motsinger-Reif
金额:
$35.34万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Alison Motsinger-Reif的其他基金

相似基金

相关文献

中文摘要
翻译
与我的研究生Tao Jiang一起,我们一直在研究最小绝对收缩和选择算子(Lasso)回归的扩展,以解决样本量与数据维度相比有限时的变量选择和建模问题。这是全基因组关联研究中的常见现象。 我们开发了一个新的上界稀疏组Lasso的正则化参数的基础上估计的置信度(1-)的假零假设的比例的下界。通过应用来自单个标记/变量分析的依赖或独立p值的经验分布来估计界限,其中使用二级显著性检验,即更高的批评统计量。Lasso中的调整参数的上界对应于虚假零假设比例的下界。因此,调谐范围是窄的,因为上界较低。非零估计的最终决定(例如,GWAS中的显著基因座)将包含更多变量,使得修改的GWAS的功效高于或等于原始稀疏组Lasso。研究了真实回归模型中变量间的不同相关水平。 我们证明了我们的方法的性能,使用模拟实验和真实的数据应用在脂质性状遗传学的行动,以控制糖尿病心血管风险(雅阁)的临床试验。 Tao Jiang的另一个项目专注于机器学习方法,用于在下一代测序分析中检测同种污染。 在这个项目中,我们开发了一种机器方法,依靠支持向量机和相应的软件来检测可能发生的相同物种的污染,因为实验室质量控制问题,或从法医应用中的混合样品。 这是第一套可以直接在.VCF文件上工作的工具,该方法可以识别肿瘤和正常细胞的混合物,以防止假阳性。 目前正在制定扩展,以估计肿瘤与正常样本的百分比,并估计法医应用中的个体数量。 在与丹尼斯·福切斯博士、他的学生杰里米·阿什和博士后梅莱恩·库内曼、我的前博士后丹尼尔·罗特洛夫博士的合作中,我一直在研究将化学结构信息整合到代谢组学分析中的方法。 开发预测性和透明的方法来分析患者队列中的代谢物谱对于理解触发或调节感兴趣的性状的事件(例如,疾病进展、药物代谢、化学风险评估)。然而,代谢物的化学结构仍然很少用于建立这些性状-代谢物关系的统计建模工作流程中。在此,我们提出了一种新的化学信息学为基础的方法,能够识别预测,解释,和可重复的性状-代谢物的关系。作为概念验证,我们利用先前发表的病例研究,包括非小细胞肺癌(NSCLC)腺癌患者和健康对照的代谢产物谱。通过使用计算的分子描述符和患者代谢物浓度曲线来表征每个结构注释的代谢物,我们表明这些互补特征增强了对与癌症相关的关键代谢物的识别和理解。最终,我们建立了多代谢物分类模型,用于使用通过化学聚类基于高结构相似性识别的特定代谢物组来评估患者的癌症状态。 此外,与NCSU和查佩尔山的Jung-Ying Tzeng博士和其他几位合作者一起,我们开发了一种测试单个罕见变异关联的方法。 由于罕见变异对人类复杂疾病的病因学贡献,其对遗传关联研究的兴趣越来越大。由于突变事件的罕见性,罕见变异体通常在总体水平上进行分析。虽然聚合分析改善了全局水平信号的检测,但它们不能在变体集中精确定位因果变体。要在局部级别上执行推理,需要额外的信息,例如,通常需要生物注释来提高罕见变异的信息含量。在观察到重要的变体可能在功能域上聚集在一起之后,我们提出了一种蛋白质结构引导的局部测试(POINT),使用结构引导的信号聚集来提供变体特异性关联信息。POINT在一个核机器框架下构建,通过以数据自适应的方式从三维蛋白质空间中的相邻变体中借用信息来执行局部关联测试。除了仅仅提供有希望的变体列表之外,POINT为每个变体分配p值以允许变体排名和优先级排序 目前正在进行的项目建立在检测基因-基因相互作用的方法上,使用方差QTL来优先考虑单核苷酸多态性以检测基因-基因相互作用。
英文摘要
With my graduate student Tao Jiang, we have been working on an extension of least absolute shrinkage and selection operator (Lasso) regression to address variable selection and modeling when sample sizes are limited compared to the data dimension. This is a common phenomenon in genome wide association studies. We developed a new upper bound of the regularization parameter in sparse group Lasso based on an estimated lower bound of the proportion of false null hypotheses with confidence (1-). The bound is estimated by applying the empirical distribution of dependent or independent p-values from single marker/variable analysis, where a second-level significance testing, the higher criticism statistic is used. An upper bound of tuning parameter in Lasso, , is decided corresponding to the lower bound of the proportion of false null hypotheses. Thus, the tuning range is narrow since the upper bound of is lower. The final decision of non-zero estimates (e.g., significant loci in GWAS) will contain more variables so that the power of modified GWAS is higher than or equal to the original sparse group Lasso. Different correlation levels among variables in true regression models are also studied. We demonstrate the performance of our method using both simulation experiments and a real data application in lipid trait genetics from the Action to Control Cardiovascular Risk in Diabetes (ACCORD) clinical trial. Another project with Tao Jiang is focused on a machine learning approach for detecting same-species contamination in next generation sequencing analysis. In this project, we have developed a machine method relying on support vector machines and corresponding software to detect same species contamination that can occur because of laboratory quality control issues, or from mixed samples in forensic application. This is the first set of tools that can work directly on the .VCF files, and the approach recognizes a mixture of tumor and normal cells to prevent false positives. Current extensions are being worked out to estimate the percentage of tumor vs. normal samples, and to estimate the number of individuals within forensic application. In collaboration with Dr. Denis Fourches, his student Jeremy Ash and post doc Melaine Kuenemann, my former postdoc Dr. Daniel Rotroff, I have been working on methods to integrate chemical structure information into metabolomics analyses. Developing predictive and transparent approaches to the analysis of metabolite profiles across patient cohorts is of critical importance for understanding the events that trigger or modulate traits of interest (e.g., disease progression, drug metabolism, chemical risk assessment). However, metabolites' chemical structures are still rarely used in the statistical modeling workflows that establish these trait-metabolite relationships. Herein, we present a novel cheminformatics-based approach capable of identifying predictive, interpretable, and reproducible trait-metabolite relationships. As a proof-of-concept, we utilize a previously published case study consisting of metabolite profiles from non-small-cell lung cancer (NSCLC) adenocarcinoma patients and healthy controls. By characterizing each structurally annotated metabolite using both computed molecular descriptors and patient metabolite concentration profiles, we show that these complementary features enhance the identification and understanding of key metabolites associated with cancer. Ultimately, we built multi-metabolite classification models for assessing patients' cancer status using specific groups of metabolites identified based on high structural similarity through chemical clustering. Additionally, with Dr. Jung-Ying Tzeng and several other collaborators at NCSU and UNC Chapel Hill, we have developed an approach for testing for associations of single rare variants. Rare variants are of increasing interest to genetic association studies because of their etiological contributions to human complex diseases. Due to the rarity of the mutant events, rare variants are routinely analyzed on an aggregate level. While aggregation analyses improve the detection of global-level signal, they are not able to pinpoint causal variants within a variant set. To perform inference on a localized level, additional information, e.g., biological annotation, is often needed to boost the information content of a rare variant. Following the observation that important variants are likely to cluster together on functional domains, we propose a protein structure guided local test (POINT) to provide variant-specific association information using structure-guided aggregation of signal. Constructed under a kernel machine framework, POINT performs local association testing by borrowing information from neighboring variants in the 3-dimensional protein space in a data-adaptive fashion. Besides merely providing a list of promising variants, POINT assigns each variant a p-value to permit variant ranking and prioritization Ongoing projects building onto methods for detecting gene-gene interactions are currently ongoing, using variance QTLs to prioritize single nucleotide polymorphisms for detecting gene-gene interactions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
Genetic Basis of Genotype-by-Environment Interactions Underlying Physiological Mo
海外基金