Data analysis tools for leveraging massive public data to improve hypothesis-driven research
Data analysis tools for leveraging massive public data to improve hypothesis-driven research
批准号:
10330636
负责人:
Jeffrey T. Leek
金额:
$2.68万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-04-01 至 2022-05-14
关键词:
AcuteBiologicalCollectionCommunitiesComputer softwareCongressesDataData AnalysesData SourcesDevelopmentDiseaseGenerationsHeartIndividualMeasurementMedicalMethodsMolecularNational Institute of General Medical SciencesPatientsProcessReproducibilityResearchResearch PersonnelRunningSample SizeSamplingSourceSpeedStatistical Data InterpretationStatistical MethodsTechnologyTrainingUnited StatesUnited States National Institutes of HealthWorkcostcrowdsourcingdata resourcedesignexperimental studyfollow-uphigh throughput technologyimprovedlarge scale datapublic repositoryrecruittool
中文摘要
项目总结
科学fic结果存在可重复性和可复制性的危机。这场危机是一种日益严重的
科学fic和大众媒体都对此表示关注。这场危机如此严重,以至于美国国会目前
研究科学fic过程的重复性。这场危机的核心是一系列问题,包括
样本量小,研究力度不足,数据分析师培训不足,无法直接利用之前的
使用高通量技术对较小的、假设驱动的实验进行统计分析的结果。
技术的进步极大地降低了收集高通量分子的成本和难度。
数据。大量原始数据越来越多地公开可用,但通常被合并到个人
NIGMS和其他调查人员在特别基础上进行的分析。与此同时,运营一家设计好的、
假设驱动的研究并没有随着技术的进步而以同样的速度减少。它仍然很贵,
识别、招募、收集和跟踪样本,即使高通量测量本身是便宜的。
尽管可用的公共数据数量惊人,但执行统计推断仍然是常见的做法
在这些假设驱动的实验中,逐个研究,仅间接包括先前的数据、估计和
结果。因此,这些研究的fi碱基可能是高度可变的、不可靠的或不可复制的。我们的团队专注于
关于开发统计方法、数据资源、软件和培训,使研究人员能够借鉴
来自公共存储库、大规模数据生成项目和众包数据的经验优势
在个体、假设驱动的研究中改进推理。我们建议在我们的工作的基础上开发
统计数据来源、方法、软件和培训,以促进和加快我们的生物和
医疗合作者。其结果将是一个研究社区,可以利用已经公开的数据
美国国立卫生研究院花费巨资收集数据,以提高功率、减少所需的样本大小并改进
许多新的假说推动了对发育和紊乱的分子研究。
英文摘要
Project summary
There is a crisis of reproducibility and replicability of scientific results. This crisis is an increasing source of
concern both in the scientific and popular press. The crisis is so acute that the United States Congress is currently
investigating reproducibility of the scientific process. At the heart of this crisis is a collection of problems including
small-sample sizes, under-powered studies, under-trained data analysts and an inability to directly leverage prior
results in the statistical analysis of smaller, hypothesis-driven experiments using high-throughput technologies.
Advances in technology have dramatically reduced the cost and difficulty of collecting high-throughput molecular
data. Large collections of raw data are increasingly publicly available but are usually incorporated into individual
analyses by NIGMS and other investigators on an ad-hoc basis. Meanwhile, the other costs of running a designed,
hypothesis-driven study have not decreased at the same speed with technological advances. It is still expensive to
identify, recruit, collect, and follow up samples even if the high-throughput measurements themselves are cheap.
Despite the incredible amount of available public data, it is still common practice to perform statistical inference
in these hypothesis-driven experiments study-by-study, only indirectly including previous data, estimates, and
results. So findings from these studies may be highly variable, unreliable, or unreplicable. Our group has focused
on developing statistical methods, data resources, and software and training that allow researchers to borrow
strength empirically from public repositories, large-scale data generation projects, and crowd-sourced data to
improve inference in individual, hypothesis driven studies. We propose to build on our work in developing
statistical data sources, methods, software and training that facilitate and speed the work of our biological and
medical collaborators. The result will be a research community that can take advantage of public data already
collected at a large cost to the NIH to improve power, reduce required sample sizes, and improve replication in
many new hypothesis driven molecular studies of development and disorder.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data analysis tools for leveraging massive public data to improve hypothesis-driven research
-
批准号:10598130
-
项目类别:
-
资助金额:$42.82万
-
财政年份:2022
-
负责人:Jeffrey T. Leek
-
依托单位:
Data analysis tools for leveraging massive public data to improve hypothesis-driven research
-
批准号:10654376
-
项目类别:
-
资助金额:$40.44万
-
财政年份:2022
-
负责人:Jeffrey T. Leek
-
依托单位:
A massive study of data science to address the scientific reproducibility crisis
-
批准号:9100338
-
项目类别:
-
资助金额:$36.45万
-
财政年份:2016
-
负责人:Jeffrey T. Leek
-
依托单位:
A massive study of data science to address the scientific reproducibility crisis
-
批准号:9244046
-
项目类别:
-
资助金额:$36.45万
-
财政年份:2016
-
负责人:Jeffrey T. Leek
-
依托单位:
Statistical models for biological and technical variation in RNA sequencing
-
批准号:8593469
-
项目类别:
-
资助金额:$30.78万
-
财政年份:2013
-
负责人:Jeffrey T. Leek
-
依托单位:
Statistical models for biological and technical variation in RNA sequencing
-
批准号:9264553
-
项目类别:
-
资助金额:$30.78万
-
财政年份:2013
-
负责人:Jeffrey T. Leek
-
依托单位:
Statistical models for biological and technical variation in RNA sequencing
-
批准号:8722575
-
项目类别:
-
资助金额:$30.78万
-
财政年份:2013
-
负责人:Jeffrey T. Leek
-
依托单位:
Core B
-
批准号:9978143
-
项目类别:
-
资助金额:$12.62万
-
财政年份:2011
-
负责人:Jeffrey T. Leek
-
依托单位:
Core B
-
批准号:9304366
-
项目类别:
-
资助金额:$13.82万
-
财政年份:--
-
负责人:Jeffrey T. Leek
-
依托单位:
Core B
-
批准号:9759993
-
项目类别:
-
资助金额:$13.01万
-
财政年份:--
-
负责人:Jeffrey T. Leek
-
依托单位:
海外基金