课题基金 / 基金详情

BIGDATA: F: DKA: CSD: Iterative Crowdsourced Hypothesis Generation

BIGDATA: F: DKA: CSD: Iterative Crowdsourced Hypothesis Generation
BIGDATA:F:DKA:CSD:迭代众包假设生成
批准号:
1447634
负责人:
James Bagrow
金额:
$59.99万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-15 至 2020-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
建立因果关系-例如,吸烟导致肺癌-是科学研究中最具挑战性的方面之一。计算机擅长计算,但无法将因果关系与纯粹的相关性分开。另一方面,人类可以根据自己的经验做出合乎逻辑的结论,但在现代大数据时代,有太多的潜在关系需要人类手动检查。本研究旨在建立一个众包网络平台,利用感兴趣的非专家的知识(Hunch)和计算机的算法能力(Crunch)来发现和测试大规模数据中的因果关系。算法识别潜在的关系,并要求用户验证它们。此外,用户能够提出自己的假设,随后可以验证,创造一个加速科学发现的反馈循环。系统地发现因果关系的目标具有广泛的社会影响力,几乎任何有网络访问权限的人都可以直接参与这项科学研究。为了支持这一目标,研究人员正在开发新的统计方法,以确定人群建议的可观测数据的数据类型。例如,“工资”和“性别”是实值变量还是二元变量?最后,人群是一种相对有限的资源。为了有效地使用它,机器学习算法将识别相关网络中的哪些子结构最有可能是因果关系,然后将人群的努力集中在它们身上。这些有效的、适应性强的方法可以将因果关系组合成更大的链条,解释越来越多的因果关系。
英文摘要
Establishing causal relationships -- for example, that cigarette smoking causes lung cancer -- is one of the most challenging aspects of scientific research. Computers excel at calculation, but are unable to separate cause-and-effect from mere correlation. Humans, on the other hand, can make logical conclusions based on their experiences but, in the modern era of Big Data, there are far too many potential relationships for humans to manually examine. This research aims to build a crowdsourcing web platform to use the knowledge of interested non-experts (Hunch) and the algorithmic power of computers (Crunch) to discover and test causal relationships in large-scale data. Algorithms identify potential relationships and users are asked to validate them. Further, users are able to propose their own hypotheses that can subsequently be validated, creating an accelerating feedback loop of scientific discovery. The goal of systematically discovering causal relationships has the potential for broad societal impact, and virtually anyone with web access can participate directly in this scientific research.To support this goal, the researchers are developing novel statistical methods that determine the data types of crowd-suggested observables on the fly. For example, are 'wages' and 'gender' real-valued or binary variables? Finally, the crowd is a relatively limited resource. To use it efficiently, machine learning algorithms would identify which substructures in the correlational network are most likely to be causal, and then focus the crowd's efforts towards them. These efficient, adaptive methods allow causal relationships to be combined into larger chains that explain growing numbers of causes and effects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
HIV-1逆转录酶/整合酶双重抑制剂DKA-DAPYs的分子设计、合成及抗HIV活性研究
  • 批准号:
    21402148
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    25.0万元
  • 批准年份:
    2014
  • 负责人:
    古双喜
  • 依托单位: