课题基金 / 基金详情

Social Science Gateway to TeraGrid

Social Science Gateway to TeraGrid
TeraGrid 的社会科学门户
批准号:
0922005
负责人:
Lars Vilhuber
金额:
$39.35万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-07-01 至 2013-06-30

项目摘要

项目成果

Lars Vilhuber的其他基金

相似基金

相关文献

中文摘要
翻译
该奖项由2009年《美国复苏和再投资法案》(公法111-5)资助。康奈尔大学的虚拟研究数据中心一直是人口普查局许多大型机密数据产品用户的成功研究支持工具,包括但不限于可通过人口普查研究数据中心网络访问的产品。超过200名计算用户和600名下载用户从VirtualRDC资源中受益。 他们的科学出版物引用了支持VirtualRDC开发的NSF赠款。 拟议的活动旨在保持这一支助网络的蓬勃发展。此外,大多数社会科学研究人员在希望利用大规模计算集群的能力时面临着巨大的障碍,特别是在使用新的、非常大的合成数据集时,这些数据集包含了关于人、工作和公司的前所未有的细节。拟议的活动旨在扩展VirtualRDC模型,以允许通过NSF赞助的TeraGrid资源支持万亿级社会科学计算。社会科学家使用最广泛的统计软件包,即,SAS、Stata和SPSS在TeraGrid本身或TeraGrid边界上的任何服务器上都不可用,无法快速连接到它。(提取、编辑和转换数据;将数据传输到计算位置;并进行分析),社会科学研究人员在接近TeraGrid上的高性能计算集群时,通常会受到这些步骤中至少一个步骤的限制。对于大多数数据编制和许多分析而言,缺乏标准的统计分析和数据编制软件包是一个严重的障碍。然而,典型的社会科学工作站或大学提供的计算基础设施没有资源来处理这些非常大的数据集。此外,社会科学工作站和大学提供的基础设施没有足够快的数据连接,无法将任何准备好的大型数据文件传输到TeraGrid进行处理。该项目旨在解决第一步和第二步中的瓶颈问题,重点扩大关键位置的资源,为社会科学提供一个非常有用的TeraGrid网关。该项目建立了一个社会科学TeraGrid网关,(i)允许研究人员使用他们的舒适级软件包执行数据准备步骤,加快数据准备阶段,(ii)在与TeraGrid快速连接的服务器上进行数据准备,从而大大加快数据传输过程。第三个瓶颈是TeraGrid缺乏社会统计数据包,本提案没有解决这一问题,因为这将需要资源,特别是许可证资源,这比我们的拟议预算大一个数量级。更广泛的影响:万亿级社会科学数据未得到充分利用。最初,严重的保密问题阻止了大多数研究人员访问这些数据。对解决大多数保密问题的项目进行了大量研究,并通过人口普查研究数据中心扩大了限制访问模式,从而开始解决这一利用不足的问题。现在,越来越多的以前保密的数据源正在进入公共领域,社会科学公共使用数据的数量再次急剧增加。该项目提出了一种解锁最近发布的数据源的方法,以允许研究界更广泛地访问。将实施以前提出但技术上不可行的研究战略,如超大规模的再结晶和合成。预期的爆炸式使用将导致许多社会科学的新成果。从运行社会科学TeraGrid网关中获得的知识将被利用并应用于未来的提案中,其中第三个确定的瓶颈--缺乏社会科学家熟悉的大规模计算资源软件--将得到解决。该提案的PI积极参与其他研究团队,这些团队正在推进此类提案的开发。这项提案的长期目标是,通过这项提案为研究界提供的工具将成为更大、更透明的机制的基石,使社会科学家能够轻松访问大规模计算设施。
英文摘要
This award is funded under the American Recovery and Reinvestment Act of 2009 (Public Law 111-5).The Virtual Research Data Center at Cornell University has been a successful research support tool for users of many of the Census Bureau large-scale confidential data products including, but not limited to, those that are accessible via the Census Research Data Center network. Over 200 computational users and 600 download users have benefited from the VirtualRDC resources. Their scientific publications cite the NSF grants that supported the development of the VirtualRDC. The proposed activity seeks to keep this support network flourishing. In addition, most social science researchers face substantial hurdles when they wish to harness the power of large-scale computational clusters, in particular when using new, very large synthetic data sets with their unprecedented detail on people, jobs, and firms. The proposed activity seeks to extend the VirtualRDC model to allow support of tera-scale social science computing via the NSF-sponsored TeraGrid resources. The most widespread statistical software packages used by social scientists, i.e., SAS, Stata, and SPSS, are not available on the TeraGrid itself or on any of the servers at the borders of the TeraGrid with fast connections to it. When viewing the problem through the lens of the typical data-driven research process (extract, edit and transform data; transfer data to a computational location; and perform analysis) social science researchers are typically constrained in at least one of these steps when approaching the high-performance computing clusters on the TeraGrid. For most data preparation, and for much analysis, the lack of standard statistical analysis and data preparation software packages is a serious impediment. However, the typical social scientist workstation or university-provided computational infrastructure does not have the resources to handle these very large data sets. Furthermore, the social science workstation and the university-provided infrastructure do not have sufficiently fast data connectivity to transfer any large prepared data files to the TeraGrid for processing there. This project aims to remedy bottlenecks in the first and second steps, with a focused expansion of resources at a critical location resulting in a highly useful gateway to the TeraGrid for the social sciences. The project builds a social science TeraGrid gateway that (i) allows researchers to perform the data preparation step using their comfort-level software packages, speeding up the data preparation phase, and (ii) do so on servers that have a fast connection to the TeraGrid, thus greatly speeding up the data-transfer process. The third bottleneck absence of social statistics packages on the TeraGrid is not addressed by this proposal, since it would require resources, in particular licensing resources, an order of magnitude larger than our proposed budget. This step is left to future proposals.Broader impacts: Tera-scale social science data are underutilized. Initially, serious confidentiality issues prevented most researchers from accessing these data. Significant research effort on projects that solve most of these confidentiality issues in combination with an expansion of the restricted-access model via Census Research Data Centers has begun to address this underutilization. Now that an increasing number of previously confidential data sources are finding their way into the public domain, the quantity of social science public-use data is once-again expanding dramatically. This project proposes a method of unlocking those recently released data sources to allow much broader access by the research community. Research strategies such as very large scale resampling and synthesis, which were previously proposed but not technically feasible, will be implemented. The expected explosion of use will lead to new results in a multitude of social sciences. The knowledge gained from running the Social Science TeraGrid Gateway will be leveraged and applied to future proposals in which the third identified bottleneck the absence of familiar software for social scientists on large-scale computing resources will be addressed. The PIs on this proposal are actively involved with other research teams that are moving forward with the development of such proposals. The long-term goal of this proposal is that the tools put together for the research community through this proposal will be the building blocks for bigger, and more transparent mechanisms, for granting social scientists easy access to large-scale computational facilities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Elements: TRAnsparency CErtified (TRACE): Trusting Computational Research Without Repeating It
  • 批准号:
    2209629
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2022
  • 负责人:
    Lars Vilhuber
  • 依托单位:
Conferences on Reproducibility and Replicability in Economics and the Social Sciences (CRRESS)
  • 批准号:
    2217493
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.0万
  • 财政年份:
    2022
  • 负责人:
    Lars Vilhuber
  • 依托单位:
RCN: Coordination of the NSF-Census Research Network
  • 批准号:
    1507241
  • 项目类别:
    Standard Grant
  • 资助金额:
    $46.29万
  • 财政年份:
    2014
  • 负责人:
    Lars Vilhuber
  • 依托单位:
RCN: Coordination of the NSF-Census Research Network
国内基金
海外基金
科学传播类:基于大科学装置“中国天眼”的AI for science新型科普平台建设
  • 批准号:
    T2241020
  • 项目类别:
    专项项目
  • 资助金额:
    10.00万元
  • 批准年份:
    2022
  • 负责人:
    毛睿
  • 依托单位:
SCIENCE CHINA: Earth Sciences
SCIENCE CHINA Chemistry
基于e-Science的民族信息资源融合与语义检索研究
  • 批准号:
    61262071
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    46.0万元
  • 批准年份:
    2012
  • 负责人:
    甘健侯
  • 依托单位: